kaeru memory
Provides Markdown export of any initiative, producing Obsidian-friendly snapshots that can be used in an Obsidian vault.
蛙 kaeru
kaeru is a cross-agent cognitive engine for LLM agents — a typed graph that agents think in, plus a recollection layer for long-term ideas and outcomes. Local-first for each agent, with an optional shared cloud tier so a whole team of agents and people build on one another's memory.
Designed for multi-session, multi-agent continuity: when an agent opens a project, it has full context of what was being thought about, can follow provenance chains, can consolidate outcomes into stable long-term knowledge, and can pull in what the rest of the team has shared.
Inspired by the LLM-wiki pattern (Karpathy, gist 442a6bf555914893e9891c11519de94f), the bi-temporal knowledge graph approach of Graphiti / Zep, the curator-driven knowledge engine of Cognee, and the reasoning-based hierarchical-summary navigation pattern from PageIndex. Two-tier design grounded in the hippocampus / cortex split.
Name: 蛙 (kaeru, "frog"; homophonic with 帰る "to return" and 変える "to change") — the agent that returns, recalls, and reshapes.

A kaeru vault rendered by kaeru-viz — each project a constellation of stars around its core, sized by memory layer, with ochre cross-project bridges, reasoning-chain replay, and a time-lapse of how the knowledge grew. Hover a node to trace its neighbours; click to pin them.
Overview

kaeru is built around a typed property graph stored in CozoDB. Two tiers, biological analogy:
cognitive (operational / hippocampus) — high-velocity working graph where the agent actively thinks: episodes, scratch, drafts, hypotheses, experiments, audit events.
recollection (archival / cortex) — settled ideas, outcomes, summaries, references; mostly read.
Every node and edge is bi-temporal — the substrate stores assertion / retraction history natively, so time-travel queries are out of the box and conflict resolution is non-destructive (the old version is invalidated, not deleted).
Per-initiative subgraphs through a junction-relation pattern: one substrate, many initiatives, multi-membership. An agent working on project A asks "what was I doing here last time?" and gets an answer scoped to A. The same node can belong to several initiatives at once.
kaeru is a facilitator, not an enforcer. The curator API exposes ~70 primitives (awake, recall, drill, claim, synthesise, at, history, settle, …) as available tools. The agent and user choose when to invoke them; the daemon hints but doesn't block.
Related MCP server: Engram
Features
Two-tier graph — operational (cognitive / hippocampus) for active thinking; archival (recollection / cortex) for settled knowledge.
settlepromotes across the boundary, preserving provenance.Bi-temporal — native assertion / retraction history.
atreads a node in full as it is now or as-of any past moment; conflicts are non-destructive (the old version is invalidated, not deleted).Per-initiative scoping + layered re-entry — one substrate, many projects.
awakerestores a project's working set by memory layer (Core → Hot → Warm) and surfaces the archival cortex (settled knowledge) alongside it, so durable facts re-enter every session;surfacereaches the archived Cold / Frozen on demand.Reasoning chains —
chainsaves the load-bearing weighted path between two nodes as a recallable trail with an agent-authored summary;whyreads a trail — give it a chain for its ordered steps, or any node to reach the chain it belongs to. Duplicates are folded at creation, andrechainrefreshes a trail after the graph changes (re-links, re-weights).Self-maintenance —
reflectcomputes a tidy-up work-list: orphan nodes to link, overdue tasks, chains gone stale, settled work to promote into cortex, and shared/cloud items whose rebalancing is escalated to the user. The agent calls it at the end of a piece of work; layer tidying runs on its own in the background (hygiene).Role slots —
slotgives an initiative a role held by exactly one live node (handoff,entrypoint,queue). Filling it archives the previous holder to Cold and linkssupersedesfrom the new holder to it, so a project cannot drift into three "current" handoffs. Nothing is deleted; the predecessor stays reachable throughat/surface.It tells you when you are behind — the daemon asks GitHub once a day whether a newer kaeru exists, and if so the next
awakeopens with what the gap is costing: the running version, how many releases back it is, that release's own one-line reason, and the upgrade command for the way this binary was actually installed (installer script,.mcpbbundle, or a source checkout). The agent is the updater and asks before running it. One public GET, no identifiers, nothing about your vault;KAERU_MCP_UPDATE_CHECK=0turns it off entirely. A warning in a daemon log nobody reads is the same as no warning at all.Automatic hygiene — a background pass keeps a project's layers honest: old unreferenced journal entries move to Cold, untouched unreferenced Core nodes drop one step, heavily-referenced nodes rise one step. It triggers on accumulation (writes, Core growth, elapsed time) rather than on a schedule, runs off the reactor in batches so it never stalls a tool call, and only ever changes a node's layer — every move reverses with one
layercall. On by default since 0.7.3: setKAERU_MCP_HYGIENE_ENABLE=0to turn it off (embedding viakaeru-rig? hygiene stays opt-in there — a library must not sweep a vault its embedder didn't ask it to;KaeruMemory::with_hygiene()starts it).hygiene <initiative>shows exactly what the next pass would move without moving it — worth a look on a vault you care about.Cross-agent sharing — local-first by default; an optional
kaeru-cloudtier lets a trusted team share settled knowledge through two safety gates (initiative policy, which names both whether an initiative may leave and which clouds it may reach, plus a deterministic secret guard).cloud_recallsearches the shared tier,unsharewithdraws a mistake, and with several clouds configured an unnamed call is refused rather than routed to a default. See Local & cloud.Structural recall — exact name lookup, typed
walk/drill/trace,between, FTS fuzzy fallback. Every read carries when each node was asserted, and a read that stops short points at the verb that goes further: an excerpt namesat, a changed node nameshistory, a node inside a saved trail nameswhy. A name that fails to resolve says which of three things happened — it lives in another initiative, something close exists, or it is nowhere.Read-back on re-entry —
awakerestores what was touched and also what is still owed: open tasks with overdue ones first, claims awaiting a verdict, and the saved trails with their summaries. On a shared initiative it says the local answer may be incomplete.Initiative management —
rename/deletean initiative (locally or team-wide), orattacha node to another initiative to repair fragmentation after the fact.Markdown export — Obsidian-friendly snapshot of any initiative.
Architecture Notes
Substrate is CozoDB with RocksDB backend; bi-temporal
Validityis native to the substrate, not bolted on.Edges carry operational semantics — each edge type is something the curator API responds to.
derived_frompowers provenance and explainability;contradictstriggers a non-destructiveunder_reviewflow;supersedesretracts the previous version through the bi-temporal substrate, and runs one way — from the replacement to what it replaced. Edges are not just associations.audit_eventis a first-class node type — every mutation writes an audit node, so changes to memory themselves become reasoning surface for the agent. Substrate-level history (Validity) and operational audit (audit-event nodes) stay separate: the substrate tracks what was, the audit nodes track who did it and why.Per-initiative scope through junction relations rather than column filtering — RocksDB prefix-scan gives O(log n + k) on the active initiative.
Retrieval is structural-first — explicit name lookup, typed graph traversal, summary views. Cozo FTS for fuzzy fallback when an exact name is forgotten. No vector/embedding layer today: Cozo supports HNSW, but kaeru wires none of it — a vector fallback is possible future work, not a current feature.
Two-tier with explicit promotion —
settle <name>moves a node that stopped changing into the archival tier as a deliberate, logged operation, carrying its name, body and manual tags over unchanged (unsettlemirrors it back). Provenance (derived_from) survives the tier boundary.Single binary, embedded substrate — the substrate runs in-process with the daemon; the vault is a local file tree, not a service. The daemon speaks MCP to the agent and, optionally, HTTP to a
kaeru-cloud. Vault on disk under a platform-specific default (Linux$XDG_DATA_HOME/kaeru, macOS~/Library/Application Support/ai.lamantin.kaeru, Windows%LOCALAPPDATA%\ai.lamantin.kaeru); override withKAERU_VAULT_PATH.
Layout
kaeru/
├── Cargo.toml ← workspace root
├── kaeru-core/ ← library: substrate, schema, primitives
├── kaeru-mcp/ ← binary `kaeru-mcp`: Model Context Protocol server (the agent's surface)
├── kaeru-cloud/ ← binary `kaeru-cloud`: shared cloud tier (Axum REST over kaeru-core)
├── kaeru-rig/ ← library: `rig` Tools that give a rig agent kaeru memory
└── skills/
└── kaeru-skill/ ← portable agent skill (Claude Code / etc.)kaeru-rig is the rig framework adapter — the full curator verb set as discrete rig Tools over an embedded Arc<Store>, so a rig agent reads and writes one vault. Future: kaeru-langchain (Python bridge), not yet started.
Install
Pre-1.0 alpha. Substrate schema may change between minor versions — export to markdown if you need to keep notes around.
See QUICK_START.md for source builds, MCP daemon setup, and the re-entry ritual.
Quick tour (MCP tools)
# See what projects exist:
initiatives
# Re-entry ritual: process state + epistemic state.
awake (initiative: "auth-rewrite")
overview (initiative: "auth-rewrite")
# Capture (auto-named):
jot (initiative: "auth-rewrite", body: "noticed token expiry differs across platforms")
# Fuzzy lookup when you forgot the exact name:
search (initiative: "auth-rewrite", query: "expiry")
# Drill into something:
drill (initiative: "auth-rewrite", name: "noticed-token-expiry-differs-across-...")
# Hypothesis cycle:
# You normally reach memory AFTER the check has run, so record the verdict
# with the claim — one call, and the status lands where every read can see it:
claim (initiative: "auth-rewrite", text: "platform-aware policy is correct",
verdict: "supported", by: "<evidence>")
# Or the prospective path, when the question is genuinely still open:
claim (initiative: "auth-rewrite", text: "platform-aware policy is correct", about: "<node>")
evidence (initiative: "auth-rewrite", hypothesis: "<hyp>", method: "compared iOS / Android TTL")
confirm (initiative: "auth-rewrite", hypothesis: "<hyp>", by: "<experiment>")
# Time-travel:
at (initiative: "auth-rewrite", name: "<name>", when: "5m")
history (initiative: "auth-rewrite", name: "<name>")
# Knowledge chains — strongest weighted path between two nodes
# (link the load-bearing edges with strong: true first):
path (initiative: "auth-rewrite", from: "<a>", to: "<b>") # preview the trail
chain (initiative: "auth-rewrite", from: "<a>", to: "<b>") # save it as a recallable chain
why (initiative: "auth-rewrite", name_or_id: "<a>") # read the trail a node sits in
# Roles — one live node each. Writing the next handoff archives the last one:
slot (initiative: "auth-rewrite", slot: "handoff", name: "handoff-tuesday")
slots (initiative: "auth-rewrite")
# What the next hygiene pass would move, and why:
hygiene (initiative: "auth-rewrite")
# Snapshot to an Obsidian-friendly markdown vault:
export (initiative: "auth-rewrite", path: "/tmp/auth-snapshot")Building without network access
Some machines cannot reach crates.io. The dependencies for each release are
vendored into a separate repository, LamantinAI/kaeru-vendor,
one tag per release:
git checkout v0.7.3
./contrib/offline/fetch-vendor.sh # ~600 MB, needs network once
# carry the whole directory across, then:
cargo build --release --offline -p kaeru-mcpRust alone is not enough on either side of that line — cozo builds RocksDB from
C++ and zstd/lz4 from C, so a C++ toolchain and libclang are needed whether or
not you are offline. See docs/offline-build.md for the
per-platform requirements and the Windows line-ending trap that fails every
checksum at once.
Connecting to an MCP-aware agent
kaeru-mcp is a long-lived HTTP service: one daemon per machine owns the substrate, any number of agent sessions (Claude Code, Opencode, Cursor, …) connect concurrently. This is intentional — RocksDB is single-writer, so a stdio MCP that forks a subprocess per session would hit lock contention. See kaeru-mcp/README.md for systemd / launchd unit templates and the full HTTP config.
Run the daemon (or set up the systemd user unit from contrib/install/):
kaeru-mcp # foreground, Ctrl-C to stopThen point your agent at it:
Claude Code:
claude mcp add --transport http kaeru http://127.0.0.1:9876/mcp— seeskills/kaeru-skill/for the portable system-prompt rules.Check memory before asking (Claude Code and Codex):
contrib/hooks/kaeru-first/is a harness hook for the moment the agent is about to ask the user something — if kaeru was not read recently, it sends the agent to search first, once, and then asks it to capture the user's answer. Stdlib Python, fails open. See its README.Opencode:
bash contrib/opencode/install-opencode.sh— wires the daemon, dropsAGENTS.kaeru.mdrules into~/.config/opencode/, and installs/kaeru//lesson//recallslash commands. Designed to coexist with your existing OSS-model provider config (Qwen / DeepSeek / GLM / Ollama). Seecontrib/opencode/README.md.Cursor and other runtimes: paste the body of
skills/kaeru-skill/SKILL.mdinto your agent's rules / system-prompt section. For MCP-aware clients the daemon URL above works directly.
After restart the agent sees tools like awake, drill, claim, at natively. Each tool takes an optional initiative parameter.
Local & cloud — sharing memory across a team
kaeru runs local-first: your vault lives on your machine and nothing leaves it by default. A second, optional tier — kaeru-cloud — is a shared store for a trusted group (a team, a family). Each initiative carries a sticky share_policy:
private(default) — nothing ever leaves; personal projects.team— nodes you explicitly marksharedmay sync to the cloud.
Sharing is never automatic and passes two gates: the initiative policy, and a deterministic pre-share secret guard that blocks anything looking like an API key, token, or private key. The guard is silent on clean content and only interrupts on a real hit.
Verbs (over MCP): policy (mark an initiative team), share (push a node), cloud_recall (see what the team has), pull (bring a shared node into your local graph), link_cloud / cloud_links (reference cloud nodes without copying), and sync_review (batch-review still-local nodes). Capture verbs (episode / jot / cite) take visibility: shared to capture-and-share in one call.
A single daemon can also reach several named clouds (e.g. a family and a work cloud) via a clouds.toml file; the cloud verbs then take an optional cloud: <name> and soft links remember which cloud they point at. See kaeru-mcp/README.md for the file format.
Per-user / per-org isolation (multi-tenant) is a future addition; today each cloud is one shared space scoped by initiative. See kaeru-cloud/README.md.
Roadmap
Initiative onboarding — a generated "what is this initiative, and what's in it" briefing the first time an agent enters a project, in the spirit of the
importguide — an instruction over the collected context, not graph machinery.PostgreSQL backend — a server-mode substrate alongside the embedded RocksDB default.
Multi-tenant + isolation — per-user / per-org separation in the shared cloud.
Cross-initiative links — opt-in edges and recall that traverse between initiatives. Per-initiative scoping is deliberate, not a gap — and
attachalready gives a node membership in several initiatives at once; this adds boundary-crossing edges/recall on top of it.
Status
Pre-1.0. Implemented and covered by a green test suite: the substrate and curator API, memory layers with layered re-entry — an operational working set plus an archival cortex that re-enters every session (awake / surface), bi-temporal time-travel with assertion time surfaced in every read (at / history), per-initiative scoping with rename / delete / attach (additive multi-membership), knowledge chains (weighted shortest-path, agent-authored summaries, creation-time dedup, rechain), a reflect maintenance pass, forward-only schema migrations, the MCP server, the shared kaeru-cloud tier (sharing, recall, soft links, sync-review) including multi-cloud, the kaeru-rig framework adapter (full curator toolset as rig Tools), and markdown export. What still needs hardening:
Multi-tenant. The cloud is one shared space scoped by initiative; per-user / per-org isolation isn't built yet.
MCP concurrency. Concurrent sessions share one
Store; the per-call initiative scope is serialized throughStore::scoped, so two sessions can't corrupt each other's scope. What remains is ordering — when an agent batch-fires async calls, a read can still land before a not-yet-applied write.Whole-second
Validityresolution. Two opposing mutations on the same node/edge within one second (e.g.linkthen an immediateunlink, or aforgetright after a write) resolve ambiguously. Interactive use is fine — human pacing always crosses the boundary; the test suite sleeps between such operations.Audit events aren't attached to the initiative junction yet (export filters them by
affected_refsintersection — a working workaround).Rig adapter shipped, LangChain not yet.
kaeru-riggives a rig agent the full memory toolset; a Python / LangChain bridge is still to come.Migrations are forward-only and add-only. A
migration_journalruns schema additions (new relations / columns) on open; there is no down-migration or destructive-change path yet. Snapshot viaexportbefore a major upgrade is still prudent pre-1.0.
Contributing
Discussion and design feedback through issues. PRs welcome on the open items above and anywhere the agent-facing surface feels rough — the verb taxonomy (awake, drill, claim, flag, settle, …) is meant to map to natural agent thinking, not just expose graph operations.
License
Business Source License 1.1 — see LICENSE, with a plain-language
summary in LICENSING.md.
kaeru is source-available: read it, modify it, run it in production, and embed it as a component or tool inside your own systems and products — including what you sell. The one thing you may not do is operate kaeru itself as a hosted / managed / SaaS platform for third parties. Each released version converts to the Apache License 2.0 four years after it ships. Contributions are accepted under the Contributor License Agreement.
Versions up to and including v0.7.0 were released under the MIT License and remain available under it.
Available Tools
71 toolsatA
Read a node IN FULL — every field plus the complete, untruncated body. drill / search / recall only show short excerpts; reach for at when you need a node's whole content. Optional when time-travels to a past moment (unix seconds, RFC-3339, or 5m / 2h ago); omit it for the node as it is now.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name or UUIDv7 id. With `when` set, it resolves as of that moment, so a node retracted since then is still reachable — by id, or by the name it carried at that time. | |
| when | No | Optional moment to time-travel to — Unix seconds, RFC-3339 (`2026-05-06T12:00:00Z`), or duration suffix (`5m`, `2h`, `3d` = "ago"). Omit to read the node as it is NOW. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden: it discloses the full-vs-excerpt behavioral distinction, time-travel semantics, and that a node retracted since the target moment is still reachable (reinforced in the schema). It implies a safe read via 'Read' but never explicitly confirms read-only/non-destructive status, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the defining distinction (full content) before the alternative-selection rule and the optional parameter. Every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema, the description usefully explains what is returned and the time-travel behavior. The notable omission is the purpose of the `initiative` parameter, which is left entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% — `name` and `when` are documented in the schema, including the same time formats the description restates. The description adds little beyond the schema for those two, and the third parameter, `initiative`, is undocumented in both places, so the description cannot fully compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and scope — 'Read a node IN FULL — every field plus the complete, untruncated body' — leaving no ambiguity about what the tool returns. It explicitly contrasts itself with siblings ('drill / search / recall only show short excerpts'), so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternatives (drill, search, recall) and the exact selecting condition ('when you need a node's whole content'), and explains the selective use of the optional `when` parameter. The when-to-use decision is fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attachA
Add a node to another initiative (additive multi-membership) — repair initiative fragmentation by giving a node captured under the wrong or a stale initiative a second home, without moving or copying it (same id, edges, history). The node is resolved across all initiatives. Idempotent. Local only. For a whole initiative rather than a few nodes, merge_initiative is the one-step version.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Target initiative to add the node to. Additive — the node keeps every initiative it already belongs to. | |
| node | Yes | Node name or UUIDv7 id to attach. Resolved across all initiatives, so the node may currently live under a different one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden and does well: it discloses that the operation is additive, non-destructive ('without moving or copying it — same id, edges, history'), idempotent, and local-only. It omits return/confirmation details and any permission requirements, but the safety and side-effect profile is unusually complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and the additive/no-copy guarantee are front-loaded, and the sibling routing comes last. It is a dense single paragraph rather than a compact sentence, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must stand alone; it covers side effects, idempotency, scope (local only), and the alternative. Confirmation/return behavior is unstated, a minor gap for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'node' and 'to'. The description reinforces the additive semantics and cross-initiative resolution, but adds little beyond what the parameter descriptions already state — baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (add a node to another initiative) with the key qualifier 'additive multi-membership' and explicitly contrasts with the sibling merge_initiative. An agent can distinguish it from merge_initiative and other initiative tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case (repair fragmentation, node captured under the wrong or stale initiative) and routes to the alternative ('For a whole initiative rather than a few nodes, merge_initiative is the one-step version'). When-to-use and when-to-use-something-else are both present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
awakeA
Restore session context: pinned set, recent episodes (24h), open reviews, plus the read-back of unfinished work — open tasks (overdue first), claims awaiting a verdict, and the saved reasoning trails. Run this when re-entering a project.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | Optional initiative to scope the operation to. When omitted, reads are cross-initiative; mutations end up un-tagged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses return content and ordering ("overdue first", 24h window), which is useful, but never states that this is a non-mutating read or describes permissions/side effects — a notable gap given the generic mutation note inherited in the parameter text.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, verb front-loaded, then a compact enumeration, then the trigger. The middle list is dense but every item earns its place by naming concrete returned artifacts.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by enumerating the returned artifacts and their ordering, so an agent knows what comes back. What remains missing is any read-only/safety assurance, keeping it below a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the single optional initiative parameter with its cross-initiative/un-tagged semantics. The description adds no parameter guidance beyond that, making baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Restore") and resource ("session context") and then enumerates exactly what is assembled: pinned set, 24h episodes, open reviews, open tasks, pending claims, reasoning trails. This enumeration sharply distinguishes it from single-purpose siblings like recent, recall, or board.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Run this when re-entering a project" is a clear, concrete trigger. It gives no exclusions or named alternatives (e.g., when to prefer recall or overview instead), so it stops short of the 5 level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
betweenA
Show every edge between two nodes, both directions, at NOW — answers "are A and B connected, and how?". Edges are typed (refers_to, causal, derived_from, contradicts, part_of, blocks, targets, supersedes, verifies, falsifies, temporal), so the answer says what KIND of connection it is. path finds a route when there is no direct edge; drill lists neighbours rather than edges.
| Name | Required | Description | Default |
|---|---|---|---|
| a | Yes | First node name (also accepts a UUIDv7 id). | |
| b | Yes | Second node name (also accepts a UUIDv7 id). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full behavioral burden. It does disclose real traits: bidirectional traversal, evaluation 'at NOW' (a point-in-time snapshot), and that edges are typed with an enumerated ontology. However, it says nothing about what happens when no edge exists, whether results are paginated or bounded, permission requirements, or error behavior for unknown node names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior before the alternatives, and only two sentences. The 11-value edge-type enumeration is long but substantive rather than filler, since it defines the kind of answer returned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain the return shape, and it partially does: typed edges, both directions, snapshot at NOW. It leaves the empty-result case and result bounds unaddressed, but the core contract for calling the tool correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions no parameters at all. Schema coverage is 67%: 'a' and 'b' are documented in the schema, but 'initiative' has no description anywhere and defaults to null, so an agent cannot learn its meaning from either source. The description does not compensate for this documented gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Show every edge between two nodes, both directions') and frames the question it answers ('are A and B connected, and how?'). It explicitly distinguishes itself from siblings 'path' and 'drill' by contrast, so an agent can identify it without reading any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'path finds a route when there is no direct edge; drill lists neighbours rather than edges.' Both the when-to-use and the when-to-use-something-else conditions are stated against named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
boardA
Show the initiative's task board: status columns (from its registry, in order, empties included) with the tasks bucketed into them. Requires initiative. Optional when rewinds the whole board — columns and cards — to a past moment (unix seconds, RFC-3339, or 5m / 2h ago).
| Name | Required | Description | Default |
|---|---|---|---|
| when | No | Optional moment to rewind the board to (unix seconds, RFC-3339, or `5m` / `2h` ago). Omit for the board as it stands now. | |
| initiative | No | Initiative whose board to show (required). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it delivers real behavioral detail: columns are pulled from the initiative's registry in order with empties included, tasks are bucketed into them, and `when` rewinds the entire board including columns. It stops short of explicitly confirming the operation is a non-mutating read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler: purpose and return shape first, then the required parameter and the optional rewind semantics. Every clause carries information the schema or annotations do not.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, two-parameter tool with no output schema, the description explains both inputs and describes the conceptual return structure (ordered columns with bucketed tasks). It lacks only a note on ordering/pagination edge cases, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented, which sets the baseline at 3. The description adds only a marginal nuance over the schema by clarifying that `when` rewinds the whole board rather than just the cards, but the accepted time formats are restated from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and resource ('the initiative's task board') and even outlines the returned structure (ordered status columns with tasks bucketed into them). It does not distinguish itself from the sibling 'board_status', which an agent could easily confuse with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Requires `initiative`' gives a prerequisite, and the optional `when` use case is explained, so usage is implied clearly. However, it never says when to pick this over 'board_status', 'task', or 'overview', leaving sibling selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
board_statusA
Customize the initiative's board columns: action add (key,label?) / remove (key) / relabel (key,label) / reorder (order = all keys permuted). The board is created from defaults [open, in-progress, done] on first edit.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Status key (stable id). Required for add / remove / relabel. | |
| label | No | Human label. Required for relabel; optional for add (defaults to key). | |
| order | No | Full ordered list of existing keys. Required for reorder. | |
| action | Yes | What to do: `add` / `remove` / `relabel` / `reorder`. | |
| initiative | No | Initiative whose board to customize (required). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses one non-obvious behavior — the board is lazily created from defaults [open, in-progress, done] on first edit — which is genuinely valuable. However, for a tool with destructive operations (`remove`, and `reorder` which rewrites the whole column set) it says nothing about what happens to existing statuses or tasks in a removed column, whether changes are reversible, or what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose, then a compact action-to-parameter digest. Every clause carries information; nothing is padding or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations and no output schema, the description covers the surface API well but omits the consequences of editing: effect on tasks occupying a removed column, behavior on invalid or partial `order` permutations, and any return/confirmation semantics. Adequate but with real gaps an agent could trip over.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter, including the per-action requirements for key/label/order. The description's `action add (key,label?) / remove (key) ...` recaps the schema rather than adding new semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource: customizing an initiative's board columns, and enumerates the four supported actions. It is clear on its own, but it never distinguishes itself from the sibling tool `board` (presumably the read side of the same resource), so an agent must guess at the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the context in which the tool applies (customizing board columns) and maps each action to its required arguments, which effectively tells the agent how to invoke each mode. There are no exclusions or named alternatives (e.g., when to use `board` instead of `board_status`), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chainA
Save the shortest weighted path between two nodes as a knowledge chain — an ordered, recallable reasoning trail. Stronger links (higher link weight) make shorter paths. Pass summary to note why the trail matters (it labels the chain for later triage). Idempotent — an identical chain is reused, not duplicated. Reports if the two are unconnected.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | End node name or UUIDv7 id. | |
| from | Yes | Start node name or UUIDv7 id. | |
| name | No | Optional name for the saved chain (auto-derived from endpoints if omitted). | |
| summary | No | Optional one-line summary of why this trail matters — having traced the path, say what it captures. Becomes the chain's body so `chains` can be triaged by name + summary without reading every trail. Auto-derived if omitted. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it discloses idempotency (identical chain reused, not duplicated) and the unconnected-endpoints outcome, which are meaningful behavioral traits beyond the schema. It omits permission/rate-limit context, but for a save operation the key behaviors are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and scope before the secondary notes about summary and idempotency. No filler, though slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param mutation tool with no annotations and no output schema, the description covers the essential behaviors (idempotency, unconnected reporting, summary purpose). The gap is the undescribed `initiative` parameter and no mention of what the returned chain looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the baseline is 3, but the description adds real meaning: it ties link weight to path length (explaining from/to semantics) and explains that summary labels the chain for later triage. It does not address the undocumented `initiative` parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (save) and resource (shortest weighted path as a knowledge chain / reasoning trail), which is clear and distinct from path-computation siblings. It stops short of naming which sibling to use instead (e.g., path, rechain, between), so it lacks explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: it explains that summary labels the chain for triage and that unconnected nodes are reported, but gives no explicit when-to-use vs alternatives guidance against siblings like path, rechain, or between.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citeA
Record an archival reference. Two flavours: external source (pass url for papers / gists / dashboards) OR persona / entity (skip url for people, places, books without links). Both land in archival tier — long-term recall. Pass visibility=shared (in a team initiative) to capture and push to the cloud in one call. Pass link_to (with weight) to connect it to an existing node in this same call — an island is found only by exact name.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Optional URL of the source. Skip for persona / entity records (a person, place, book without a link). | |
| body | Yes | One-paragraph summary — what's at the link, or who this entity is. | |
| name | Yes | Short, recallable name. | |
| after | No | Optional: hold this back until `after` (YYYY-MM-DD), then surface it in `awake` as a debt. For something TRUE LATER, NOT NOW — a certificate that expires, a number to re-measure before quoting it again, work to pick up after the next release. Requires `for_days`. | |
| cloud | No | Which cloud `visibility=shared` publishes to. Only consulted when sharing. Required when several clouds are configured — the push is refused rather than sent to a default you did not name. | |
| layer | No | Optional memory layer stamped at creation: `core`, `hot`, `warm`, `cold`, or `frozen`. Defaults to `warm`. | |
| weight | No | How load-bearing the `link_to` edge is, 0..1. REQUIRED with `link_to`: it is the only signal knowledge chains route on, and there is no default because an unweighted graph makes every chain rank on noise. The capture still lands without it; the edge does not. | |
| link_to | No | Optional: the node this one connects to, by name or id — the edge is made in THIS call, while both ends are still in mind. Linking as a second step is the step nobody takes: one real vault reached 23 nodes and 0 edges with the nudge asking every time. Needs `weight`. | |
| for_days | No | Optional: how many days the reminder keeps appearing once it surfaces. REQUIRED with `after`, and there is no default — state it deliberately, because "the certificate expires" and "re-measure this" want completely different windows. The window starts when the reminder is first actually seen, so it cannot expire while nobody is looking. | |
| edge_type | No | Edge type for `link_to` — same closed vocabulary as `link` (`refers_to` by default, `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`). | |
| initiative | No | ||
| visibility | No | Optional visibility. `shared` marks team knowledge and — in a `team` initiative with the secret guard clear — pushes it to the cloud in this one call. Defaults to `local` (stays private). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the archival tier, that shared visibility pushes to cloud in the same call, that link_to creates the edge inline, that islands are matched only by exact name, and that cloud is refused rather than defaulted. These are real behavioral traits beyond the schema text.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the primary two-flavour branch are front-loaded, followed by the optional capabilities. Each sentence carries information, though the density makes it a slightly heavy read for a single description block.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, annotation-free tool with no output schema, the description covers the tricky interactions (weight required with link_to, for_days required with after, cloud conditional on visibility) that a schema alone would not make obvious. Coverage is strong, with only the sibling-routing dimension left thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 92%, so the schema already documents nearly every parameter. The description adds framing for the url flavour and reinforces weight/link_to coupling, but it is largely redundant with the structured parameter docs. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Record an archival reference') and cleanly splits the tool into two flavours (external source with url vs persona/entity without). The intended target is clear, though it does not explicitly distinguish itself from nearby siblings like jot, link, or attach.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional guidance for the main branches (pass url vs skip; visibility=shared to push to cloud; after for 'true later, not now'), which is useful. However, it never frames when to choose cite over competing tools (jot, link, attach), leaving the sibling-selection decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claimA
Record a hypothesis. Auto-named. If you ALREADY know how it turned out — the usual case, since you reach memory after the check has run — pass verdict (supported/refuted/inconclusive) and by (the evidence node) and it lands settled in this one call. Without a verdict it is an open question, and awake will keep surfacing it until one arrives. Optional about links via refers_to.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | Evidence node (name or id) the verdict rests on — linked `verifies` / `falsifies`. Optional, but a verdict without one is a claim with no citation. | |
| text | Yes | The hypothesis text. Auto-named from first words + id suffix. | |
| about | No | Optional existing node this claim is about (refers_to edge). | |
| layer | No | Optional memory layer stamped at creation: `core`, `hot`, `warm`, `cold`, or `frozen`. Defaults to `warm`. | |
| verdict | No | The verdict, when you ALREADY KNOW IT — the usual case, since you normally reach memory after the check has run. Omit for a genuinely open question. A CLOSED vocabulary, one of exactly these: `supported`, `refuted`, `inconclusive` (`confirmed` / `falsified` / `partial` are accepted as aliases). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does well: it discloses auto-naming, the settled-vs-open lifecycle, and that open claims are repeatedly resurfaced by `awake` until a verdict arrives. It does not state return value or whether late verdict attachment is possible, but the core behavioral model is conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the auto-naming rule are front-loaded, followed by the verdict branch and the open-question consequence. Three sentences with little waste; minor redundancy between the verdict sentence and the schema's own verdict description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema and no annotations, the description covers the primary workflow paths (settled claim, open claim, optional linking) well enough to call correctly. It stops short of describing what the created node looks like on return or the effect of `initiative`/`layer`, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema already documents most parameters including the accepted verdict vocabulary and the `by` citation semantics. The description still adds meaning by framing `verdict`/`by` as the one-call settled path and `about` as a relationship link, tying the parameters to an outcome rather than just restating types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ("Record a hypothesis") and immediately adds the auto-naming behavior, so an agent knows exactly what this tool creates. It does not explicitly contrast itself with nearby siblings like `evidence`, `cite`, or `refute`, which would push this to a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear conditional: pass `verdict` + `by` when you already know the outcome ("the usual case"), omit them for a genuinely open question. It also states the downstream consequence ("`awake` will keep surfacing it"). What's missing is routing against alternative tools such as `settle`, `refute`, or `inconclusive`, which likely overlap with the verdict path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_reviewA
Close an open review on a node — retracts its contradicts edge(s) so it leaves the review queue; the doubt stays in history. Optional resolution note is recorded as provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | Target node name (or id) whose open review to close. | |
| initiative | No | ||
| resolution | No | Optional note on how it was settled — recorded as a resolution episode that supersedes the closed review. Omit for a bare close. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the mutation (retracting contradicts edges), the queue state change, and that the doubt persists in history, plus how resolution is stored as provenance. It omits auth/permission needs and whether a close can be undone (e.g. via un/settle), leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with the core action first, a clarifying em-dash clause on the side effect, and the optional-parameter note last. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description explains the core behavior well, but it never mentions the `initiative` scoping parameter, leaving an agent without guidance on a third of the inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the description's note on `resolution` being recorded as provenance largely repeats the schema's own description. The `initiative` parameter is undocumented in both the schema and the description, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Close an open review on a node') with the actual mechanism spelled out: it retracts the contradicts edge(s) so the node leaves the review queue. This clearly distinguishes it from siblings like sync_review, resolve, refute, and settle.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the triggering condition ('open review') and the effect ('leaves the review queue'), giving clear context for when to call it. It does not explicitly name alternatives (e.g. sync_review, resolve, settle) or say when not to use it, so routing is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cloud_initiativesA
List the initiatives a cloud holds, with how many nodes each has shared. The map of the second tier: use it when you don't know what the team has, before cloud_recall on one of them. Note an initiative in the cloud is INDEPENDENT of the local one with the same name — same name, different contents. (clouds lists the clouds this daemon can reach; this lists what is inside one.)
| Name | Required | Description | Default |
|---|---|---|---|
| cloud | No | Which cloud to ask. Required when several are configured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a critical behavioral trait: a cloud initiative is INDEPENDENT of the local one with the same name (same name, different contents). It does not cover auth/permission needs or the exact return shape, but for a read-oriented listing this is solid disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first sentence, then usage, then the independence caveat. Slightly dense with backticked references and a nested parenthetical, but each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, usage routing, and the key independence semantic for a no-annotation, no-output-schema tool. Return format is only implied ('how many nodes each has shared'), but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single `cloud` parameter already has 100% schema description coverage ('Which cloud to ask. Required when several are configured.'), so the schema does the heavy lifting. The description reinforces the cloud scope but adds no format or syntax detail beyond the schema, matching the baseline 3 for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the initiatives a cloud holds') and adds the returned detail ('how many nodes each has shared'). It explicitly differentiates itself from the sibling `initiatives` (local) and from `clouds` (which lists reachable clouds), so an agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('use it when you don't know what the team has, before `cloud_recall` on one of them') and names the alternative it precedes. The closing parenthetical also routes the agent between `clouds` and this tool, covering both alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cloud_linksA
Resolve a node's cloud soft links — fetches the cloud nodes they point at, live, so you see what they say NOW rather than what they said when the link was made. The read half of link_cloud. Routes each link to the cloud it was created against.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Local node name or UUIDv7 id whose cloud soft links to resolve. | |
| initiative | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a genuinely useful trait: links are resolved live, so results reflect current cloud node content rather than link-creation-time snapshots. It also notes per-link routing to the originating cloud. However, it says nothing about permissions, dangling/missing links, or failure behavior for what is otherwise a read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and scope, then the freshness behavior, then the sibling relationship. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-param read tool with no output schema and no annotations, the description covers the core behavior (live resolution, per-cloud routing) but leaves the `initiative` parameter unexplained and omits error/edge behavior for dangling links. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: `name` is documented in the schema, but `initiative` has no description anywhere. The tool description talks about 'a node's cloud soft links' but never clarifies what `initiative` scopes or requires, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Resolve a node's cloud soft links' with the scope clarified as fetching the cloud nodes they point at. It explicitly positions itself as 'The read half of `link_cloud`', which lets an agent distinguish it from the write-side sibling without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'read half of `link_cloud`' framing tells the agent when to pick this over the write tool, and the live-fetch emphasis implies use when current cloud state matters. No explicit when-not or mention of other read alternatives like `cloud_recall`/`clouds`, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cloud_recallA
SEARCH or list what the cloud holds for an initiative — the second tier your local search cannot reach. Pass query to match shared names and excerpts; omit it to list everything. Then pull <id> brings one into the local vault. Reach for this whenever a team initiative's local answer looks complete: the cloud may hold nodes this machine has never seen. Paged at 25, reporting the true total and the exact call for the next page. In a multi-cloud setup cloud is required.
| Name | Required | Description | Default |
|---|---|---|---|
| cloud | No | Cloud name to query in a multi-cloud setup. Required when several are configured. | |
| limit | No | Page size, default 25, ceiling 500. A shared initiative can hold hundreds of nodes — more than a context window — so this read is bounded like every other list on the surface. | |
| query | No | Search the cloud instead of listing it: a case-insensitive substring over shared node names and excerpts. Omit to list everything shared. | |
| offset | No | How many to skip, for the next page. The result says when there is one. | |
| initiative | Yes | Initiative to list shared cloud nodes for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden: it discloses pagination behavior (paged at 25, reports the true total and the exact next-page call), the bounded-read rationale, and the multi-cloud `cloud` requirement. It does not state auth/permission needs or confirm read-only semantics explicitly, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded with the core verb, and each sentence adds routing or behavioral value (search vs list, pull, pagination, multi-cloud). Slightly wordy with the extended pagination and vault commentary, but nothing is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe returns — and it does, covering pagination size, true total, and the next-page call. Read-only/permission behavior is left implicit, a minor gap for a bounded read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, setting the baseline at 3. The description adds little beyond it — it restates the query-omit-to-list behavior and the cloud-required rule, which the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search or list) plus resource (shared cloud nodes for an initiative), and explicitly distinguishes itself from the sibling `search` as 'the second tier your local search cannot reach.' An agent can tell it apart from local recall/search without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('reach for this whenever a team initiative's local answer looks complete'), the query-vs-list condition ('pass query... omit it to list everything'), the follow-up alternative (`pull <id>`), and the multi-cloud prerequisite for `cloud`. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cloudsA
List the clouds this daemon can reach, with their endpoints and which one is default. Ask this before any cloud verb in an unfamiliar setup: with more than one cloud configured, share / pull / cloud_recall and the initiative verbs require cloud named explicitly, and nothing is routed to a default you did not choose.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the return shape (endpoints + which cloud is default) and the important routing semantic that nothing is routed to a default the agent did not choose. It stops short of stating failure modes or what happens with zero clouds configured, so it falls just short of complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and then the usage rule, with zero filler. Every clause adds information an agent needs before deciding to call it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly compensates by describing the return content (clouds, endpoints, default). Combined with the routing warning it gives an agent everything needed to call and interpret this zero-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document and the baseline is 4. The empty schema is consistent with a pure enumeration call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the clouds this daemon can reach') and even names the payload fields the agent will get back (endpoints, default marker). It is immediately distinguishable from the many sibling verbs like share, pull, and cloud_recall, which it explicitly contrasts itself with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition: 'Ask this before any cloud verb in an unfamiliar setup.' It then names the exact alternative tools that require a named cloud (share / pull / cloud_recall and the initiative verbs) and the condition under which they require it (more than one cloud configured).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
configB
Show resolved configuration: vault path, the configured clouds and which is default, and every cap (initiative not relevant).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but adds little beyond implying a read-only display. It does disclose the returned fields (vault path, clouds, default cloud, caps), which is useful, but says nothing about freshness, permissions, or whether values are cached vs computed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The trailing parenthetical '(initiative not relevant)' is slightly cryptic but does carry a scoping signal, so it roughly earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument read tool with no output schema, the description usefully enumerates what is returned (vault path, clouds, default, caps). Remaining gap is the lack of routing relative to sibling configuration/list tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is no parameter surface to document, and the description correctly adds nothing about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and resource ('resolved configuration'), then enumerates the contents: vault path, configured clouds and default, and caps. Clear on its own, but it does not distinguish itself from siblings like 'clouds', 'path', or 'policy', which appear to cover overlapping surfaces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and no named alternatives. The only routing hint is the parenthetical '(initiative not relevant)', which scopes out one case but leaves the relationship to 'clouds'/'path' unexplained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirmA
Mark a hypothesis as supported. by (the verifying evidence node) is optional — record the verdict even with nothing to point at yet rather than leaving the claim open with the answer buried in its text.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | Evidence node name or id — linked `verifies` for `confirm`, `falsifies` for `refute`. OPTIONAL: record the verdict even with nothing to point at yet, rather than leaving the claim tagged `open` with the answer buried in its prose. (`inconclusive` writes no edge at all, so it never needs one.) | |
| hypothesis | Yes | Hypothesis name or id. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose a meaningful state transition — that skipping the verdict leaves the claim 'open' with the answer buried in its text — but says nothing about permissions, whether a prior verdict is overwritten, or how to reverse a confirmation (e.g. unsettle/supersede).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the optional-parameter caveat immediately after. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no annotations and no output schema, the description covers the verdict concept and the optional evidence link but ignores `initiative` entirely and omits any reversibility or interaction detail with sibling verdict tools. Adequate but with clear holes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description's explanation of `by` (the verifying evidence node, optional) largely restates the schema's own `by` description, so it adds little beyond it. The third parameter, `initiative`, is undocumented in both the schema and the description, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Mark a hypothesis as supported.' An agent can infer this is the affirmative-verdict tool, but the description never names its obvious siblings (refute, inconclusive), so the agent must rely on the tool list to disambiguate the trio.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a usable decision rule for the optional `by` parameter ('record the verdict even with nothing to point at yet'), but offers no explicit guidance on when to choose confirm over refute or inconclusive, nor any prerequisites. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_initiativeA
Delete an initiative — drops its scoping and forgets nodes exclusive to it (bi-temporal: recoverable via at at a past time). Nodes shared with other initiatives only lose this membership. Local by default. Pass cloud="<name>" to ALSO delete it from that shared cloud, which removes it for everyone and CANNOT be undone there; with several clouds configured the name is required rather than defaulted.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Initiative to delete. | |
| cloud | No | Name of a cloud to ALSO delete this initiative from — removes it for everyone, and the cloud has no undo. Omit for a local-only delete. Never defaulted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it explains that local deletion drops scoping and forgets exclusive nodes, that shared nodes only lose membership, that local deletion is recoverable via `at`, and that cloud deletion is irreversible for everyone. This is exactly the kind of destructive-effect disclosure an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and packed with distinct, non-redundant information. Every sentence adds a new behavioral constraint (scope, shared nodes, local default, cloud irreversibility), with no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description is complete: it covers default behavior, local versus cloud effects, reversibility, shared-node handling, and when the cloud name is required. Nothing an agent needs to call it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds extra meaning beyond the schema by stating that with several clouds configured the name is required rather than defaulted, and it reinforces the irreversible cloud-wide effect, which helps an agent decide whether to pass `cloud` at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) and resource (initiative), and immediately describes the scoping and node-level effects. It distinguishes local from cloud deletion and names the bi-temporal recovery mechanism, so an agent can tell exactly what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage context for the `cloud` parameter ('Local by default. Pass cloud=... to ALSO delete it...'), including when the name is required with multiple clouds. However, it does not explicitly say when to choose this tool over alternatives like merge_initiative, forget, or rename_initiative, so guidance is implied rather than comparative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
doneA
Mark a task done — the id, name and manual tags survive; only the status moves, and history shows the transition. Accepts a task name or a UUIDv7 id. For any other column of the initiative's board use set_status; done is the shortcut for the terminal one.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose real mutation semantics: id, name and manual tags survive, only status moves, and `history` records the transition. It stops short of stating whether the change is reversible or whether any permission applies, which for a no-annotation mutation tool is a remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the observable effect of the call and then the sibling routing and accepted input forms. No filler, and the most decision-relevant information comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description usefully explains the mutation's effects, the identifier formats, and where the change is recorded. It leaves the `initiative` parameter and any return behavior unexplained, but that is a minor omission given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%. The description reinforces the `name` parameter by stating it accepts a task name or UUIDv7 id, which adds meaning. However, the `initiative` parameter (with default null) is undocumented in both the schema and the description, so the description only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Mark a task done' names both the action and the object, and the follow-up sentence distinguishes it from `set_status` ('the shortcut for the terminal one'). An agent can identify what it does and how it differs from the sibling that does essentially the same job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'For any other column of the initiative's board use `set_status`' tells the agent exactly when to pick the alternative, and implicitly when not to use this one. Also clarifies the acceptable identifier forms (task name or UUIDv7 id) so the agent doesn't need to guess about resolution.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drillA
Drill into a node — its brief plus one hop of children (sources via derived_from, parts via part_of). The fast way to see what a memory is attached to. Bodies come back as EXCERPTS: use at <name> when you need the whole text, and between a b when you want the edges rather than the neighbours.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the important behavioral trait that bodies come back as EXCERPTS rather than full text, plus the exact edge types included. It stops short of permissions or any error behavior, but the key return-shape caveat is surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose first, then the excerpt caveat, then the sibling routing. No filler and well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema and no annotations, the description adequately conveys what is returned and its excerpted form. The main gap is the unexplained `initiative` parameter, which an agent may need before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `name` is documented in the schema and reinforced by the description's node/name context, but the `initiative` parameter is undocumented in both the schema and the description. The description adds no syntax or meaning beyond what the schema already gives.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (drill) and resource (a node) and precisely what comes back: its brief plus one hop of children, with the edge semantics named (sources via derived_from, parts via part_of). It is immediately distinguishable from siblings like at and between.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use `at <name>` for the whole text and `between a b` for the edges rather than the neighbours. Both alternatives and the conditions that select them are stated, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
episodeA
Write a deliberately-named operational episode. Use when you know you'll want to recall by exact name. Pass visibility=shared (in a team initiative) to capture and push to the cloud in one call. Pass link_to (with weight) to connect it to an existing node in this same call — an island is found only by exact name.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Free-form body. | |
| name | Yes | Short, recallable name. | |
| after | No | Optional: hold this back until `after` (YYYY-MM-DD), then surface it in `awake` as a debt. For something TRUE LATER, NOT NOW — a certificate that expires, a number to re-measure before quoting it again, work to pick up after the next release. Requires `for_days`. | |
| cloud | No | Which cloud `visibility=shared` publishes to. Only consulted when sharing. Required when several clouds are configured — the push is refused rather than sent to a default you did not name. | |
| layer | No | Optional memory layer stamped at creation: `core`, `hot`, `warm`, `cold`, or `frozen`. Defaults to `warm`. | |
| weight | No | How load-bearing the `link_to` edge is, 0..1. REQUIRED with `link_to`: it is the only signal knowledge chains route on, and there is no default because an unweighted graph makes every chain rank on noise. The capture still lands without it; the edge does not. | |
| link_to | No | Optional: the node this one connects to, by name or id — the edge is made in THIS call, while both ends are still in mind. Linking as a second step is the step nobody takes: one real vault reached 23 nodes and 0 edges with the nudge asking every time. Needs `weight`. | |
| for_days | No | Optional: how many days the reminder keeps appearing once it surfaces. REQUIRED with `after`, and there is no default — state it deliberately, because "the certificate expires" and "re-measure this" want completely different windows. The window starts when the reminder is first actually seen, so it cannot expire while nobody is looking. | |
| edge_type | No | Edge type for `link_to` — same closed vocabulary as `link` (`refers_to` by default, `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`). | |
| initiative | No | ||
| visibility | No | Optional visibility. `shared` marks team knowledge and — in a `team` initiative with the secret guard clear — pushes it to the cloud in this one call. Defaults to `local` (stays private). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful side effects: visibility=shared pushes to the cloud in the same call, link_to creates an edge immediately, and orphaned captures are only retrievable by exact name. It does not cover permissions, reversibility, or failure modes, so it is good but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with purpose and usage, and each carries an actionable instruction. The exact-name point is stated twice (once for recall, once for linking), which is mild redundancy keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter write tool with 91% schema coverage and no output schema, the description adequately frames the core workflow (named capture, optional cloud push, optional inline edge). Remaining gaps around non-core parameters are covered by the schema, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 91%, so the baseline is 3; the description adds real meaning on top by clarifying the link_to/weight coupling and by implying that `name` is the exact-match recall key. It does not explain most parameters (after, layer, edge_type, cloud), leaving the schema to do that work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Write a deliberately-named operational episode") and adds a differentiating qualifier: it is the capture path for content you will later find by exact name. This is clearly distinct from generic capture siblings, though no sibling tool is named explicitly for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use trigger ("Use when you know you'll want to recall by exact name") and two conditional workflows (visibility=shared for team/cloud, link_to for edge creation in the same call). No explicit exclusions or named alternatives, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evidenceB
Record what you actually checked, and attach it to a hypothesis. Past tense — this documents a check that already ran, it does not schedule one (and it is not cargo test). Pass method to write the result up as a new experiment node, or node to point at something you already captured.
| Name | Required | Description | Default |
|---|---|---|---|
| node | No | An existing node (name or id) to register as the evidence instead of writing a new one — an episode you already captured, say. Give this OR `method`. | |
| method | No | What you actually did and what came out of it — past tense. Creates the experiment node. Give this OR `node`. | |
| hypothesis | Yes | Hypothesis name or id this evidence bears on. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the mutation semantics: passing `method` creates a new experiment node, passing `node` registers something already captured. Beyond that it says nothing about permissions, reversibility, failure if the hypothesis is absent, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, and the parameter routing follows logically. The parenthetical `cargo test` aside is slightly gratuitous but the whole thing stays to three tight sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no annotations and no output schema, the description covers purpose and the method/node choice but leaves gaps: the unexplained `initiative` param, the return/error behavior, and any prerequisites around the hypothesis are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the schema already documents most parameters, including the 'Give this OR `node`' mutual exclusivity. The description adds the past-tense framing for `method`, but largely restates schema content and does not explain the undocumented `initiative` parameter. Baseline 3 fits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('record... evidence') and its scope ('attach it to a hypothesis'), which sets it apart from generic linking tools like attach/link. It further clarifies what it is not (scheduling, `cargo test`). It doesn't explicitly contrast against close siblings such as cite/claim/episode, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real usage context — past tense, documents a check that already ran, does not schedule — which implies when to reach for it. However it never names an alternative sibling (cite, claim, episode, attach) or states when-not to use this versus those, so the when/when-not guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exportB
Snapshot the substrate as an Obsidian-friendly markdown vault (README + INDEX + LOG + pages). Output dir is created if missing.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | ||
| output_dir | Yes | Output directory. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses a side effect ('Output dir is created if missing') and the shape of the output (vault with README/INDEX/LOG/pages), but omits critical traits such as whether existing files are overwritten, permissions required, or whether the operation is read-only on the substrate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and its output, with the side-effect note appended where it belongs. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple export tool with no output schema and no annotations, the description covers the core action and one side effect, but leaves the 'initiative' parameter's meaning and overwrite/idempotency behavior unexplained. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% and the one documented parameter (output_dir) is described minimally as 'Output directory.' The description does add meaning about output_dir behavior (created if missing), but the 'initiative' parameter is undocumented in both schema and description, leaving half the surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (snapshot/export) and a specific resource (the substrate as an Obsidian-friendly markdown vault), even naming the artifacts produced (README + INDEX + LOG + pages). The action is clear, but it never references the obvious sibling 'import' or any other tool to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what is produced but gives no when-to-use guidance, no prerequisites, and no mention of alternatives such as 'import' or other export-adjacent tools. An agent must infer the usage context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flagA
Flag a node as doubtful — writes a review episode carrying your REASON and a contradicts edge to the target. The target itself is untouched: the doubt is recorded beside it, not written into it. It then shows up in awake's under-review list until close_review or resolve settles it. Not the same as link contradicts, which records the edge without the reason.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Reason / description of the concern. | |
| target | Yes | Target node name to flag. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does well: it discloses the side effects (review episode + edge), that the target is untouched and the doubt recorded beside it, and the lifecycle visibility until settled. It omits edge cases such as flagging a nonexistent target or repeated flags, so not quite exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and side effect, then the non-destructive guarantee, then lifecycle, then the sibling distinction. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations, so the description must cover behavior and it largely does: side effects, target immutability, and where the flag surfaces. The only real hole is the undocumented `initiative` parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: `reason` and `target` are described in the schema, and the description reinforces that the reason is carried into the episode. The third parameter `initiative` (nullable, default null) has no description in either the schema or the description, so the gap is not compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (flag) and the exact effect: writes a review episode carrying the reason plus a contradicts edge to the target. It names and distinguishes the closest sibling ('not the same as `link contradicts`'), so an agent can choose correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when the flag shows up (in `awake`'s under-review list) and how it is cleared (`close_review` or `resolve`), plus the explicit contrast with `link contradicts`. It stops short of stating when not to flag at all, but the alternative-routing is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forgetA
Bi-temporal forget — retract a node and every edge connected to it. Historical reads still see it; reads at NOW skip.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the cascade (node plus every connected edge), the bi-temporal semantics (historical reads still see it, reads at NOW skip), implying a soft/non-physical retraction. It omits auth/permission requirements, reversibility, and effect on the edges' other endpoints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and then the temporal behavior. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with no annotations and no output schema, the description covers the most decision-relevant facts (cascade scope, temporal visibility). It could say more about reversibility and permissions, but nothing essential for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% — 'name' is documented but 'initiative' has no description in either schema or text. The description adds no parameter-level detail (e.g., UUIDv7 acceptance or initiative scoping), so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('retract') and resource ('a node and every edge connected to it'), plus a distinctive bi-temporal qualifier. It's clear what the tool does, but it doesn't explicitly differentiate itself from near-siblings like delete_initiative or supersede.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no routing to alternatives. The description never says when to choose 'forget' over 'supersede', 'delete_initiative', or 'unlink', leaving selection to inference from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
historyA
Print every assertion / retraction recorded for a node, chronologically — + asserted, - retracted. This is how you see that a node CHANGED, and when. Accepts a former name too: a node renamed by revise or supersede is still reachable by the name it used to carry. Pair with at <name> when=<t> to read any of those versions in full.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose meaningful traits: chronological output, `+`/`-` semantics, and that renamed nodes remain reachable by former names (via `revise`/`supersede`). It doesn't state permissions or output size/pagination behavior, but the read-only nature is strongly implied by 'Print'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences, each doing work: purpose, the change-tracking rationale, and the former-name behavior plus the pairing hint. Slightly dense but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no annotations and no output schema, the description covers output shape (chronological, `+`/`-`) and a non-obvious input behavior (former names). It leaves the `initiative` parameter and any output limits unexplained, but the essentials to call it correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `name` is documented and the description usefully extends it (former names accepted, beyond the schema's name/UUID note). However, the `initiative` parameter is undocumented in both schema and description, leaving a real gap that the description does not compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Print) and resource (assertions/retractions for a node, chronologically), and even decodes the output symbols. It distinguishes itself from the sibling `at` by framing itself as the change-over-time view rather than a version reader.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tells the agent when this is the right lens ('how you see that a node CHANGED, and when') and explicitly routes to the companion tool `at <name> when=<t>` for reading a specific version in full. No explicit when-not, but the boundary with `at` is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hygieneA
Hygiene status for an initiative: node and core counts, when the last pass ran, whether one is due, and exactly what the next pass would move. Passes run on their own — when writes accumulate, when core grows past its threshold, or on the sweep timer — and only ever change a node's layer, reversibly. force=true runs one now.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Run a pass now instead of reporting what one would do. | |
| initiative | Yes | Initiative to report on, or to sweep when `force` is set. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so well: it discloses that passes are self-triggering, names the three trigger conditions, scopes the mutation to 'only ever change a node's layer', and explicitly states the change is reversible. That is precisely the safety-relevant context an agent needs before calling with force=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the report contents before the behavioral mechanics and the force override. Every clause earns its place — counts, timing, due state, projection, triggers, mutation scope, reversibility — with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by enumerating the report's contents, and it fully explains the one non-obvious parameter. It stops short of describing the report's structure or any permission/auth requirements, but for a two-parameter read-mostly tool this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema, and the description's '`force=true` runs one now' largely restates the schema's own wording. It adds no format, default, or interaction detail beyond what is structured, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Hygiene status for an initiative') and enumerates exactly what the report contains: node and core counts, last pass time, due status, and projected next-pass movement. It is unambiguous on its own, though it never explicitly contrasts itself with nearby siblings such as `lint` or `policy`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains that passes run automatically (on write accumulation, core threshold, or sweep timer) and that `force=true` triggers one now, which gives the agent context for the force flag. However, it never states when to reach for this tool versus `lint`, `overview`, or `history`, so the base read use case is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ideasA
List the initiative's archival IDEAS — proposals that settled without yet becoming results. Part of the cortex awake loads every session, so this is the deliberate deep read when you want them all rather than the layered slice. outcomes is the sibling for results; settle is how a node gets here.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | Optional initiative to scope the operation to. When omitted, reads are cross-initiative; mutations end up un-tagged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description does meaningful work: it labels the operation as a read/list, explains that these are archival settled proposals, and ties them to the `awake` session context. However, it still omits operational details like pagination, return shape, and any permission behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it defines the resource first, then gives routing context, then names sibling alternatives. Every sentence contributes to selection or interpretation without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter list tool with no output schema, the description is nearly complete: it explains what IDEAS are, how they relate to `awake`, and which siblings to consider. The remaining gap is that it does not describe the return payload or pagination behavior, though the schema handles the only parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the sole optional `initiative` parameter is fully documented in the schema. The description does not add syntax, edge-case behavior, or scoping details beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (List) and resource (archival IDEAS), and explicitly defines them as proposals that settled without becoming results. It distinguishes the tool from `outcomes` (results) and `settle` (how a node gets here), so an agent can tell it apart from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use this tool: as a deliberate deep read when you want all IDEAS rather than the layered slice loaded by `awake`. It also names the relevant alternatives (`outcomes` for results, `settle` for how a node gets here). This gives clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
importC
Read FIRST before bulk-importing knowledge. Returns the import playbook: scope by initiative, pick the verb by epistemic status, stamp the memory layer at creation by importance, and link after capturing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The description does not state side effects, permissions, rate limits, or whether calling this tool mutates state. It only promises a 'playbook' with no explanation of what that entails or how to use it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with multiple clauses, fairly front-loaded with the 'Read FIRST' instruction. However, it is dense and jargon-heavy ('epistemic status', 'memory layer', 'stamp') without explaining terms, which reduces navigability for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, no annotations, and no parameter detail. The description provides only a high-level procedural outline but omits what the returned playbook contains, how to act on it, or any prerequisites. For a tool that appears to gate bulk operations, this is insufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete. Per the rubric, 0 params gives a baseline of 4. The description does not and need not describe parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The tool name 'import' and description 'Read FIRST before bulk-importing knowledge' suggest a bulk ingestion tool, but the description never states what the tool actually does. It only says the tool 'Returns the import playbook', which is a meta-instruction rather than a clear verb+resource. An agent cannot tell whether this tool imports data, returns documentation, or does both.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The guidance is one prescriptive step ('Read FIRST before bulk-importing knowledge'), but it does not say when to use this tool vs alternatives, nor does it name any sibling tool. There is no condition for use beyond an imperative ordering.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inconclusiveA
Mark a hypothesis as inconclusive — the check ran and did not decide. A real third verdict, not a failure to answer: it closes the claim out of the open queue while recording that the question stayed open on the merits. Writes no verdict edge, so by is not needed.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | Evidence node name or id — linked `verifies` for `confirm`, `falsifies` for `refute`. OPTIONAL: record the verdict even with nothing to point at yet, rather than leaving the claim tagged `open` with the answer buried in its prose. (`inconclusive` writes no edge at all, so it never needs one.) | |
| hypothesis | Yes | Hypothesis name or id. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key side effects: it 'closes the claim out of the open queue' and 'writes no verdict edge'. It stops short of stating permissions, reversibility, or how related nodes are affected, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then the semantic distinction, then the mechanical consequence. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a verdict-marking tool with no output schema and no annotations, it covers what the tool does and its side effects on the claim queue and edges. It leaves out reversibility/undo and the role of `initiative`, which keeps it just short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description usefully reinforces that 'by' is not needed for this verdict — though the schema's own `by` description already says this. The `initiative` parameter is undocumented in both schema and description, so the description does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Mark a hypothesis as inconclusive') and immediately positions it as 'a real third verdict, not a failure to answer', which distinguishes it from the confirm/refute siblings without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage condition — 'the check ran and did not decide' — and clarifies the distinction from leaving a claim open. It implies but never explicitly names confirm/refute as the alternatives, so the routing is clear but not fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
initiativesB
List initiatives that have at least one node attached. Use this first when re-entering, then pick one for subsequent calls.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the filter (only initiatives with at least one attached node) and hints results feed later calls, but says nothing about read-only nature, return shape, or that initiatives with no nodes are omitted from view.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no waste, and the filter condition is front-loaded ahead of the usage hint. Efficient, though not maximally information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no annotations, and no output schema, the description is the only source of behavioral context. It conveys purpose and a usage cue but omits return format and selection guidance needed to act on the listed initiatives, leaving clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document and the baseline of 4 applies. The description adds no parameter detail, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (initiatives) with a meaningful filter condition (at least one node attached). It is clear, but it does not distinguish itself from the sibling cloud_initiatives, leaving the local-vs-cloud ambiguity unresolved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this first when re-entering' gives a real usage cue and implies this is an entry-point/orientation call. However, it names no alternatives (e.g., cloud_initiatives) and gives no when-not-to-use guidance, so routing is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jotA
Low-friction episode write — auto-named from body's first words plus a unique id suffix. Defaults to observation/low. Pass visibility=shared (in a team initiative) to capture and push to the cloud in one call. Pass link_to (with weight) to connect it to an existing node in this same call — an island is found only by exact name.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Free-form body. Name is auto-derived from first words + id suffix. | |
| after | No | Optional: hold this back until `after` (YYYY-MM-DD), then surface it in `awake` as a debt. For something TRUE LATER, NOT NOW — a certificate that expires, a number to re-measure before quoting it again, work to pick up after the next release. Requires `for_days`. | |
| cloud | No | Which cloud `visibility=shared` publishes to. Only consulted when sharing. Required when several clouds are configured — the push is refused rather than sent to a default you did not name. | |
| layer | No | Optional memory layer stamped at creation: `core`, `hot`, `warm`, `cold`, or `frozen`. Defaults to `warm`. | |
| weight | No | How load-bearing the `link_to` edge is, 0..1. REQUIRED with `link_to`: it is the only signal knowledge chains route on, and there is no default because an unweighted graph makes every chain rank on noise. The capture still lands without it; the edge does not. | |
| link_to | No | Optional: the node this one connects to, by name or id — the edge is made in THIS call, while both ends are still in mind. Linking as a second step is the step nobody takes: one real vault reached 23 nodes and 0 edges with the nudge asking every time. Needs `weight`. | |
| for_days | No | Optional: how many days the reminder keeps appearing once it surfaces. REQUIRED with `after`, and there is no default — state it deliberately, because "the certificate expires" and "re-measure this" want completely different windows. The window starts when the reminder is first actually seen, so it cannot expire while nobody is looking. | |
| edge_type | No | Edge type for `link_to` — same closed vocabulary as `link` (`refers_to` by default, `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`). | |
| initiative | No | ||
| visibility | No | Optional visibility. `shared` marks team knowledge and — in a `team` initiative with the secret guard clear — pushes it to the cloud in this one call. Defaults to `local` (stays private). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses auto-naming, the observation/low default, that shared visibility triggers a cloud push, and that link_to requires weight. It omits auth/permission requirements and is vague on what 'observation/low' actually refers to, but covers the notable side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first clause, then the two optional behaviors. Each sentence carries information and none is filler, though the second sentence packs visibility, team initiative, and cloud push tightly together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter write tool with no annotations and no output schema, the description covers the high-value usage paths (shared/cloud push, in-call linking, weight requirement) while the schema covers the rest. The `initiative` parameter is left undocumented in both places, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 90%, so the schema already documents nearly all parameters, including the weight/link_to and after/for_days dependencies. The description restates the key cross-parameter couplings rather than adding new syntax or semantics, which is the baseline-3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('episode write') with the distinguishing trait of being 'low-friction' and auto-naming from the body. It is clear what the tool does, though it never names or contrasts with the sibling `episode` tool an agent might reasonably confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete conditional guidance: pass visibility=shared in a team initiative to capture-and-push in one call, and pass link_to (with weight) to connect within the same call. It explains the why behind in-call linking ('an island is found only by exact name') but does not explicitly exclude alternatives or route the agent away from separate link/episode calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
layerB
Set a node's memory layer — controls recall priority (injected Core → Hot → Warm → Cold → Frozen). Accepts name or id; layer one of core/hot/warm/cold/frozen.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name or UUIDv7 id. | |
| layer | Yes | Target memory layer: `core`, `hot`, `warm`, `cold`, or `frozen`. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully reveals the layer semantics and injection order (Core → Hot → Warm → Cold → Frozen), but says nothing about permissions, reversibility, side effects, or how an invalid target is handled. The enum ordering is real added value, but the mutation profile remains thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight clauses with the action front-loaded and no wasted framing. Minor redundancy in re-listing the layer values already enumerated in the schema, but the piece is well-sized for its content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations means the description must stand alone, and it does convey the core purpose and layer semantics. It leaves the undocumented 'initiative' parameter and all side-effect/error behavior unexplained, so it is adequate but not complete for a 3-param mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description restates the name/layer meanings already present in the schema plus the priority ordering. The third parameter, 'initiative', is undocumented in both schema and description, so the gap is not compensated. Marginal added meaning over the schema justifies the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Set a node's memory layer') and even clarifies the downstream effect ('controls recall priority'). That is enough to tell what it does and how it differs functionally from siblings like set_status or reweight. It stops short of naming an alternative sibling explicitly, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (change a node's layer to adjust recall priority/injection order), which gives context for when to reach for it. However, it offers no when-not guidance and never names alternatives such as set_status or reweight that also mutate memory state, leaving the routing choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linkA
Create a typed edge between two nodes (by name or id). Endpoints resolve in the active initiative first, then across all initiatives, so a link may span initiatives. Edge type defaults to refers_to. weight (0..1) is REQUIRED — it is HOW LOAD-BEARING the edge is and the only signal knowledge chains route on (path cost is 1−weight); there is no default because an unweighted graph makes every chain rank on noise. State it by the scale: 0.9–1.0 load-bearing (a cause, a source a conclusion rests on, a supersession — the edges a chain should follow); 0.6–0.8 supporting but not decisive; 0.3–0.5 loose / associative. DIRECTION matters for supersedes: link <new> <old> --edge_type supersedes — the replacement points at what it replaced, so an INBOUND supersedes means this node is obsolete.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Destination node name or id — resolved in the active initiative first, then across all initiatives. | |
| from | Yes | Source node name or id — resolved in the active initiative first, then across all initiatives, so an edge may span initiatives. | |
| cloud | No | Which cloud to mirror this edge change to, when BOTH endpoints are already shared. Omit with one cloud configured (it is unambiguous) or when the endpoints are local. With several configured and none named, the local edit still happens and the result says the edge was not mirrored. | |
| weight | Yes | Connection strength in `0..1` — REQUIRED, no default. How load-bearing this edge is, and the only signal knowledge chains route on (path cost is `1 − weight`). State it deliberately by how much the connection matters: 0.9–1.0 load-bearing (a cause, a source a conclusion rests on, a supersession — the edges a chain should follow); 0.6–0.8 supporting, not decisive; 0.3–0.5 loose / associative. There is no neutral fallback on purpose: an unweighted graph is what made every chain rank on noise. | |
| edge_type | No | Edge type — a CLOSED vocabulary, one of exactly these: `refers_to` (default), `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`. `supersedes` is directed: `src` supersedes `dst` — the replacement is the `src`, so an inbound `supersedes` means the node is obsolete. Nothing else is accepted (`related_to` and friends are not edge types). Snake_case or kebab-case both accepted. | refers_to |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does much of it: it discloses endpoint resolution order (active initiative first, then all), the possibility of cross-initiative edges, that weight has no default and why, and the directed semantics of supersedes. It omits permission/auth requirements, failure behavior when a name fails to resolve, and whether an existing edge is overwritten or duplicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then layered with resolution rules, weight guidance, and direction — a sensible order. It is dense but mostly earns its length; the duplication of the weight-scale explanation (already in the schema) is the main wasted sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with no annotations and no output schema, the description covers the decisive behaviors: resolution order, weight semantics, edge vocabulary, direction. Remaining gaps are the undocumented `initiative` param, failure/resolution-error behavior, and what the call returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already 83%, so the baseline is 3; the description goes beyond it with the actionable weight bands and the explicit supersedes direction rule. That said, much of the weight and edge_type guidance is near-verbatim repetition of the schema text, and the `initiative` parameter is left undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Create a typed edge between two nodes') and immediately pins down scope with the name-or-id resolution rule and cross-initiative possibility. An agent can distinguish this from siblings like unlink, reweight, or link_cloud without opening the schema, since the description itself contrasts edge creation against edge mutation and cloud mirroring.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong in-tool guidance — the weight band rubric, the required-vs-default rationale, and the explicit direction convention for supersedes (
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_cloudA
Soft-link a local node to a cloud node by id — a reference, with NO copy in your vault. Use it instead of pull when the cloud node is someone else's to maintain and you only need to point at it: a pull makes a copy that silently goes stale when the owner revises it, while a soft link resolves live through cloud_links. Pull when you need the content locally; link when you need the citation. Edge type defaults to refers_to; in a multi-cloud setup pass cloud to record where the dst lives.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Local node name or UUIDv7 id to soft-link from. | |
| cloud | No | Cloud the dst lives in (multi-cloud). Omit for the default cloud — the soft link records the name so resolution routes to the right endpoint. | |
| cloud_id | Yes | UUIDv7 id of the cloud node to link to. | |
| edge_type | No | Edge type for the soft link — closed vocabulary, one of: `refers_to` (default), `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`. | |
| initiative | Yes | Initiative scope (both sides share the same initiative name). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the burden, and it does well: it discloses that no copy is made, that resolution happens live through `cloud_links`, and that the edge type defaults to refers_to. It does not state idempotency, permission requirements, or what happens on a re-link of the same pair, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and the copy-free guarantee before moving to the sibling comparison. The second sentence is long and clause-dense, but each clause (staleness, live resolution, defaults, multi-cloud) carries distinct information, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must cover behavior, and it explains the reference semantics and live-resolution path. It is nearly complete for a 5-param tool; only idempotency, error behavior, and return shape are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all five parameters including the closed edge_type vocabulary and the multi-cloud `cloud` semantics. The description reinforces the edge-type default and the cloud routing rationale but adds no syntax or format detail beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'soft-link a local node to a cloud node by id' — and immediately clarifies the nature of the operation ('a reference, with NO copy in your vault'). It explicitly distinguishes itself from the sibling `pull`, so an agent can route without reading either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('Use it instead of pull when the cloud node is someone else's to maintain') plus the deciding condition ('Pull when you need the content locally; link when you need the citation'). It even explains the failure mode that motivates the choice — pull copies silently go stale when the owner revises.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lintA
Diagnostic snapshot of graph hygiene: orphan nodes (no edges at all), unresolved reviews, and dangling edges whose endpoint was retracted. Read-only. reflect is the fuller version — it pairs each finding with what to do about it, and adds overdue tasks, stale chains and cortex candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | Optional initiative to scope the operation to. When omitted, reads are cross-initiative; mutations end up un-tagged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral load — and it does declare 'Read-only', which is the essential safety trait. It also discloses the scope of what is scanned. It stops short of return format details or any rate/auth considerations, but the key trait is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the tool's identity and finding categories, then routes to the sibling in one clause. Every sentence earns its place with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only diagnostic with no output schema and no annotations, the description supplies the finding categories the agent needs to interpret results plus the read-only guarantee. Slightly incomplete only in not characterizing output shape beyond those categories.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one optional parameter with 100% schema description coverage, so the schema already documents the initiative scoping and its cross-initiative default. The description adds nothing about this parameter, which is the correct baseline when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('diagnostic snapshot of graph hygiene') and then enumerates exactly the three finding classes it reports (orphan nodes, unresolved reviews, dangling edges). It also names the sibling `reflect` and how it differs, so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly contrasts with the alternative: `reflect` is 'the fuller version' that pairs findings with remedies and adds overdue tasks, stale chains and cortex candidates. This tells the agent when to pick lint versus reflect, though it never states an explicit 'use this when' condition or any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
merge_initiativeA
Merge one initiative INTO another — the fix for a project whose memory ended up split across two names (a typo, a separator, a translated alias). Re-homes every node and edge of source into target in ONE step and removes source. Nothing is forgotten: a merge only moves memberships, so unlike attach-then-delete_initiative it cannot lose a node you missed. target must already exist (use rename_initiative to move to a fresh name). target keeps its own share policy. Local only. reflect lists the candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Initiative to merge FROM. It stops existing; every node and edge in it becomes a member of `target`. Nothing is forgotten. | |
| target | Yes | Initiative to merge INTO — must already exist. It keeps its own share policy when both have one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that source is removed, that only memberships move, that target keeps its own share policy, and that it is 'Local only'. It omits auth/permission requirements and any reversibility statement, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then layers rationale, contrast with alternatives, and prerequisites. Dense but every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param mutation with no output schema, the description covers the full lifecycle: what is destroyed, what is preserved, preconditions on target, and the local-only scope. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the 'target must already exist' constraint and the share-policy behavior, but adds little beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (merge) and resource (initiative) with directionality (INTO), and explicitly frames the use case (split memory across two names). It distinguishes itself from attach, delete_initiative, and rename_initiative by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use it (split project memory from a typo/separator/alias), contrasts it against attach-then-delete_initiative, routes fresh-name cases to rename_initiative, and points to reflect for finding candidates.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
neighboursA
List every node ONE HOP from a node, in BOTH directions, across ALL edge types — the way to discover what a memory is connected to. drill follows only derived_from and part_of; this follows all twelve (refers_to, contradicts, supersedes, causal, temporal, blocks, targets, verifies, falsifies, part_of, derived_from, consolidated_to), so a contradiction or a supersession shows up here even when drill reports nothing. Each line names the edge TYPE and which way it points: —[type]→ outgoing, ←[type]— incoming. Optional edge_type (one or more, comma-separated) narrows to certain types; omit for all. between a b answers about a specific PAIR; drill shows the derived_from/part_of tree.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name or UUIDv7 id whose neighbours to list. | |
| edge_type | No | Optional edge-type filter — one or more, comma-separated; omit for ALL types. Closed vocabulary: `refers_to`, `derived_from`, `supersedes`, `causal`, `temporal`, `contradicts`, `part_of`, `blocks`, `targets`, `verifies`, `falsifies`, `consolidated_to`. Snake_case or kebab-case. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the return shape ('each line names the edge TYPE and which way it points' with the `—[type]→` / `←[type]—` notation) plus the complete edge vocabulary. It stops short of stating read-only safety, result-count limits, or pagination behavior, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then layers differentiation and output format. Every sentence carries information, though the dense middle sentence enumerating all twelve edge types is heavy and partially duplicated by the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a listing tool with no output schema and no annotations, the description covers output format and sibling differentiation well. The only real gap is the undocumented `initiative` parameter, which neither the schema nor the description explains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the description restates that `edge_type` accepts comma-separated values and that omitting it returns all types, but the schema already says this. The `initiative` parameter has no schema description and is never mentioned in the description, so the gap is not compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List every node ONE HOP from a node'), and pins the scope precisely: both directions, all twelve edge types. It explicitly distinguishes itself from sibling tools `drill` and `between`, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternatives with conditions: `drill` covers only the derived_from/part_of tree, `between a b` answers about a specific pair, while this tool surfaces contradictions and supersessions that `drill` would miss. The when-to-use decision is fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outcomesA
List the initiative's archival OUTCOMES — what the work actually concluded. This is the highest-value read for a fresh agent: results, not the working notes that produced them. trace walks any of them back to its sources; ideas lists the proposals that have not become results yet.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | Optional initiative to scope the operation to. When omitted, reads are cross-initiative; mutations end up un-tagged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'List' and 'highest-value read' imply a read-only operation, but it does not state permissions, side-effect guarantees, pagination, or return shape. It adds conceptual context (results vs working notes) but leaves important behavioral details unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences with zero waste, front-loading what the tool lists before adding selection guidance and sibling routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool, the description supplies strong purpose and routing context. It omits return-shape details, which is a minor gap given no output schema exists, but the conceptual output ('what the work actually concluded') is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents the optional initiative parameter. The description adds no additional parameter semantics, matching the baseline for well-described schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
State a specific verb 'List' and resource 'archival OUTCOMES', explains that these are final results, and explicitly distinguishes from siblings trace and ideas. An agent can identify what this tool returns versus adjacent tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says this is the highest-value read for a fresh agent and routes to alternatives: trace for walking sources back, ideas for proposals that are not yet results. The when-to-use and sibling alternatives are directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
overviewB
Print a terminal-readable map of the substrate: counts by tier/type, provenance forests, open questions, edge stats.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | Optional initiative to scope the operation to. When omitted, reads are cross-initiative; mutations end up un-tagged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. "Print" and "counts/stats" strongly imply a read-only, non-mutating display operation, which is meaningful transparency, but it says nothing about cost, output size, whether it respects the initiative scope, or whether anything is written.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, and the core action ("Print a... map of the substrate") is front-loaded before the content list. The tier/type/forest jargon is dense but every token carries content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does list the map's contents, which helps an agent anticipate the return. However, it omits any hint of output size, formatting beyond "terminal-readable," or how the optional initiative scope affects the result, leaving a read tool under-explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single optional "initiative" parameter is fully documented in the schema, including the cross-initiative default behavior. The description adds no parameter-level syntax or semantics beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ("Print") and names the resource ("map of the substrate"), then enumerates exactly what the map contains: counts by tier/type, provenance forests, open questions, edge stats. It is not a tautology of the name "overview," but it never distinguishes itself from sibling summary/inspection tools like board, history, or recent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no condition selecting this over the many siblings (board, board_status, history, recent, ideas), and no prerequisites or exclusions. The agent must infer the use case purely from the output list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pathA
Compute the shortest weighted path between two nodes WITHOUT writing anything — a preview. Edge weight is the cost, so stronger links (higher weight) make shorter paths. chain saves the same path as a recallable trail; use path to look first when you are not sure the two are meaningfully connected.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | End node name or id. | |
| from | Yes | Start node name or id. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the key behavioural trait ('WITHOUT writing anything — a preview') and explains the unusual weight semantics (higher weight = shorter path, since weight is cost), which an agent cannot infer from the schema. It stops short of stating permissions or pagination/return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the non-mutating guarantee, then adds the weight caveat and the routing hint. Three compact clauses that each carry information, with only slight density overhead in the weight explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must cover behaviour and result shape. It covers non-mutation and edge-weight semantics well, but omits what a path result contains and leaves the `initiative` parameter unexplained, so an agent has gaps for a 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: `from` and `to` are documented in the schema, but `initiative` has no description anywhere. The description adds no parameter-level meaning (no format, id-vs-name guidance, or explanation of `initiative`), so it neither extends nor compensates for the coverage gap. Baseline 3 for mid coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Compute the shortest weighted path between two nodes') and immediately differentiates itself from the sibling `chain` by stating it is a non-writing preview. An agent can distinguish it from neighbouring graph tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (`chain` saves the path as a recallable trail) and gives a selection condition: use `path` first 'when you are not sure the two are meaningfully connected.' Clear when-to-use guidance with one named alternative, though no explicit when-not-to-use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pinC
Pin a node to the active window. Accepts either a name or a UUIDv7 id.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name or UUIDv7 id. | |
| reason | Yes | Why the node deserves a place in the active window. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but only says what the input can be. It does not disclose what pinning actually does to state, whether it is reversible, what happens on repeat pins, or how it interacts with the active window — all relevant for a mutation-style verb.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. Brevity is appropriate, though the second sentence spends space repeating the schema rather than adding new meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must define the operation's effect and context. It leaves 'active window', the purpose of the required 'reason' field, and the outcome of pinning entirely unexplained, so an agent cannot fully reason about invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents 'name' as 'Node name or UUIDv7 id', so the description's second sentence is largely a restatement. More importantly, the required 'reason' parameter and the optional 'initiative' parameter are never explained in the description, leaving semantical gaps at 67% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('pin') and resource ('a node to the active window'), which lets an agent distinguish it from the inverse sibling 'unpin'. It does not clarify what the 'active window' concept entails or how it differs from related siblings like 'slot' or 'settle'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this rather than alternatives such as 'slot', 'settle', or 'attach'. The mirrored 'unpin' sibling implies a symmetry, but the description never states the condition that selects pinning.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
policyA
Read or set an initiative's cloud sharing policy (Gate 1). Omit both arguments to read. policy says WHETHER it may leave: private (default for any initiative — never leaves), team (shared nodes may sync), ask. clouds says WHERE TO: a comma-separated list restricting the initiative to those clouds, empty string to clear. An initiative with no list may go to any configured cloud, so this changes nothing until you ask for it.
| Name | Required | Description | Default |
|---|---|---|---|
| clouds | No | Restrict this initiative to named clouds — comma or space separated. `policy` says WHETHER an initiative may leave; this says WHERE TO. An initiative with no list may go to any configured cloud, which is how every initiative behaves until this is set. Pass an empty string to clear the restriction. Set independently of `policy`, so restricting does not re-open and re-opening does not un-restrict. | |
| policy | No | New policy: `private` (default, never leaves), `team` (shared nodes may sync), or `ask`. Omit to leave it as it is. | |
| initiative | Yes | Initiative whose cloud sharing policy to read or set. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the default (private = never leaves), the meaning of each policy value, that empty string clears the cloud list, and that a no-list initiative may go to any cloud so setting clouds 'changes nothing until you ask for it.' It omits auth requirements and return format, but the state-changing semantics are unusually well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and every sentence carries information; there is no filler. The only blemish is the unexplained '(Gate 1)' jargon and the density of backticked terms, which slightly reduce clarity without adding bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read/write tool with no annotations and no output schema, the description covers the read trigger, the defaults, and the clearing behavior. The one real gap is that it never says what a read returns, which an agent selecting it in read mode would want, but otherwise it is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema descriptions are themselves detailed and largely restate the same 'WHETHER vs WHERE TO' distinction. The description adds little beyond what the schema already provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb pair and resource: 'Read or set an initiative's cloud sharing policy.' It clearly separates the read mode from the set mode, so the agent knows exactly what the tool operates on. It does not, however, differentiate itself from likely siblings such as share/unshare/clouds, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage rule — 'Omit both arguments to read' — which tells the agent how to trigger read vs write. But there is no guidance on when to choose this tool over alternatives like share/unshare, nor any exclusions or prerequisites. Implied usage only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pullB
Pull a shared node from the cloud into the local vault by id, attaching it to the given initiative — the recall mechanism for team knowledge you don't have locally yet. In a multi-cloud setup pass cloud to target a specific cloud.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | UUIDv7 id of the cloud node to materialise into the local vault. | |
| cloud | No | Source cloud name in a multi-cloud setup. Omit for the default cloud. | |
| initiative | Yes | Initiative to attach the pulled node to locally. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It discloses the core effect (materialising a cloud node locally and attaching it to an initiative), but says nothing about idempotency, what happens if the node already exists locally, permission/auth requirements, or the return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the verb and resource, then the multi-cloud caveat. Little waste; the 'recall mechanism' framing adds context without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema mutation tool, the description covers the core flow but omits conflict/already-exists behavior, auth needs, and the local result shape. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three params are already documented. The description restates that `cloud` targets a specific cloud in multi-cloud setups, adding no format or behavioral detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Pull a shared node from the cloud into the local vault by id, attaching it to the given initiative') with a clear directional scope. However, calling it 'the recall mechanism' risks confusion with the sibling tools `recall` and `cloud_recall`, which are not explicitly distinguished.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies when to use it ('team knowledge you don't have locally yet'), which is useful context, but names no alternatives (`recall`, `import`, `cloud_recall`) and gives no when-not guidance. Usage is inferred rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recallA
Look up a node id by EXACT name — no fuzziness, no stemming. Returns the id alone, so follow it with at <name> for the full text or drill <name> for its neighbours. When you don't know the exact name, search is the verb; a miss here tells you whether the name lives in another initiative, is spelled differently, or is absent.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it does disclose real behavior: exact-match-only resolution, that the return is the id alone, and how to chain for full text or neighbours. It omits whether the call is read-only and what permissions/scope it needs, which for a no-annotation tool is a remaining gap, but the substantive behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with zero filler: match semantics, return shape plus chaining, then alternative/failure routing. Each sentence earns its place and the most decision-relevant fact (exact match) leads.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by stating the return is just the id and how to follow up. Combined with alternatives and miss semantics, an agent has what it needs to call correctly; only the `initiative` parameter's role and any auth requirement remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: the `name` param is documented in-schema (including UUIDv7 acceptance) but `initiative` has no description. The description's note that a miss reveals whether the name 'lives in another initiative' hints at scoping behavior but never states what the `initiative` parameter does, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Look up a node id') and pins down the matching semantics ('EXACT name — no fuzziness, no stemming'). It explicitly contrasts with the sibling `search`, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('by EXACT name'), the alternative for the opposite case ('When you don't know the exact name, `search` is the verb'), and follow-up actions (`at <name>`, `drill <name>`). Even the failure case is explained (a miss disambiguates spelling/initiative/absence).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recentA
List what was CAPTURED within the time window (defaults 24h) — every kind of write, not only episodes: references, claims, tasks and jots all count. This is the verb that answers "did what I just write land?", so a zero here really means nothing was captured. Use since like 30m, 3h, 2d, or raw seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Time window (e.g. `30m`, `3h`, `2d`, raw seconds). Defaults to 24h. | 24h |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: the 24h default window, the fact that all write kinds count, and crucially the empty-result semantics ('a zero here really means nothing was captured'). It stops short of ordering, result caps, or pagination, which matters for a list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the scope statement and the default window before the framing and the `since` format note. Slight redundancy between the 'every kind of write' list and the 'did what I just write land' restatement, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no annotations and no output schema, the description covers purpose, timing default, and empty-result meaning well. It omits any explanation of the `initiative` filter and any hint of return shape or ordering, so an agent still has gaps before calling it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `since` is already documented in the schema with the same `30m`/`3h`/`2d` examples, so the description largely repeats it. `initiative` has no schema description and the description says nothing about it, leaving half the parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (everything captured in the time window), then explicitly widens scope beyond the obvious sibling: 'every kind of write, not only episodes: references, claims, tasks and jots all count.' An agent can distinguish this from `episode` without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the use case concretely — 'the verb that answers "did what I just write land?"' — which tells the agent when to reach for it. It implicitly contrasts with the `episode` sibling but never names a when-not case or an alternative such as `search`/`recall` for historical lookups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rechainB
Refresh a chain the graph has outgrown. With no to, regenerate it — recompute the shortest path between its current endpoints (picks up new edges / re-weights). With to, extend the trail out to that node. Keeps the chain's id, name, and summary.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Omit to regenerate (recompute the shortest path between the chain's current endpoints). Provide a node name/id to instead extend the trail out to it. | |
| chain | Yes | Chain name or UUIDv7 id to refresh. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden. It discloses real behavior (recomputes shortest path, picks up new edges/re-weights) and what is preserved (id, name, summary), but says nothing about permissions, reversibility, or what edges/nodes are lost during a refresh.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in one sentence, then the two modes, with no filler. Tight and readable, though the dual-mode branching takes slightly more parsing than a single-clause tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, it covers the core behavior and preservation guarantee but leaves gaps: no `initiative` semantics, no permission/reversibility context, and no indication of what the refresh discards.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67% and the description mostly restates what the schema already documents for `to` (regenerate vs extend), so it adds little beyond structured data. The `initiative` parameter is left unexplained in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('refresh a chain') and splits it into two clearly named modes (regenerate vs extend). It's clear what the tool does, though it never distinguishes itself from siblings like 'chain' or 'reweight'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the condition that selects each mode (omit `to` = regenerate, provide `to` = extend), which is useful. But it offers no guidance on when to reach for this tool versus the many alternatives (chain, reweight, path, between).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reflectA
Reflect on the store: a computed maintenance work-list with how to act on each part — orphans to link, open reviews to resolve, chains gone stale (rechain), settled work to promote into cortex or drop to cold, and shared/cloud items that need YOUR sign-off (never auto-rebalanced). The last two are candidates to judge one at a time, not a batch to apply: cortex loads whole into every session, so the report prices the move (cortex N -> M). Run it when a piece of work ends, and before a session stops — that is the moment nothing else marks.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | No | Optional initiative to scope the operation to. When omitted, reads are cross-initiative; mutations end up un-tagged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that items are never auto-applied ("never auto-rebalanced", "not a batch to apply") and that the report prices the move (cortex N -> M). However, it never states explicitly whether reflect itself mutates the store or is strictly read-only, which is the key behavioral fact an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool does, but the single dense paragraph uses heavy em-dash clauses and unglossed domain jargon, making it harder to parse than necessary. Nothing is obviously redundant, yet the structure is not clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey return content; it does describe the report's categories and the priced cortex move reasonably. It stops short of describing the report's actual shape or ordering, leaving a small gap for a report-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the lone optional 'initiative' param is already fully documented in the schema, including the cross-initiative read / un-tagged mutation behavior. The description adds nothing beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ("Reflect on the store") and enumerates the concrete contents of the work-list: orphans, open reviews, stale chains, settled work, cloud sign-offs. It does not explicitly distinguish itself from maintenance-adjacent siblings like lint, hygiene, or overview, and undefined jargon (cortex, cold) slightly blurs the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ("Run it when a piece of work ends, and before a session stops") and clarifies how to act on the last two categories (judge one at a time, not a batch). It does not name alternative sibling tools to use instead for those follow-up actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refuteA
Mark a hypothesis as refuted. by (the falsifying counter-evidence node) is optional, same as for confirm.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | Evidence node name or id — linked `verifies` for `confirm`, `falsifies` for `refute`. OPTIONAL: record the verdict even with nothing to point at yet, rather than leaving the claim tagged `open` with the answer buried in its prose. (`inconclusive` writes no edge at all, so it never needs one.) | |
| hypothesis | Yes | Hypothesis name or id. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It reveals that `refute` mutates state (marks a hypothesis) and that `by` is optional, allowing recording without evidence. However, it does not disclose permissions, reversibility, side effects, or what happens to existing verdicts, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: two sentences that front-load the core action and then address the optional parameter. No wasted words; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a mutation tool, the description is incomplete: it lacks details on permissions, side effects, return values, and how it differs from other verdict-setting tools like `confirm` or `inconclusive`. It covers the basic action and optional parameter but leaves much unsaid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description clarifies that `by` is the falsifying counter-evidence node and is optional, which adds semantic meaning beyond the schema. The schema already describes `by` in detail, but the description reinforces that it's optional and same as `confirm`. The `initiative` parameter is not mentioned, but schema provides some description (67% coverage), so this is not a major gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Mark a hypothesis as refuted'), which is clear. However, it does not differentiate from siblings like `confirm` or `inconclusive` beyond the verb itself, leaving the agent to infer that this is the falsification counterpart of confirmation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions `by` is optional 'same as for `confirm`', implying a parallel usage, but does not explicitly state when to use `refute` vs `confirm` or `inconclusive`. No prerequisites or exclusions are given; it's implied that this is for falsifying a hypothesis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_initiativeA
Rename an initiative — moves all its nodes, edges, and sharing policy to the new name (fails if the new name already exists). Local by default. Pass cloud="<name>" to ALSO rename it in that shared cloud, which is team-wide and affects everyone; with several clouds configured the name is required, because this cannot be undone in the wrong one.
| Name | Required | Description | Default |
|---|---|---|---|
| new | Yes | New initiative name (must not already exist). | |
| old | Yes | Current initiative name. | |
| cloud | No | Name of a cloud to ALSO rename this initiative in — team-wide, and it affects everyone using it. Omit for a local-only rename. Never defaulted: with several clouds configured, an unnamed cloud rename would be a guess at which team it disrupts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the failure mode ('fails if the new name already exists'), the blast radius of the cloud path ('team-wide and affects everyone'), and the irreversibility risk ('cannot be undone in the wrong one'). It stops short of stating permission/auth requirements or whether a local rename is reversible, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the local-vs-cloud distinction in a single compact block with no filler. The cloud clause is somewhat detailed and partially overlaps the schema, but every sentence carries a distinct constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no output schema and no annotations, the description covers behavior, side effects, failure mode, and the risky cloud path well. The remaining gap is auth/permission context and local-rename reversibility, which are minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters, including the cloud semantics. The description reinforces the 'ALSO' behavior and the name-required-with-multiple-clouds rule, but this largely restates the schema rather than adding new syntax or format detail. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Rename an initiative') and even specifies the scope of what moves: nodes, edges, and sharing policy. An agent can immediately tell this apart from siblings like merge_initiative or delete_initiative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the default behavior ('Local by default'), the condition for the alternative ('Pass cloud="<name>" to ALSO rename it in that shared cloud'), and a hard prerequisite ('with several clouds configured the name is required'). Both when-to-use and when-name-is-required are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolveA
Resolve an open question by recording that by answers it — a supersedes edge from the answer to the question, so the question stays readable as history instead of being deleted. For a doubt raised by flag, close_review is the matching verb.
| Name | Required | Description | Default |
|---|---|---|---|
| by | Yes | Answer / resolution node name. | |
| question | Yes | Question node name. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well on the key trait: it makes explicit that the question is not deleted but kept 'readable as history'. It does not address permissions, whether the edge is reversible, or what the response contains, leaving some behavioral gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and mechanism, then the sibling routing. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The purpose, mechanism, and one alternative are covered, which is reasonable for a 3-parameter tool. But with no annotations, no output schema, and an undocumented `initiative` parameter, the definition leaves an agent guessing about scoping and return behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67% and the description compensates partially by explaining the roles of `question` and `by` ('recording that `by` answers it'). However, the `initiative` parameter is described nowhere in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resolve), resource (open question), and the exact mechanism: a supersedes edge from the answer to the question. It also implicitly distinguishes itself from siblings by naming close_review as the matching verb for flag-raised doubts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use context ('resolve an open question') and routes the agent away for one case: doubts raised by flag should use close_review. It does not cover other nearby siblings such as supersede or settle, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reviseA
Rewrite a node's body and/or rename it IN PLACE — the id survives, and history shows both versions. This is the verb for correcting or extending a node. Use supersede instead when the change is big enough to deserve a new identity, and settle when the node is not changing at all, just finished.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | New body. If omitted, keeps current. | |
| name | Yes | Node name. | |
| rename | No | New name. If omitted, keeps current. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses the key behavioral traits: the mutation is IN PLACE, the id survives, and `history` retains both versions (auditable/recoverable). It omits permission requirements and any failure modes, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, zero filler, with the core in-place semantics front-loaded before the sibling routing. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the essential semantics, reversibility via history, and alternatives well. The gaps are the undocumented `initiative` param and lack of permission/error context, neither of which is fatal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the schema already documents body, name, and rename individually. The description only reaffirms that body and name are the things being changed; it adds no syntax, format, or meaning for the undocumented `initiative` param, so the baseline 3 holds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Rewrite a node's body and/or rename it') and immediately distinguishes itself from two named siblings. An agent can tell revise apart from supersede and settle without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules: use this to correct/extend, use `supersede` when the change deserves a new identity, use `settle` when nothing changes. Both alternatives and their selecting conditions are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reweightB
Set an existing edge's connection strength (weight 0..1) in place. Stronger edges make shorter knowledge-chain paths; use to tune which links matter after the fact.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Destination node name or id — resolved in the active initiative first, then across all initiatives. | |
| from | Yes | Source node name or id — resolved in the active initiative first, then across all initiatives. | |
| cloud | No | Which cloud to mirror this edge change to, when BOTH endpoints are already shared. Omit with one cloud configured (it is unambiguous) or when the endpoints are local. With several configured and none named, the local edit still happens and the result says the edge was not mirrored. | |
| weight | Yes | New connection strength in `0..1` (1 = strong → shorter chain paths). | |
| edge_type | No | Edge type — closed vocabulary, one of: `refers_to` (default), `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`. | refers_to |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. 'In place' and 'existing edge' disclose that this mutates an already-present edge, and it explains the downstream effect of weight on chain paths, but it omits failure behavior (edge not found), whether weight 0 removes the edge, and any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the mutation and its effect front-loaded; nothing is wasted. The parenthetical 'weight 0..1' duplicates the schema slightly but costs almost nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter mutation with no annotations and no output schema, the description covers purpose, effect and rough timing but leaves error handling and the weight-0 boundary case unaddressed. The cloud mirroring behavior is only covered by the schema, not the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already documents all six parameters including the cloud-mirroring nuance and the edge_type vocabulary. The description only restates the weight range and its chain-path effect, adding no syntax or format detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb ('Set') and resource ('an existing edge's connection strength'), and the word 'existing' implicitly distinguishes it from the create-style sibling `link`. It never names a sibling outright, so an agent must infer the boundary rather than being told it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'use to tune which links matter after the fact' implies the post-hoc adjustment scenario, i.e. after edges exist. There is no explicit when-not guidance, no prerequisite statement, and no named alternative such as `link` for creating or `unlink` for removing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Full-text search across name and body via Cozo FTS. No stemming — search the form you wrote. For inflection-tolerant matching across any language append *: утечк* finds утечку/утечке, token* finds tokens/tokenize. Search in the SAME language as the original capture, not in English. Results are ordered by score, then newest-first within equal scores.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results. Capped at 50 internally. | |
| query | Yes | Cozo FTS query (`AND`/`OR`/`NOT`, `"phrase"`, `prefix*`). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the search engine behavior, lack of stemming, wildcard inflection behavior, language expectations, and result ordering. It stops short of covering permissions, rate limits, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core capability, then gives compact, practical query guidance and ordering behavior. Every sentence earns its place and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no annotations and no output schema, the description adequately covers matching behavior and ordering. It is incomplete because it never explains the optional `initiative` filter, which could matter when an agent scopes a search.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds real meaning beyond the schema for the `query` parameter, especially no-stemming and prefix matching semantics. But schema coverage is 67%, and the `initiative` parameter is undocumented in both schema and description, leaving one of three parameters without semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Full-text search across name and body via Cozo FTS.' That is clear enough to distinguish from generic retrieval siblings, but it does not explicitly name which sibling to use instead of, or alongside, this one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful query guidance such as 'No stemming — search the form you wrote' and when to append `*` for inflection-tolerant matching. However, it gives no explicit when-to-use/when-not-to-use guidance relative to alternatives like recall, trace, or surface.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_statusA
Move a task to a board column: sets its status:<key>, strictly validated against the initiative's board registry (unknown status is refused). The general form of done.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Task node name or id to move. | |
| status | Yes | Target status (board column key) — must exist in the initiative's board. | |
| initiative | No | Initiative whose board defines the valid statuses (required). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden, and it does disclose one real behavioral trait: strict validation against the initiative's board registry with unknown statuses refused. It omits permissions/authorization needs, reversibility, and what happens to any prior status, which are meaningful gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, front-loaded with the action and the field set, then the validation constraint and the sibling relationship. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param mutation tool with no output schema and no annotations, the description covers purpose, the key parameter's validation semantics, and the sibling relationship. It leaves out authorization requirements and return/confirmation behavior, so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds genuine meaning: it explains that `status` is a board column key validated against the initiative's board registry and that `initiative` supplies that registry. This goes beyond restating the schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ('Move a task to a board column') and the exact field it sets ('status:<key>'), and explicitly positions itself against a sibling ('The general form of `done`'). An agent can distinguish this from `done` without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the tool as the generalized version of `done`, which gives an agent a clear routing cue between the two. It stops short of an explicit when-to-use/when-not-to-use rule or naming other alternatives, so it is clear context rather than full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
settleA
Promote a node that stopped changing from the operational tier into the archival one — this is how knowledge hardens. settle <name> ALONE IS ENOUGH: with no new_* the node keeps its name and its full body, and the type is derived (episode/task/experiment/hypothesis → outcome; draft/scratch → idea). Provenance via derived_from is replicated across the tier boundary, and manual tags come with it. Don't demote finished work to a cold layer instead — a layer is how eagerly a node loads, a tier is whether it is still in flight.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Source node name (or id). | |
| new_body | No | Body for the successor. OPTIONAL — omit to carry the source's body over unchanged (in full, not the excerpt). | |
| new_name | No | Name for the successor. OPTIONAL — omit to carry the source's name over unchanged. | |
| new_type | No | Type for the successor. OPTIONAL — omit to promote in place, and the type is derived from the source (`episode`/`task`/`experiment`/ `hypothesis` → `outcome`; `draft`/`scratch` → `idea`; anything already archival keeps itself). The chosen type is always printed back. A CLOSED vocabulary, one of exactly these: `episode`, `task`, `checklist`, `roadmap`, `experiment`, `hypothesis`, `scratch`, `draft`, `audit_event`, `chain`, `board`, `idea`, `outcome`, `reference`, `concept`, `entity`, `summary`. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that no new_* fields means name, body, provenance (derived_from) and manual tags are all carried across the tier boundary, and states the type derivation rules. It omits reversibility and any permission/irreversibility caveats for this mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the 'settle <name> ALONE IS ENOUGH' rule before elaborating. A few stylistic flourishes ('this is how knowledge hardens') add voice but not operational value, keeping it just shy of maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no annotations and no output schema, it covers purpose, defaults, type derivation, and tier-vs-layer semantics well. The undocumented `initiative` parameter and absence of any return/reversibility note are the only meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents defaults for new_body/new_name/new_type. The description largely restates those defaults and the derivation rules rather than adding new meaning, and the `initiative` parameter is left unexplained in both places. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (promote) and resource (a node) plus the direction of movement (operational tier → archival tier). It explicitly contrasts the concept against the sibling notion of a 'layer' tool, so an agent can tell 'settle' apart from demotion tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('a node that stopped changing') and an explicit exclusion ('Don't demote finished work to a cold layer instead'), clarifying a likely confusion with the `layer` sibling. It falls short of naming the inverse `unsettle` tool or stating prerequisites, so it's strong but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slotA
Make a node the live holder of a ROLE in an initiative — handoff, entrypoint, queue, prod-state. A role holds exactly one node: taking it archives the previous holder to cold and links supersedes from the new holder to it, so a project can never end up with three current handoffs. Nothing is deleted; the predecessor stays readable via at / surface layers=cold.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node to install as the holder. Name or UUIDv7 id. | |
| slot | Yes | The role — `handoff`, `entrypoint`, `queue`, `prod-state`, or any name you keep to one live node. | |
| initiative | Yes | Initiative the role belongs to. Slots are per-initiative: the same role name in another initiative is a different slot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses the side effect (previous holder archived to `cold`), the link created (`supersedes`), the reversibility guarantee (nothing is deleted), and how to recover the predecessor (`at` / `surface layers=cold`). This is exactly the mutation semantics an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with no waste: what it does, the exclusivity invariant plus side effect, and the non-destructive recovery path. Dense but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-required-param mutation tool with no annotations and no output schema, the description covers behavior and recovery well. It omits return/failure semantics and any permission prerequisites, but the core semantics are complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds real meaning beyond the schema by stating that a role holds exactly one node and that taking it archives the prior holder, which is not captured in the field descriptions. The role names and per-initiative scoping are, however, largely duplicative of the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: making a node the live holder of a named role, with concrete role examples (handoff, entrypoint, queue, prod-state). It implicitly routes to `at`/`surface layers=cold` for reading the predecessor, but never names `slots` or `unslot` as alternatives, so it stops short of sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through behavior rather than stated: the exclusivity rule (a role holds exactly one node) tells the agent this is the way to establish/replace a role holder. There is no explicit when-to-use, when-not, or alternative (e.g. `supersede`, `link`) named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slotsB
List the filled roles of an initiative and which node holds each.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | Yes | Initiative to report on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose a meaningful trait — only *filled* roles are returned, implying empty slots are omitted — and it sketches the return shape (roles plus the claiming node). It stops short of covering permissions, pagination, or the empty-initiative case.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler that front-loads the action and the subject. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, the description conveys both the operation and the conceptual return shape, which is enough to call it correctly. It could still clarify how 'roles' and 'nodes' relate or what an empty result means, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already documents the 'initiative' argument. The description's 'of an initiative' phrasing aligns with it but adds no format, scope, or identifier details beyond the schema. Baseline 3 for a well-covered single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ('List') with a concrete resource ('filled roles of an initiative') and states the payload ('which node holds each'). It clearly distinguishes a read/list operation from the mutation-oriented siblings 'slot' and 'unslot', though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to reach for this versus related tools like 'slot', 'unslot', 'initiatives', or 'overview'. The read-only nature is only implied by the verb 'List', and no conditions or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supersedeA
Replace a node with a fresh one carrying new content, connected by a supersedes edge from the successor to the node it replaces. Use when the change is large enough to warrant a new identity; revise edits in place instead. new_type is optional — it defaults to the old node's.
| Name | Required | Description | Default |
|---|---|---|---|
| old | Yes | Old node name (or id). | |
| tier | No | Tier override. A CLOSED vocabulary, one of exactly these: `operational`, `archival`. Defaults from the type. (Not the memory *layer* — that is `core`/`hot`/`warm`/`cold`/`frozen`, set by `layer`.) | |
| new_body | Yes | New node body. | |
| new_name | Yes | New node name. | |
| new_type | No | Type for the successor. OPTIONAL — omit to keep the old node's own type, which is what a straight replacement wants. A CLOSED vocabulary, one of exactly these: `episode`, `task`, `checklist`, `roadmap`, `experiment`, `hypothesis`, `scratch`, `draft`, `audit_event`, `chain`, `board`, `idea`, `outcome`, `reference`, `concept`, `entity`, `summary`. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does explain the core mutation mechanics: a new node is created and a supersedes edge is added from successor to the old node, so old is retained as the edge target. It also states the new_type default. However, it does not mention permission requirements, whether the old node's status changes, or any failure behavior, leaving some behavioral gaps for a mutation tool without annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with zero waste. The primary operation is front-loaded, the alternative routing follows immediately, and the parameter note is a brief final fragment. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, the description covers the essential semantics: what is created, how it is linked, and when to prefer the alternative. It omits return-value or error behavior, but without an output schema that is less critical; the remaining gap is minor side-effect detail on the old node.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already documents most parameters including the new_type default. The description repeats the optionality and default of new_type but adds no syntax, format, or constraint details beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: replace a node with a new identity, connected by a supersedes edge. It explicitly contrasts with the sibling `revise`, so an agent can distinguish the two without inspecting both schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('when the change is large enough to warrant a new identity') and names the alternative tool (`revise`) with its opposite condition. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
surfaceA
Explicit layered recall — surface nodes from specific memory layers (default cold,frozen) that awake does not load. Use when you deliberately need archived/not-surfaced material. layers is a comma/space list; scoped to initiative when given.
| Name | Required | Description | Default |
|---|---|---|---|
| layers | No | Comma/space-separated memory layers to surface, e.g. `cold,frozen` or `cold`. Defaults to `cold,frozen` when omitted. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the default layers (cold,frozen), the fact that awake won't load them, and the initiative scoping behavior, but it never states the operation is a safe read, nor describes the return shape or any limits. Partial disclosure for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences that lead with the core concept and then cover usage and parameters, with no filler. The em-dash clauses add density but each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should ideally cover the read-only nature and return shape. It covers purpose, usage, and parameter scoping adequately, but the lack of any statement about what is returned or the safety profile leaves a real gap for a tool with zero structured support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, and the undocumented 'initiative' parameter is the one the description compensates for by explaining scoping ('scoped to initiative when given'). The layers format is repeated from the schema rather than added to, so net value is moderate — the baseline 3 for half coverage fits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('surface') and resource ('nodes from specific memory layers'), and explicitly distinguishes itself from the sibling 'awake' by noting it loads layers 'that awake does not load'. An agent can tell it apart from awake and recall without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use condition ('Use when you deliberately need archived/not-surfaced material') and names the contrasting alternative (awake). It lacks an explicit when-not-to-use or a pointer to a more general sibling like recall, but the routing context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_reviewA
Batch sync-review of a team initiative's still-local nodes: splits them into PROPOSE SHARE (guard-clean) vs KEEP LOCAL (secret-guard flagged). Review once, then share the approved ones — low-friction periodic sharing instead of deciding per capture.
| Name | Required | Description | Default |
|---|---|---|---|
| initiative | Yes | Team initiative to review still-local nodes for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It describes the classification concept (guard-clean vs secret-guard flagged) but never states whether the call mutates state, whether review results persist, whether it requires specific permissions, or whether it is idempotent/safe to re-run. For a tool whose name suggests state reconciliation, these gaps matter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core operation front-loaded and the workflow guidance trailing. No filler, though the all-caps bucket names and jargon make it slightly harder to parse at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully sketches what the tool yields (two classified buckets) and how the result feeds `share`. With a single well-documented parameter, the main remaining gap is the undisclosed state/persistence behavior, which is a transparency issue more than a completeness one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% — the schema already says `initiative` is the team initiative to review. The description's framing ('still-local nodes for') mirrors rather than extends the schema text, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (batch sync-review of a team initiative's still-local nodes) and describes the concrete output split (PROPOSE SHARE vs KEEP LOCAL). It also distinguishes itself from the sibling `share`, which is named as the follow-up action rather than a duplicate. It stops short of a clean 'use this instead of X' framing, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: 'Review once, then `share` the approved ones — low-friction periodic sharing instead of deciding per capture.' This tells the agent this is a periodic batch triage step feeding `share`, contrasted with per-capture decisions. No explicit exclusions or prerequisites (e.g., when not to run it) are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
synthesiseA
Many-to-one consolidation — write one durable node from several seeds, with a derived_from edge to each so trace can walk back to them. Use when scattered observations have converged into a single finding; settle is the one-to-one version, which promotes a single node in place.
| Name | Required | Description | Default |
|---|---|---|---|
| from | Yes | Seed node names. | |
| tier | No | Tier override. A CLOSED vocabulary, one of exactly these: `operational`, `archival`. Defaults from the type. (Not to be confused with the memory *layer* — `core`/`hot`/`warm`/`cold`/`frozen` — which `layer` sets.) | |
| new_body | Yes | Body for the synthesised node. | |
| new_name | Yes | Name for the synthesised node. | |
| new_type | No | Type of the synthesised node (defaults `summary`). A CLOSED vocabulary, one of exactly these: `episode`, `task`, `checklist`, `roadmap`, `experiment`, `hypothesis`, `scratch`, `draft`, `audit_event`, `chain`, `board`, `idea`, `outcome`, `reference`, `concept`, `entity`, `summary`. | summary |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that edges (derived_from) are written so `trace` can walk back, which is real graph-side-effect info. However it says nothing about whether the seed nodes are consumed/retained, whether the operation is reversible, or any permission/conflict requirements for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the core operation and its graph effect are front-loaded, and the sibling contrast is placed last where it is cheapest to skip. Nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with an output schema absent but 83% schema coverage, the description covers the operation, graph effect, and sibling routing adequately. It is slightly thin on the optional parameters (tier/new_type/initiative) and on seed-node lifecycle, which keeps it below a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema already documents most parameters well (including the closed tier/new_type vocabularies). The description only implicitly maps 'seeds' to `from` and 'durable node' to `new_name`/`new_body`, adding no syntax or constraint detail beyond what the schema carries. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (write one durable node from several seeds) and adds the structural consequence (derived_from edges to each seed). It explicitly names and contrasts the sibling `settle` as the one-to-one variant, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('when scattered observations have converged into a single finding') and names the nearest alternative (`settle`) with the discriminator (one-to-one, in-place promotion). Both the trigger and the routing rule are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
taggedA
List nodes whose tags array contains the given tag — exact match. Common tag families: kind:<type> (observation, experiment, idea, reference, …), sig:<level> (low/medium/high), role:<role> (jot/review/synthesise/revised), lang:<code> (ru/en/mixed/other — auto-detected from body), topic:<word> (up to 5 auto-derived tokens — a node's MOST-MENTIONED words, weighted toward a name somebody chose; a compound like figma-макет is also tagged by its parts), status:<state> (hypotheses and tasks). Exact match, no stemming — but a miss comes back with the near tags that DO exist in scope, so an empty answer tells you what to ask instead. For loose matching over text use search prefix*. Newest-first when multiple match.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | Yes | Tag value (case-sensitive). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so well: it discloses exact-match/no-stemming behavior, that misses return near tags in scope, and that results are newest-first. That is meaningful operational context beyond the schema, though it says nothing about result size, pagination, or scoping by initiative.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first clause, then layers detail. It is one long sentence and the tag-family enumeration is verbose, but nearly every clause conveys a distinct, useful fact rather than restating the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema and no annotations, the description covers matching semantics, fallback behavior, and ordering well. The main gap is the undocumented `initiative` scoping parameter, which an agent would need to understand before constraining a query.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `tag` is documented in the schema as case-sensitive, but the description enriches it substantially by enumerating tag families (kind:, sig:, role:, lang:, topic:, status:) and their value shapes. The `initiative` parameter is undocumented in both schema and description, keeping this below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List nodes') plus the exact matching semantics ('tags array contains the given tag — exact match'). It also distinguishes itself from the sibling `search` by noting that loose text matching belongs there. An agent can tell this apart from `search` or `recall` without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative for loose matching ('For loose matching over text use `search prefix*`') and describes the miss/fallback behavior that tells the agent what to query instead. It gives clear usage context but does not enumerate when-not to use it against the many other retrieval siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
taskA
Capture a todo as a Task node. Auto-named from body. Tags: kind:task, status:open, optional due:YYYY-MM-DD. due accepts ISO date, RFC-3339, or future duration like 3d/2w. Pass link_to (with weight) to connect it to an existing node in this same call — an island is found only by exact name.
| Name | Required | Description | Default |
|---|---|---|---|
| due | No | Optional deadline. Accepts an ISO date (`2026-05-15`), an RFC-3339 datetime, or a future duration (`3d`, `2w`). Omit for tasks without a deadline. | |
| body | Yes | Free-form task description. | |
| layer | No | Optional memory layer stamped at creation: `core`, `hot`, `warm`, `cold`, or `frozen`. Defaults to `warm`. | |
| weight | No | How load-bearing the `link_to` edge is, 0..1. REQUIRED with `link_to`: it is the only signal knowledge chains route on, and there is no default because an unweighted graph makes every chain rank on noise. The capture still lands without it; the edge does not. | |
| link_to | No | Optional: the node this one connects to, by name or id — the edge is made in THIS call, while both ends are still in mind. Linking as a second step is the step nobody takes: one real vault reached 23 nodes and 0 edges with the nudge asking every time. Needs `weight`. | |
| edge_type | No | Edge type for `link_to` — same closed vocabulary as `link` (`refers_to` by default, `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: auto-naming, tags stamped at creation, and the link_to/weight coupling being made in one call rather than two. It omits permissions, reversibility, default layer effect, and what the call returns, so it is helpful but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and each sentence carries actionable content (tags, due formats, linking). Slight duplication of the due-format detail already in the schema is the only waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter capture tool with no output schema and no annotations, the description covers creation, tagging, deadline formats, and in-call linking well. The main gap is absence of return/error expectations, which the missing output schema would otherwise carry.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the schema already documents most parameters, including the due format and weight/link_to relationship. The description mostly restates these, only adding the edge-name-exactness caveat, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource: "Capture a todo as a Task node," and adds concrete side effects (auto-naming from body, default tags kind:task/status:open). It does not explicitly differentiate from near-siblings like jot or claim, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies when to use it via the linking nudge ("pass link_to ... in this same call") and the note that islands resolve only by exact name, which steers behavior. But there is no explicit when-not guidance or named alternative among the many capture-adjacent siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
traceA
Walk derived_from ancestors of a node back to its sources — where a conclusion CAME FROM. Use it before trusting a synthesised or settled node: it shows the raw material the claim was built on. why is the sibling verb for the reasoning trail; trace is the material one.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It conveys that this is a read-only graph walk over derived_from edges and that it surfaces 'raw material,' but it says nothing about return shape, depth limits, cost, or idempotency — modest disclosure for an annotation-free query tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and payoff, then the usage trigger, then the sibling contrast. Every sentence earns its place with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and 50% parameter coverage, the description does the core job of explaining purpose but leaves the `initiative` parameter's meaning and any return/edge semantics unaddressed. Adequate but with clear gaps for a two-parameter graph-traversal tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `name` is documented in the schema and merely implied by the word 'node' in the description, while `initiative` is undocumented in both schema and description. The description adds no format, resolution, or scoping detail beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: walking derived_from ancestors back to sources, framed as 'where a conclusion CAME FROM'. It explicitly distinguishes itself from the sibling verb `why`, so an agent can select between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('use it before trusting a synthesised or settled node') and names the alternative (`why` for the reasoning trail, `trace` for the material one), which effectively defines when this tool is the right pick versus its closest sibling. It stops short of stating when not to use it or prerequisite states.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlinkB
Retract a previously-asserted edge. Bi-temporal — historical reads still see it.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Destination node name or id — resolved in the active initiative first, then across all initiatives. | |
| from | Yes | Source node name or id — resolved in the active initiative first, then across all initiatives, so an edge may span initiatives. | |
| cloud | No | Which cloud to mirror this retraction to, when BOTH endpoints are already shared. Omit with one cloud configured (it is unambiguous) or when the endpoints are local. With several configured and none named, the local edit still happens and the result says the edge was not mirrored. | |
| edge_type | No | Edge type to retract — the same CLOSED vocabulary `link` writes: `refers_to` (default), `causal`, `derived_from`, `contradicts`, `part_of`, `blocks`, `targets`, `supersedes`, `verifies`, `falsifies`, `temporal`, `consolidated_to`. Snake_case or kebab-case both accepted. | refers_to |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden. It does add real value by disclosing the bi-temporal semantics — historical reads still observe the retracted edge — which tells the agent this is a soft/retractable rather than hard delete. It omits error behavior (nonexistent edge), idempotency, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the core action is front-loaded ahead of the qualifying temporal detail. Nothing can be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation with no annotations and no output schema, the description covers the key semantic (retraction is not erasure) but leaves gaps: what happens when the edge does not exist or is ambiguous, whether the call is idempotent, and what the result reports on partial cloud mirroring (only covered inside the schema).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents `from`, `to`, `cloud`, and `edge_type` in detail. The top-level description adds no parameter-level meaning at all, and `initiative` carries no description anywhere. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Retract") and resource ("a previously-asserted edge"), which clearly reads as the inverse of the sibling `link`. It does not explicitly name `link` or contrast with other removal-ish siblings (forget, supersede, refute), so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative-tool guidance is given. With siblings such as `supersede`, `refute`, `forget`, and `rechain` also removing or invalidating relationships, the agent gets no help choosing `unlink` over them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unpinC
Unpin a node. Accepts name or id.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Node name (also accepts a UUIDv7 id where the verb supports polymorphic resolution). | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a state mutation (removing a pin) but says nothing about idempotency, reversibility, permission requirements, or what happens if the node was not pinned. Only the identifier-acceptance note adds any behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences with zero filler, and the core action is front-loaded. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and an undocumented 'initiative' parameter, this is too thin. The agent cannot tell what unpinning entails in this domain or how the optional scoping parameter affects the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: 'name' is documented but 'initiative' has no description at all in either the schema or the description. 'Accepts name or id' largely repeats the schema's own note on 'name', and the optional 'initiative' parameter is never explained, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Unpin a node'), which is a clear operation. The inverse sibling 'pin' exists in the tool list, and the name/verb pairing makes the distinction inferable, but the description never explicitly names or contrasts it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not, or alternative guidance is given. 'Accepts name or id' is a parameter note, not usage context. The agent gets no signal about when unpinning is appropriate versus other state-changing siblings like forget or delete_initiative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unsettleA
Bring an archival node back into the operational tier — settle's mirror, for settled knowledge that turned out to still be in flight. unsettle <name> alone is enough: name, body and type all carry over unless you say otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Source node name (or id). | |
| new_body | No | Body for the successor. OPTIONAL — omit to carry the source's body over unchanged (in full, not the excerpt). | |
| new_name | No | Name for the successor. OPTIONAL — omit to carry the source's name over unchanged. | |
| new_type | No | Type for the successor. OPTIONAL — omit to promote in place, and the type is derived from the source (`episode`/`task`/`experiment`/ `hypothesis` → `outcome`; `draft`/`scratch` → `idea`; anything already archival keeps itself). The chosen type is always printed back. A CLOSED vocabulary, one of exactly these: `episode`, `task`, `checklist`, `roadmap`, `experiment`, `hypothesis`, `scratch`, `draft`, `audit_event`, `chain`, `board`, `idea`, `outcome`, `reference`, `concept`, `entity`, `summary`. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It discloses the state transition, that name/body/type carry over by default, and the 'unless you say otherwise' override, which is useful. However, it never says what happens to the source node (consumed, renamed, preserved?) even though the schema calls the output a 'successor', leaving the mutation's side effects ambiguous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the purpose and then the minimal-invocation default. Nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param mutation with no annotations and no output schema, the description covers purpose and defaults but omits the source node's fate, permission/auth requirements, and reversibility, leaving an agent with open questions about the operation's effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 80% (baseline 3), and the description adds real value by explaining the default carry-over behavior of new_name/new_body/new_type with no arguments needed. The `source` parameter is only loosely referenced via `<name>`, so it does not fully compensate for every param.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('bring an archival node back into the operational tier') and explicitly frames itself as `settle`'s mirror, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use condition ('settled knowledge that turned out to still be in flight') and names the mirror sibling `settle`, implying when each applies. It stops short of an explicit when-not or a direct 'use settle instead when…' statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unslotB
Free a role without touching the node that held it — its layer stays as it is.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | The role name. | |
| initiative | Yes | Initiative whose slots to act on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the whole burden. It does disclose a genuinely useful side-effect boundary — the holding node and its layer are preserved — but says nothing about reversibility, permissions, or what happens to the freed role's prior assignment.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the scope qualifier after the dash. Nothing is wasted, though the elliptical phrasing costs a little immediate clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating tool with no annotations and no output schema, the description covers what changes and what does not, which is the most important thing. It still omits reversibility, error conditions, and the return/confirmation behavior an agent would want before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: 'slot' (the role name) and 'initiative' (whose slots to act on) are both documented in the schema. The description adds no syntax, format, or scoping detail beyond those definitions, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('free a role') and, crucially, scopes it: the node and its layer are left untouched. That distinguishes it from the general 'slot'/'slots' family, though it never names a sibling directly and the domain terms 'role'/'node'/'layer' are left ungrounded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'without touching the node' clause implies when this is the right choice versus a variant that does modify the node or layer, but no alternative tool is named and no prerequisite (e.g. the role must currently be slotted) is stated. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
whyA
Why is this here? Reads the saved reasoning that leads to a node — the state → reasoning → decision trail, not an isolated record. Give it a chain to read its ordered steps, or any node to see the chain it belongs to (read directly when there is only one, listed for triage when there are several). Replaces the former chains + read_chain pair.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | A chain (read its steps) or any node (see the reasoning it belongs to). Name or UUIDv7 id. | |
| initiative | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the read nature (verb "Reads") and a genuine output behavior (direct read for a single chain, triage list for several), but omits permissions, side effects, pagination, and any explicit read-only declaration.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose after a short rhetorical hook, then covers usage and the replaced tools in a compact block. Every sentence carries information, though the opening "Why is this here?" is stylistic rather than substantive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description adequately covers the primary (required) `name` parameter's dual behavior and the read/triage semantics. The unexplained optional `initiative` parameter is the main residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `name` is documented in the schema and the description largely restates it (chain vs node), while the optional `initiative` parameter is undocumented in both places. The description adds the ordered-steps/triage nuance but does not compensate for the gap on `initiative`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource: "Reads the saved reasoning that leads to a node" and the state → reasoning → decision trail, explicitly distinguishing it from reading "an isolated record." It does not, however, distinguish itself from near-siblings like `chain` or `trace`, so an agent must still infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit operating modes: pass a chain to read its ordered steps, or any node to see the chain it belongs to, with single-vs-multiple handling described ("read directly when there is only one, listed for triage when there are several"). It stops short of naming which sibling to use instead when a caller wants something else.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
71 tool updates
v0.7.5- First observed
at - First observed
attach - First observed
awake - First observed
between - First observed
board - First observed
board_status - First observed
chain - First observed
cite - First observed
claim - First observed
close_review - First observed
cloud_initiatives - First observed
cloud_links - First observed
cloud_recall - First observed
clouds - First observed
config - First observed
confirm - First observed
delete_initiative - First observed
done - First observed
drill - First observed
episode - First observed
evidence - First observed
export - First observed
flag - First observed
forget - First observed
history - First observed
hygiene - First observed
ideas - First observed
import - First observed
inconclusive - First observed
initiatives - First observed
jot - First observed
layer - First observed
link - First observed
link_cloud - First observed
lint - First observed
merge_initiative - First observed
neighbours - First observed
outcomes - First observed
overview - First observed
path - First observed
pin - First observed
policy - First observed
pull - First observed
recall - First observed
recent - First observed
rechain - First observed
reflect - First observed
refute - First observed
rename_initiative - First observed
resolve - First observed
revise - First observed
reweight - First observed
search - First observed
set_status - First observed
settle - First observed
share - First observed
slot - First observed
slots - First observed
supersede - First observed
surface - First observed
sync_review - First observed
synthesise - First observed
tagged - First observed
task - First observed
trace - First observed
unlink - First observed
unpin - First observed
unsettle - First observed
unshare - First observed
unslot - First observed
why
TDQS
Scored across 71 tools
There is heavy thematic overlap in the read family (at vs recall vs search vs tagged; drill vs neighbours vs between; path vs chain vs why) and in the write/lifecycle family (jot vs episode; settle vs synthesise; link vs link_cloud; reflect vs lint), but the descriptions go to unusual lengths to draw explicit boundaries ('drill follows only derived_from and part_of; neighbours follows all twelve'). An agent can tell them apart with care, though the sheer number of near-parallel verbs invites occasional misselection.
Casing and separators are consistent (all lowercase snake_case), but predicate style is mixed: bare imperative verbs (link, settle, pin, search), verb_noun compounds (set_status, close_review, merge_initiative, cloud_recall), and bare nouns (board, slots, ideas, outcomes, history, config). No camelCase chaos, but no single predictable convention either.
71 tools is far beyond any comfortable surface, and the rubric flags 25+ as heavy and 50+ as extreme. The domain (bi-temporal graph, tiers, layers, chains, roles, reviews, initiatives, multi-cloud sync) genuinely justifies more than a typical server, so it is not a pure mismatch, but the count is still excessive for reliable tool selection.
Coverage is remarkably exhaustive for a knowledge substrate: capture (jot/episode/cite/claim/task/evidence), layered read (at/recall/search/tagged/drill/neighbours/between/path/trace/why/history/surface), graph editing (link/unlink/reweight/chain/rechain), lifecycle (settle/unsettle/supersede/revise/synthesise/forget), task board, reviews/verdicts, roles, initiative management, cloud sync, and maintenance (reflect/lint/hygiene) plus export/import. Almost no dead ends in the domain flow.
Maintenance
Related MCP Connectors
Shared project memory for AI coding agents: decisions, lessons, risks and tasks in one graph.
Graph-native persistent memory for AI agents — 33 MCP tools, zero-LLM writes.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Shared long-term memory for AI agents: save and recall context as a searchable knowledge graph.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA local memory engine for AI agents. Stores conversation episodes, consolidates knowledge through a neuroscience-inspired lifecycle, and builds a personal knowledge graph — all in a local SQLite database.17MIT
- AlicenseNot gradedqualityCmaintenancePersistent memory for AI agents with Ebbinghaus forgetting curve, semantic graph retrieval, and engineering state tracking. Local-first, DuckDB, zero cloud dependency.1MIT
- AlicenseAqualityAmaintenanceLocal-first memory engine for AI-agent teams: private/team/project ACL, associative recall, and federated sync across nodes. One SQLite file, no LLM required.126Apache 2.0
- AlicenseNot gradedqualityAmaintenanceLocal-first, single-file, knowledge-graph memory layer for AI agents.1MIT