hive
Hive is an MCP server that turns an Obsidian markdown vault into persistent, on-demand knowledge for AI coding assistants via read, search, write, and worker tools.
Browse & read vaults —
vault_listlists projects/files with glob filters;vault_queryloads context, tasks, roadmap, lessons, or any file by path.Search —
vault_searchdoes full-text, regex, ranked, recent-changes, and metadata-filtered search, plus lesson-ranking viarank_by.Session startup —
session_briefingbundles tasks, lessons, git activity, and health in one call.Write & edit —
vault_write(append/replace/create),vault_patch(surgical find-and-replace),vault_commit(batch flush), andvault_delete, all git-backed and idempotency-aware.Health & diagnostics —
vault_healthreports identity, drift checks, usage stats, and runtime metadata.Lessons —
capture_lessonwrites, batch-extracts, or looks up lessons whose confidence grows with each read.Model delegation —
delegate_taskoffloads work or summarizes large files to cheaper models;worker_statusprobes worker reachability and models.Semantic Q&A (optional) —
vault_askanswers natural-language questions with source-cited synthesis when embeddings/synthesis are configured.
Automatically performs Git commits when creating, appending, replacing, or patching files within the vault.
Manages vault content stored in Markdown, allowing tools to browse, read, and update files while preserving project documentation structure.
Connects an AI assistant to an Obsidian vault to enable on-demand querying, full-text searching, and management of project context, tasks, roadmaps, and lessons.
Integrates with Ollama as a worker to delegate tasks or summarize vault files using local AI models.
Utilizes YAML frontmatter within vault files for metadata filtering, structured data retrieval, and health metrics.
hive-vault
Your AI coding assistant forgets everything between sessions. Hive fixes that.
Hive is an MCP server that connects your AI assistant to an Obsidian vault. Instead of loading everything upfront, it queries only what's needed — on demand.
Metric | Without Hive | With Hive |
Context loaded per session | ~800 lines (static) | ~50 lines (on demand) |
Token cost for context | 100% every session | 6% average per query |
Knowledge retained between sessions | 0% | 100% (in vault) |
Measured on a real vault with 19 projects, 200+ files. See benchmarks.
Quick Start
Hive runs without a vault — vault tools return a friendly error until VAULT_PATH is set, so you can install first and configure later.
# Minimal — uses default vault path ~/Projects/knowledge
claude mcp add -s user hive -- uvx --upgrade hive-vault
# With a custom vault path
claude mcp add -s user hive -e VAULT_PATH=$HOME/path/to/vault -- uvx --upgrade hive-vault
# Gemini CLI
gemini mcp add -s user -e VAULT_PATH=$HOME/path/to/vault hive-vault uvx -- --upgrade hive-vaultDefault vault path:
~/Projects/knowledge. Override withVAULT_PATH(orHIVE_VAULT_PATH) as shown above.
For Codex CLI, GitHub Copilot, Cursor, Windsurf, and other clients, see Getting Started.
Then ask your assistant: "Use vault_list to see my vault"
Related MCP server: Obsidian MCP
Requirements
Hive degrades gracefully — every recommended or optional dependency reveals more capability without breaking the baseline.
Required
Python 3.12+ (works on 3.13).
A directory of markdown files. The vault structure used by
00_meta/10_projects/50_work/80_agentsis optional — without it, vault tools still operate but the scope routing is flat.
Recommended
gitinitialised inside the vault. Without it,vault_write/vault_patchstill write to disk; they just skip the per-write commit (andvault_commitreports the working tree as untracked).The Obsidian desktop app to author the vault by hand.
The obsidian-git plugin with auto-commit set to 5–10 minutes. Pair it with
vault_write(commit=False)/vault_patch(commit=False)to push the git workload off the synchronous tool path; see Recommended configuration below.
Optional
Ollama running
qwen2.5-coder:7b(or compatible) for local, freedelegate_task/capture_lessonworker calls.An OpenRouter API key (
OPENROUTER_API_KEY) as a free-tier and paid fallback worker.A backup git remote (e.g. private GitHub repo) so vault history survives a disk loss.
Recommended configuration
Per ADR-006 (commit policy), the recommended pairing for write-heavy flows is:
Install and enable the obsidian-git plugin in your vault.
Set its auto-commit interval to 5 or 10 minutes.
Call
vault_write(..., commit=False)andvault_patch(..., commit=False)for all bulk operations.Optionally call
vault_commit(message="...")at the end of a session to force a checkpoint sooner than the obsidian-git tick.
vault_health reports a ## external_committer block when it detects obsidian-git in the vault. The commit=False durability contract is explicit: files are persisted to disk regardless; only the commit is deferred. A crash before the next flush loses the commit, not the content.
When a tool call is cancelled mid-flight (slow worker, client timeout), the server may have already mutated the disk before the cancel ack reaches the wire. vault_health surfaces a ## ghost_responses counter and emits a mcp.ghost_response.suppressed_after_cancel_ack WARNING for each event — verify state via vault_query rather than retrying, since the ErrorData ack does not imply rollback (ADR-007).
Daemon mode (optional)
The default uvx hive-vault runs a fresh server per session. Daemon mode
instead runs one long-lived hive serve that owns the vault, with each
stdio-only session connecting through hive client. The adapter starts without
importing the Hive server or FastMCP, then relays JSON-RPC to the daemon's stable
loopback endpoint. If the daemon or credential is unavailable, it fails
explicitly rather than starting a competing in-process owner.
The default endpoint is a deterministic per-user port in 49152..65535, so an
ordinary daemon restart does not invalidate client configuration. Override it
with HIVE_DAEMON_PORT in both the daemon and every client environment when
the derived port conflicts with another local service. hive serve --port
changes only the daemon's bind port; it does not reconfigure hive client or
hive delegate. Hive fails closed rather than silently moving to a different
port. The owner-only bearer token persists across ordinary restarts in the daemon state
directory. daemon.port remains diagnostic migration metadata, not client
discovery state. See ADR-022.
Security hold: Until #456 is resolved, do not deploy daemon mode on untrusted multi-user hosts. The endpoint serves TLS with a per-user certificate that Hive's clients pin, so an impostor on the fixed port fails the handshake before any bearer is sent; the hold stays until the Linux and Windows evidence, including the cross-user check, is recorded.
uv tool install --upgrade hive-vault # >= 1.32.0
hive service install # supervise hive serve (systemd --user / Task Scheduler)Run hive client from your MCP host. A host that can verify the daemon's
certificate may connect directly to
https://127.0.0.1:<derived-or-overridden-port>/mcp (for Node-based hosts,
NODE_EXTRA_CA_CERTS pointing at daemon.crt in the state directory); never
disable verification, and note that http:// no longer works. Never print or
copy the token into logs or shell history.
To install a newer release, use the platform-specific command:
# Linux / macOS
uv tool upgrade hive-vault
# Windows (the version is optional; omitted selects the latest PyPI release)
hive self-upgrade [version]On Windows, self-upgrade builds the release beside the running files and atomically switches
the managed runtime, avoiding in-use-file conflicts. Open a new terminal after the first managed
upgrade so its PATH change is available. Once supervised, the daemon detects the new installed
version, exits 75, and the supervisor restarts it into the new code. See the
daemon mode guide and the
activation runbook.
Tools
Tool | What it does |
| Load project context, tasks, roadmap, lessons — or any file by path |
| Full-text search with metadata filters, regex, ranked results, recent changes, lesson-usage ranking ( |
| Browse projects and files with glob filtering |
| Server identity (version, vault path, backends), health metrics, drift detection, usage stats, opt-in runtime block |
| Create, append, or replace vault files. |
| Surgical find-and-replace. |
| Flush pending |
| Capture lessons inline / batch-extract from text / look up existing lessons by keyword ( |
| Tasks + lessons + git log + health in one call |
| Route tasks to cheaper models or summarize vault files |
| Budget, connectivity, available models |
Plus 5 resources and 4 prompts for guided workflows.
Lesson reinforcement
Every read of a lesson via vault_query, vault_search, or capture_lesson(find=…) increments a counter and grows that lesson's confidence asymptotically toward 1.0. Validated lessons rank higher than one-shot captures over time.
# Surface the top-ranked lessons matching a keyword
capture_lesson(project="hive", find="multi-process")
# Search lessons ranked by usage signal (not BM25)
vault_search(query="timeout", rank_by="reinforcements") # most-reinforced first
vault_search(query="timeout", rank_by="confidence") # highest decayed confidence
vault_search(query="timeout", rank_by="hybrid") # α=0.7 BM25 + 0.3 confidenceStorage: SQLite side-table at HIVE_LESSON_DB_PATH (default ~/.local/share/hive/lesson_reinforcement.db). WAL mode + busy_timeout make it cross-process safe.
Architecture
MCP Host (Claude Code, Gemini CLI, Codex CLI, Cursor, ...)
└── hive-vault (MCP server, stdio)
├── Vault Tools (7) ── Obsidian vault (Markdown + YAML frontmatter)
├── Session Tools (1) ── Adaptive context assembly
└── Worker Tools (2) ── Ollama (free) → OpenRouter free → paid ($1/mo cap) → rejectDocumentation
Full documentation at mlorentedev.github.io/hive:
Getting Started — install for all MCP clients
Configuration — all 19 environment variables
Vault Structure — how to organize your vault
Use Cases — real-world workflows
Architecture — module map and design decisions
Troubleshooting — common issues and fixes
Project-bound knowledge (docs-as-code) lives in docs/:
docs/adr/— Architecture Decision Recordsdocs/runbooks/— operational proceduresdocs/troubleshooting/— known issues and root-cause write-upsdocs/lessons.md— accumulated gotchas and post-mortems
Contributing
See CONTRIBUTING.md for setup and PR workflow.
git clone https://github.com/mlorentedev/hive.git && cd hive
make install # create venv + install deps
make check # lint + typecheck + test (478 tests, 90% coverage)License
Available Tools
13 toolscapture_lessonCapture LessonA
Capture lessons: inline / batch write, or lookup by keyword.
Inline mode (default): provide title, context, problem, solution.
Batch mode: provide text to extract lessons automatically via worker.
Lookup mode: provide find to surface top-ranked existing
lessons whose heading matches the keyword.
| Name | Required | Description | Default |
|---|---|---|---|
| find | No | Keyword to look up in existing lesson headings (lookup mode). | |
| tags | No | Optional tags (e.g. ["python", "testing"]). | |
| text | No | Raw text to extract lessons from (batch mode). | |
| title | No | Short descriptive title (inline mode). | |
| context | No | What you were doing (inline mode). | |
| problem | No | What went wrong or what decision was needed (inline mode). | |
| project | Yes | Project slug (directory under 10_projects/). | |
| rank_by | No | Lookup ranking — 'reinforcements' (default), 'confidence', or 'hybrid'. Ignored unless ``find`` is set. | reinforcements |
| solution | No | What fixed it or what was decided (inline mode). | |
| max_lessons | No | Maximum lessons to extract / surface. Default 5. | |
| min_confidence | No | Minimum confidence for batch extraction. Default 0.7. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only providing readOnlyHint=false, idempotentHint=false, and destructiveHint=false, the description adds useful behavioral context: batch mode 'extract[s] lessons automatically via worker' suggests asynchronous/delegated processing, and lookup mode 'surface[s] top-ranked existing lessons' clarifies read behavior. This goes beyond the sparse annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line summary followed by three bullet-like mode explanations. Every sentence earns its place, and the core purpose is front-loaded. No redundancy or filler exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 11 parameters and three modes, the description provides sufficient high-level context to choose and invoke the tool correctly, especially with an output schema present. It explains the main parameter roles and mode selection. A possible improvement is explicitly stating mode-selection precedence (e.g., if find is set, lookup wins), but the schema defaults and parameter descriptions make this inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaningful mode-based grouping: inline mode maps to title/context/problem/solution, batch mode to text, and lookup mode to find. It also reinforces that rank_by is ignored unless find is set. This is above baseline because it organizes parameters semantically rather than just repeating schema field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Capture lessons: inline / batch write, or lookup by keyword,' which names a specific verb and resource. It then enumerates three distinct modes (inline, batch, lookup), making it clearly distinguishable from sibling vault_* tools that handle generic vault content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear guidance on when to use each mode: inline for title/context/problem/solution, batch for text extraction via worker, and lookup for keyword search. However, it does not explicitly compare this tool to alternatives like vault_write or vault_query, so some competitive routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegate_taskDelegate TaskA
Offload work to a cheaper model or summarize vault files.
When project is provided, reads a vault file. Small files (≤50 lines) are returned directly. Large files are auto-delegated to a worker for summarization — falls back to raw content if workers are unavailable.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Relative path to a .md file. Overrides section. | |
| model | No | Concrete model id. Empty uses the configured worker model. The 4.0.0 removal retired 'auto', 'ollama', 'openrouter-free' and 'openrouter'; passing one is rejected rather than ignored. | |
| prompt | No | The task description or code to process. | |
| context | No | Optional system context for the model. | |
| project | No | Project slug for vault summarization mode. | |
| section | No | Shortcut name for summarization. Ignored if path is set. | context |
| timeout_s | No | Per-dispatch deadline in seconds. 0 uses the ambient tool timeout. A value ABOVE the ambient one raises the ceiling rather than being clamped by it — a deadline a 60s default can silently cap is not a deadline (HIVE-384 AC3). | |
| max_tokens | No | Maximum tokens in the response. | |
| structured | No | Return a JSON record instead of prose. Prose is the default so every existing caller's contract is unchanged; the dispatcher asks for JSON because it needs the status as a VALUE. Exception types do not survive the JSON-RPC boundary between the daemon and its clients, so "the pool refused" and "the worker answered badly" cannot be told apart by type on the far side — and a dispatcher that cannot tell them apart turns a rate limit into a silent retry against a different model. | |
| max_summary_lines | No | Target summary length for summarization. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations are neutral (readOnlyHint=false, idempotentHint=false, destructiveHint=false), so the description carries the burden of behavior disclosure. It reveals genuinely non-obvious traits: the ≤50 line direct-return threshold, automatic delegation of large files, and the fallback to raw content when workers are unavailable. It stops short of stating side effects (whether summaries persist back to the vault) or cost/latency implications of delegation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero filler. The one-line purpose statement is front-loaded, and the conditional flow is compressed into a tight, scannable second sentence. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter dual-mode tool with an output schema, the description captures the core decision logic (mode selection, size threshold, fallback) while the schema documents parameters and return values. The main gaps are the relationship between the two modes and whether summarization has persistent side effects, but the rich schema compensates for most of what the description omits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with exceptionally rich per-parameter descriptions (model-value rejection semantics, timeout ceiling behavior with HIVE-384 reference, structured-mode rationale). Per the rubric, high coverage sets a baseline of 3. The tool description adds connective flow logic (project triggers file reading; size threshold drives delegation) but no parameter-level detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: "Offload work to a cheaper model or summarize vault files." This clearly distinguishes it from the vault_* siblings and worker_status. However, "offload work" is slightly vague about what the work actually is, and the dual-purpose framing splits focus, so it just misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains mode selection implicitly: "When project is provided, reads a vault file" and the small/large file branching. It does not explicitly name alternatives or say when to prefer this tool over vault_ask or other siblings — an agent could not tell from this alone whether to pick delegate_task or vault_ask for a summarization request.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_briefingSession BriefingARead-onlyIdempotent
Call at the start of every new session to load project context.
Without a project, returns the available project list with a usage hint — discoverability parity with vault_health() and worker_status(). With a project, assembles active tasks, recent lessons, git activity, and project health into a single response (replaces 3-4 manual calls).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project slug (directory under 10_projects/). Empty = list available projects so the caller can pick one. This is the only parameter — there is no `days` argument (the briefing window is fixed). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true and the description's 'loads/returns/assembles' language is fully consistent with them, so there is no contradiction. The description adds real behavioral value beyond the annotations: the empty-project vs. project-set branches, the aggregation of four data sources, and the discoverability convention shared with vault_health() and worker_status(). It does not cover failure behavior for an invalid project slug, but that is minor for a safe, read-only call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences with the purpose front-loaded, followed by the two parameter-dependent branches and a cost comparison. Every sentence earns its place — including 'replaces 3-4 manual calls,' which conveys value, and the sibling mention, which conveys response-shape parity — with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-optional-parameter, read-only, idempotent tool with an output schema, the description covers when to call it, what each branch returns, and how it relates to alternatives. The only gaps are minor: no statement of behavior on an unknown project slug and no note on data freshness/caching for the aggregated git and lesson data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — the schema already explains the `project` slug format, the empty-string meaning, and even notes there is no `days` argument. The description's 'without a project / with a project' phrasing mirrors the schema rather than adding new semantic content, so it lands at the high-coverage baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete directive — 'Call at the start of every new session to load project context' — then details the two behavioral branches (project list vs. assembled briefing of tasks, lessons, git activity, health). It distinguishes itself from the vault_* and worker_status siblings by explicitly naming vault_health() and worker_status() and by describing aggregation semantics no sibling offers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition ('start of every new session') and frames the tool as replacing '3-4 manual calls,' which tells the agent when the aggregated call is preferable. It names vault_health() and worker_status() for the discovery-parity case, but it never states an explicit when-not-to-use condition — e.g., what to call when only one slice of context (git activity alone) is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_askVault AskARead-onlyIdempotent
Ask a natural-language question; get a source-cited synthesized answer (semantic retrieval / RAG) or relevant vault sections when no synthesis model is configured.
OPTIONAL — disabled by default. Requires the [semantic] extra plus an
embeddings backend (HIVE_EMBED_BASE_URL); until then it returns a
short how-to-enable message and never errors. Set HIVE_SYNTH_MODEL to
enable LLM synthesis on top of retrieval. For keyword / regex lookups
use vault_search instead.
| Name | Required | Description | Default |
|---|---|---|---|
| question | No | The natural-language question to answer. Use `question`, not `query` or `prompt`. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly and idempotent annotations, the description discloses meaningful behavior: it returns a how-to-enable message when not configured, never errors, and changes behavior when HIVE_SYNTH_MODEL is set. This gives the agent accurate expectations for unconfigured environments.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then efficiently covers fallback behavior, configuration prerequisites, and the sibling alternative. Every sentence serves a distinct purpose without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the tool has optional synthesis behavior, the description fully covers prerequisites, degraded mode, configuration, and when to use a sibling tool. An agent has everything needed to decide whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by emphasizing use of the `question` parameter over `query` or `prompt`, which helps agents avoid common naming mistakes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Ask') and resource ('vault') with a concrete output ('source-cited synthesized answer' or 'relevant vault sections'). It clearly distinguishes itself from sibling vault_search by framing this as natural-language / semantic retrieval versus keyword / regex lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states exactly when to use the tool (natural-language questions, semantic retrieval/RAG) and explicitly says to use vault_search instead for keyword/regex lookups. It also covers the disabled-by-default state, prerequisites, and configuration needed for synthesis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_commitVault CommitA
Stage everything in the vault and create one commit.
Companion to vault_write(commit=False) and
vault_patch(commit=False): callers that opt out of per-write
commits batch many writes and then flush with a single
vault_commit call.
Returns the new commit SHA on success, a clean-tree notice when there is nothing to commit, or a human-readable error.
| Name | Required | Description | Default |
|---|---|---|---|
| message | No | Commit message. Empty defaults to "vault: batch update". This is the only parameter — there is no `project` argument; the commit spans the whole vault working tree. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false, idempotentHint=false, destructiveHint=false, so the description carries most of the behavioral burden. It discloses the no-op case ('clean-tree notice when there is nothing to commit'), the commit-wide scope, and the possible return states. It does not detail failure causes or permission requirements, but the behavior is well covered for a simple commit tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded. The first sentence states the action, the second gives the companion context, and the third describes return behavior. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema, the description is complete: it explains what the tool does, when to use it instead of per-write commits, its whole-tree scope, and its return values. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the schema already explains the message default value and explicitly notes this is the only parameter, with no project argument. The main description adds no further parameter-level meaning beyond referring to the commit message, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Stage everything in the vault and create one commit.' It clearly differentiates itself from siblings by positioning it as the batched flush counterpart to vault_write(commit=False) and vault_patch(commit=False), and it states the output (commit SHA or clean-tree notice).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the companion sibling tools and the exact condition for using this tool: callers that opt out of per-write commits should batch writes and then 'flush with a single vault_commit call.' The schema also reinforces scope by stating there is no project argument and the commit spans the whole vault working tree.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_deleteVault DeleteA
Delete a single file from the vault (destructive; recoverable via git).
Removes one file and, by default, commits the deletion so it stays
recoverable from git history (git revert / git show). Files
only — directories are rejected. A non-existent path is an error,
unless idempotency_key is set (then a retry against an already-gone
file is a no-op success).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Relative path to the file within the project. | |
| commit | No | Must be True (the default). Unlike ``vault_write``, this tool has no deferred mode: it neither uses the commit queue (a delete and a recreate inside one tick would collapse to a single state) nor leaves the removal uncommitted, which is the indefinite deferral ADR-018 §4 removed. ``commit=False`` is rejected with an explanation rather than silently upgraded — see the ADR's 2026-08-09 amendment. | |
| project | Yes | Project slug or '_meta' for cross-project content. | |
| idempotency_key | No | Optional at-most-once token. If set, a retry with the same key is a no-op after the first delete (ADR-013), which also makes deleting an already-removed file succeed. Empty (default) disables idempotency. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, which is odd for a delete tool. The description compensates by explicitly stating the operation is destructive, recoverable via git, commits by default, rejects directories, and errors on non-existent paths unless idempotency_key is set. It also discloses the commit=False rejection behavior. This adds substantial behavioral context beyond the annotations, though the annotation contradiction is notable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear opening sentence, then a compact paragraph covering behavior and edge cases. It is dense but not bloated; every sentence adds value. The only minor issue is the heavy use of ADR references and parentheticals, which could be slightly streamlined, but overall it is efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (destructive operation, commit behavior, idempotency, edge cases) and the presence of an output schema, the description covers all essential behavioral aspects: what gets deleted, what happens to the commit, how errors are handled, and how retries behave. The output schema handles return values, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description adds meaningful context beyond the schema: it explains the commit parameter's relationship to vault_write and ADR-018, and clarifies the idempotency_key's retry semantics. This goes beyond baseline schema coverage, earning a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Delete a single file from the vault', and immediately flags the destructive nature and recoverability via git. It distinguishes itself from siblings by noting it handles files only and rejects directories, and by contrasting with vault_write's deferred mode. This is unambiguous and clearly scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool (delete a single file), what it does not do (directories rejected, no deferred mode), and how it differs from vault_write. It also explains the idempotency_key behavior for retries, giving an agent clear conditions for when a non-existent path is acceptable. This is strong usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_healthVault HealthARead-onlyIdempotent
Return vault health metrics, validation, and optional usage analytics.
Always emits the ## server identity block (version, python,
vault path, backend presence, started_at) at the top.
Without parameters, returns a health summary for all projects. When checks are specified, runs drift detection (frontmatter, stale, links). When include_usage is True, appends tool usage analytics. When include_runtime is True, appends runtime metadata (uptime, tools, budget).
| Name | Required | Description | Default |
|---|---|---|---|
| checks | No | Validation checks to run. Empty = health summary only. Options: frontmatter, stale, links. | |
| project | No | Project slug to validate. Empty = all projects. | |
| max_issues | No | Maximum validation issues to report. Default 50. | |
| usage_days | No | Usage look-back window in days. Default 30. | |
| include_usage | No | Append vault tool usage analytics. Default False. | |
| include_runtime | No | Append runtime metadata block. Default False. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnlyHint and idempotentHint, and the description adds beyond that: the guaranteed server identity block, the all-projects default, and the parameter-gated behavior for drift detection and analytics. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the one-line purpose is followed by the guaranteed output block and then a clean list of conditional behaviors. Every sentence carries necessary information without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the fully documented schema, the presence of an output schema, and read-only/idempotent annotations, the description covers what the tool returns, how parameters change the response, and what the default behavior is. Nothing essential is missing for safe, correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented with defaults and meaning in the input schema. The description reinforces the conditional relationships between parameters and behavior, but does not add substantial meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'Return vault health metrics, validation, and optional usage analytics.' It clearly distinguishes this tool from content-oriented siblings like vault_list, vault_query, and vault_ask by focusing on health, drift detection, and runtime/usage metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description spells out the main behavioral modes: no parameters means a health summary for all projects, checks trigger drift detection, and include_usage/include_runtime append specific blocks. It does not explicitly name alternative tools, but the conditional usage is clear enough for an agent to know when this tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_listVault ListBRead-onlyIdempotent
List vault projects, or files within a project.
When called without arguments, lists all available projects. When called with a project, lists files in that project directory.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Subdirectory within the project. Empty = project root. (Use `path`, not `subpath` — `subpath` is accepted as an alias.) | |
| pattern | No | Glob pattern to filter files (e.g. 'adr-*', '*.md'). | |
| project | No | Project slug. Empty = list all projects. | |
| subpath | No | Alias of `path` (#151). Prefer `path`. Note: there is no `scope` parameter here — `scope` lives on `vault_search`. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds the core behavior of listing projects/files but does not disclose additional details like sorting, pagination, or the effect of the pattern parameter (though the schema covers it). It is consistent with annotations and adds some context, but not rich behavioral depth beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero fluff. The core action is front-loaded, followed by two clear usage examples. Every sentence earns its place, making it efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (list projects/files) and has an output schema, so the return format is defined. Annotations cover side effects, and the description covers the primary invocation modes. It doesn't mention the subpath alias or pattern parameter, but those are fully documented in the schema. Overall, nothing critical is missing for a read-only list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description implicitly explains the project parameter (with/without project) and path usage via the two invocation modes, but it does not add syntax or format details beyond what the schema already provides. It adds marginal value by clarifying the default behavior, but not enough to exceed the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists vault projects or files within a project, with a specific verb and resource. It distinguishes between the two modes (no args vs. with a project). However, it does not explicitly differentiate from sibling tools like vault_search or vault_query, which might also list/filter content, so it's not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to invoke the tool with or without a project, which is internal usage. It provides no guidance on when to choose this tool over alternatives such as vault_search (for content search) or vault_query (for structured queries). The absence of any reference to siblings leaves the agent to infer the appropriate tool, which is a significant gap given the overlapping toolset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_patchVault PatchA
Surgical find-and-replace in a vault file with auto git commit.
Supports single or multi-replacement. For a single replacement, provide
find and replace. For multiple replacements, provide patches
— a list of {"find": "...", "replace": "..."} dicts applied in
sequence. Do not mix both modes.
Each find value must appear exactly once in the file (after prior
patches in the list have been applied). If any patch fails validation,
no changes are written.
Uses 3-pass cascading match: exact → body-only → whitespace-normalized.
| Name | Required | Description | Default |
|---|---|---|---|
| find | No | Exact text to find (single mode). Empty = not set. (Use `find`/`replace`, NOT `old_string`/`new_string` — those are accepted as aliases.) | |
| path | Yes | Relative path to the file within the project. | |
| commit | No | If True, commit synchronously before returning. Defaults to False, which queues the path for the reconciler. See ``vault_write`` docstring for the durability contract. | |
| patches | No | List of {"find", "replace"} dicts (multi mode). | |
| project | Yes | Project slug or '_meta' for cross-project content. | |
| replace | No | Replacement text (single mode). Empty = not set. | |
| new_string | No | Alias of `replace` (#151). Prefer `replace`. | |
| old_string | No | Alias of `find` (#151). Prefer `find`. | |
| idempotency_key | No | Optional at-most-once token. If set, a retry with the same key is a no-op after the first apply (ADR-013). Empty (default) disables idempotency. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only carrying false boolean hints, the description adds substantial behavioral detail: 3-pass cascading match, sequential patch application, all-or-nothing validation, and the auto-commit side effect. The commit parameter description further clarifies the default queue-for-reconciler path, so there is no misleading gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short paragraphs, each earning its place: summary, mode selection, validation constraint, and matching algorithm. It front-loads the core purpose and avoids restating schema definitions or adding filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with an output schema, the description covers all invocation-critical behavior: mode selection, uniqueness constraint, atomic failure, matching semantics, and commit behavior (with a pointer to the durability contract). An agent has enough to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description goes beyond per-field docs by explaining how find/replace relate to patches, that patches apply in sequence, that each find must be unique, and that modes cannot be mixed. This adds real cross-parameter meaning rather than repeating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Surgical find-and-replace in a vault file with auto git commit,' which names a specific operation, resource, and side effect. This clearly differentiates it from the sibling vault_write/vault_delete tools by emphasizing targeted in-place edits, and the mode details reinforce the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage rules: single mode via find/replace, multi mode via patches, 'Do not mix both modes,' and the validation requirement that each find must appear exactly once. It lacks explicit when-not guidance naming alternatives (e.g., vault_write for whole-file overwrites), so it doesn't earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_queryVault QueryBRead-onlyIdempotent
Read content from a vault project — use instead of direct filesystem access.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Relative path to a specific .md file within the project. Overrides section. (Use `path` for a file, not `identifier` — `identifier` is accepted as an alias.) | |
| project | Yes | Project slug (directory under 10_projects/), or '_meta' for 00_meta/. | |
| section | No | Shortcut name (context, tasks, roadmap, lessons). Ignored if path is set. | context |
| max_lines | No | Maximum lines to return. 0 = unlimited. | |
| identifier | No | Alias of `path` (#151). Prefer `path`; for a section shortcut use `section` instead. | |
| include_metadata | No | Prepend a structured metadata line from YAML frontmatter. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, which cover the safety profile of this read-only operation. The description adds no behavioral detail beyond 'read content' — nothing about truncation, metadata prepending, or section-versus-path behavior — though those are documented in the schema. There is no contradiction between the description and the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler, and the 'instead of direct filesystem access' clause adds meaningful guidance. It is concise and scannable. However, it is minimal enough that it misses a natural opportunity to point to alternative vault tools.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich input schema, output schema, and read-only annotations, an agent can determine how to invoke the tool correctly without extra description. The description alone, though, does not situate vault_query among the twelve sibling tools, particularly vault_ask and vault_search, so selection context is incomplete. The overall package is adequate for invocation but not fully complete for optimal tool choice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter already has a detailed description covering path, project, section, max_lines, identifier, and include_metadata. The tool description itself contributes nothing about parameters. Baseline 3 is appropriate because the schema carries the full semantic burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific action ('Read content') and resource ('vault project'), so the core function is clear. It does not, however, differentiate vault_query from siblings like vault_list, vault_search, or vault_ask, all of which also involve retrieving information from the vault. It avoids tautology and is more informative than the title alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'use instead of direct filesystem access' gives a useful policy context: when an agent would otherwise read files directly, it should use this tool instead. It does not say when to prefer vault_query over vault_list, vault_search, or vault_ask, nor does it provide any exclusions. This is a clear but incomplete usage signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_searchVault SearchBRead-onlyIdempotent
Search the vault: full-text, ranked, or recent changes.
Default mode: flat full-text search across all vault files. Ranked mode (ranked=True): results scored by relevance. Recent mode (since_days>0): files changed in the last N days. rank_by mode (rank_by != 'bm25'): lessons-only, ranked by usage.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Alias of `max_results` (#202). Prefer `max_results`. When both are given the tighter (smaller) cap wins; 0 = unset. | |
| query | No | Text to search for (case-insensitive). | |
| regex | No | Alias of `use_regex` (#151). Prefer `use_regex`. To narrow by location use `scope` / `project`, not `path_filter` / `path_prefix`. | |
| scope | No | Restrict search to a scope (e.g. 'work', 'projects'). Empty = all. | |
| ranked | No | Score results by relevance. Default False. | |
| project | No | Filter to this project (recent mode only). | |
| rank_by | No | Lesson ranking ('bm25' default keeps current behaviour; 'reinforcements', 'confidence', 'hybrid' filter to 90-lessons.md only and rank by usage signal). | bm25 |
| max_lines | No | Maximum output lines. Default 500. | |
| use_regex | No | Treat query as regex. Default False. (Use `use_regex`, not `regex` — `regex` is accepted as an alias.) | |
| since_days | No | Show recent changes (0 = disabled). Default 0. | |
| tag_filter | No | Only files that have this frontmatter tag. | |
| max_results | No | Max result files. Default 10. Caps the file count in all modes (flat, ranked, recent); in flat/recent the cap is by path order (alphabetical) — use ranked=True for relevance order. | |
| type_filter | No | Only files whose frontmatter type matches. | |
| status_filter | No | Only files whose frontmatter status matches. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, covering the safety profile. The description adds mode-specific behavior (e.g., rank_by != 'bm25' limits to lessons-only) but does not disclose other behavioral traits like performance, pagination details, or the interaction of aliases. Given the annotations, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the core purpose, and uses a compact bullet-like structure to list modes. Every line adds value, and there is no fluff. It could be slightly more structured, but it is efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and 100% schema coverage, the description does not need to explain return values or every parameter. It covers the primary modes and their key parameters, which is sufficient for an agent to understand how to invoke the tool correctly for common use cases. It does not mention all filters (tag, status, type), but those are self-explanatory in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage, the baseline is 3. The description adds meaningful semantics by explaining the effects of ranked, since_days, and rank_by parameters, including that rank_by != 'bm25' restricts to lessons-only and ranks by usage. This goes beyond the schema's property descriptions and clarifies intended usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Search the vault' and enumerates three distinct modes (flat, ranked, recent) plus a rank_by variant, giving a specific verb and resource with mode distinctions. However, it does not explicitly differentiate from sibling tools like vault_query, so it doesn't fully separate itself from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage guidance for its own modes (when to use ranked, recent, rank_by) but never mentions when to use this tool versus other vault-related tools like vault_query or vault_list. There is no explicit 'use this for X, use that for Y' guidance, so the tool-level selection context is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vault_writeVault WriteA
Write to the vault: append, replace a section, or create a new file.
Modes:
append/replace: Update a project section. Requires section.
create: Create a new file with auto-generated frontmatter. Requires path; doc_type defaults to "note". Inferred automatically when you pass a path with no section, so operation may be omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Relative path for the file. Setting this with no section creates the file (create mode). | |
| commit | No | If True, commit synchronously before returning — the escape hatch for a caller that needs the commit to exist by the time the call ends. Defaults to False, which queues the path for the reconciler to commit on its next tick (a few seconds). Durability contract: the file is persisted to disk regardless; only the *commit* is deferred, so a crash before the next flush loses the commit, not the content. | |
| content | Yes | Markdown content to write (body only for create mode). | |
| project | Yes | Project slug or '_meta' for cross-project content. | |
| section | No | Section shortcut (context, tasks, roadmap, lessons). For append/replace. | |
| doc_type | No | Document type for frontmatter (create mode). Optional; defaults to "note". | |
| operation | No | 'append', 'replace', or 'create'. Default 'append'. 'create' is inferred when path is set and section is empty. | append |
| idempotency_key | No | Optional at-most-once token. If set, a retry with the same key is a no-op (safe for transparent retries after a daemon restart cuts an in-flight write — ADR-013). Empty (default) disables idempotency. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the annotations: auto-generated frontmatter, mode-specific requirements, and operation inference. However, it does not disclose the deferred-commit behavior (only present in the schema) or explicitly note that 'replace' overwrites existing section content. No contradiction with the annotations was found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then uses a compact bulleted mode list that maps cleanly to parameters. Every sentence earns its place; there is no filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters and a rich output schema, the description covers the key mode-selection logic well. However, it omits the important deferred-commit/durability behavior and does not contrast the tool with siblings like vault_patch, which would help an agent choose the correct tool at a glance. The schema fills many gaps, but the description itself is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by clarifying that section is behaviorally required for append/replace and path for create, even though these are not schema-level required fields, and by explaining the operation inference rule. This exceeds what the schema alone provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Write to the vault') and enumerates the three operation modes (append, replace, create), making the core purpose clear. It does not explicitly distinguish itself from siblings like vault_patch, so it misses the top-tier differentiator required for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context within the tool: append/replace requires a section, create requires a path, and operation is inferred when path is set with no section. It does not state when to prefer an alternative sibling tool, but the internal mode guidance is strong and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
worker_statusWorker StatusARead-onlyIdempotent
Show worker health: configuration, reachability, model, and usage.
HIVE-384 reshaped this tool, and the reshape is the point rather than a side effect. The old output led with a dollar budget and reported two providers by configuration: it said "Ollama: offline / OpenRouter: no API key" for an unknown length of time while every caller treated the worker as a working capability. A status surface that cannot distinguish "configured" from "answers" is how a dead backend stays invisible.
So reachability is probed, not inferred, and reported separately from configuration. The dollar figures are gone: on a flat subscription they would read zero forever, and a gauge that always says the same thing looks like a working gauge.
| Name | Required | Description | Default |
|---|---|---|---|
| include_models | No | Probe the provider for its model list. Default True. Set False to report configuration without a network call. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and idempotent, and the description adds material behavioral context: reachability is probed, not inferred; configuration and reachability are reported separately; and dollar-budget fields were removed as misleading. This tells an agent not to treat 'configured' as 'answering' and explains why no budget gauge appears. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is crisp and front-loaded. However, most of the text is a historical rationale about HIVE-384 and the old dollar budget; it supports understanding but is longer than needed for a one-parameter status tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent tool with one optional parameter and an output schema, the description is sufficient: it states what is shown, clarifies the key semantic guarantee of probed reachability, and the schema covers parameter details and return shape. An agent has what it needs to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter include_models is fully documented in the input schema (100% coverage), including its default and effect of avoiding a network call when false. The description itself adds no parameter-specific meaning, so the schema carries the burden. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb ('Show') and resource ('worker health') and enumerates the four aspects reported: configuration, reachability, model, and usage. The 'worker' framing clearly separates it from the vault_* siblings, including vault_health, so there is no ambiguity about what it targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended use context clear — checking worker health/status — but never states when to prefer this tool over an alternative or when not to use it. There is no explicit exclusion or sibling routing, and guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v4.1.0- Changed
delegate_task2 fields changed- added
Input schema / properties / structuredAdded value: +{ + "default": false, + "description": "Return a JSON record instead of prose. Prose is the\ndefault so every existing caller's contract is unchanged; the\ndispatcher asks for JSON because it needs the status as a\nVALUE. Exception types do not survive the JSON-RPC boundary\nbetween the daemon and its clients, so \"the pool refused\" and\n\"the worker answered badly\" cannot be told apart by type on the\nfar side — and a dispatcher that cannot tell them apart turns a\nrate limit into a silent retry against a different model.", + "type": "boolean" +} - added
Input schema / properties / timeout_sAdded value: +{ + "default": 0, + "description": "Per-dispatch deadline in seconds. 0 uses the ambient\ntool timeout. A value ABOVE the ambient one raises the ceiling\nrather than being clamped by it — a deadline a 60s default can\nsilently cap is not a deadline (HIVE-384 AC3).", + "type": "number" +}
2 tool updates
v4.0.0- Changed
delegate_task3 fields changed- removed
Input schema / properties / max_cost_per_requestRemoved value: -{ - "default": 0, - "description": "Max USD. 0 = free models only.", - "type": "number" -} - changed
Input schema / properties / model / defaultPrevious value: -"auto"New value: +"" - changed
Input schema / properties / model / descriptionPrevious value: -"'auto', 'ollama', 'openrouter-free', 'openrouter' (paid), or model ID."New value: +"Concrete model id. Empty uses the configured worker model.\nThe 4.0.0 removal retired 'auto', 'ollama', 'openrouter-free'\nand 'openrouter'; passing one is rejected rather than ignored."
- Changed
worker_status1 field changed- changed
Input schema / properties / include_models / descriptionPrevious value: -"Include available model list from all providers. Default True."New value: +"Probe the provider for its model list. Default True.\nSet False to report configuration without a network call."
3 tool updates
v3.0.0- Changed
vault_delete1 field changed- changed
Input schema / properties / commit / descriptionPrevious value: -"If True (default), stage + commit the deletion. If False,\nunlink on disk but leave the removal staged for a later\n``vault_commit`` (or obsidian-git). Same durability contract as\n``vault_write``."New value: +"Must be True (the default). Unlike ``vault_write``, this\ntool has no deferred mode: it neither uses the commit queue\n(a delete and a recreate inside one tick would collapse to a\nsingle state) nor leaves the removal uncommitted, which is\nthe indefinite deferral ADR-018 §4 removed. ``commit=False``\nis rejected with an explanation rather than silently\nupgraded — see the ADR's 2026-08-09 amendment."
- Changed
vault_patch2 fields changed- changed
Input schema / properties / commit / defaultPrevious value: -trueNew value: +false - changed
Input schema / properties / commit / descriptionPrevious value: -"If True (default), auto-commit. If False, write to disk\nwithout committing — useful for batching many patches into\none ``vault_commit`` flush. See ``vault_write`` docstring for\nthe durability contract."New value: +"If True, commit synchronously before returning.\nDefaults to False, which queues the path for the reconciler.\nSee ``vault_write`` docstring for the durability contract."
- Changed
vault_write2 fields changed- changed
Input schema / properties / commit / defaultPrevious value: -trueNew value: +false - changed
Input schema / properties / commit / descriptionPrevious value: -"If True (default), auto-commit to git. If False, write to\ndisk but leave the file dirty so the caller can batch many\nwrites into one commit via the ``vault_commit`` tool, or let\nobsidian-git's auto-commit pick it up. Durability contract:\nfiles are persisted to disk regardless; only the *commit* is\ndeferred. A crash before the next flush loses the commit, not\nthe file content."New value: +"If True, commit synchronously before returning — the\nescape hatch for a caller that needs the commit to exist by\nthe time the call ends. Defaults to False, which queues the\npath for the reconciler to commit on its next tick (a few\nseconds). Durability contract: the file is persisted to disk\nregardless; only the *commit* is deferred, so a crash before\nthe next flush loses the commit, not the content."
4 tool updates
v1.41.1- Added
vault_ask - Added
vault_delete - Changed
vault_search2 fields changed- added
Input schema / properties / limitAdded value: +{ + "default": 0, + "description": "Alias of `max_results` (#202). Prefer `max_results`. When\nboth are given the tighter (smaller) cap wins; 0 = unset.", + "type": "integer" +} - changed
Input schema / properties / max_results / descriptionPrevious value: -"Max files when ranked. Default 10."New value: +"Max result files. Default 10. Caps the file count in\nall modes (flat, ranked, recent); in flat/recent the cap is by\npath order (alphabetical) — use ranked=True for relevance order."
- Changed
vault_write3 fields changed- changed
Input schema / properties / doc_type / descriptionPrevious value: -"Document type for frontmatter. For create mode."New value: +"Document type for frontmatter (create mode). Optional;\ndefaults to \"note\"." - changed
Input schema / properties / operation / descriptionPrevious value: -"'append', 'replace', or 'create'. Default 'append'."New value: +"'append', 'replace', or 'create'. Default 'append'.\n'create' is inferred when path is set and section is empty." - changed
Input schema / properties / path / descriptionPrevious value: -"Relative path for new file. For create mode."New value: +"Relative path for the file. Setting this with no section\ncreates the file (create mode)."
2 tool updates
v1.32.2- Changed
vault_patch1 field changed- added
Input schema / properties / idempotency_keyAdded value: +{ + "default": "", + "description": "Optional at-most-once token. If set, a retry with\nthe same key is a no-op after the first apply (ADR-013). Empty\n(default) disables idempotency.", + "type": "string" +}
- Changed
vault_write1 field changed- added
Input schema / properties / idempotency_keyAdded value: +{ + "default": "", + "description": "Optional at-most-once token. If set, a retry with\nthe same key is a no-op (safe for transparent retries after a\ndaemon restart cuts an in-flight write — ADR-013). Empty\n(default) disables idempotency.", + "type": "string" +}
TDQS
Scored across 13 tools
Most vault_* tools have clearly distinct purposes, but vault_query, vault_search, and vault_ask are all read-oriented and could occasionally be confused by an agent. The descriptions do enough to separate exact-content reads from full-text search and RAG, so this is only a minor concern.
The dominant vault_* verb pattern is consistent and readable, covering commit, list, query, search, write, patch, delete, health, and ask. The auxiliary tools (session_briefing, worker_status, capture_lesson, delegate_task) break the prefix pattern, though they are still reasonably named and not chaotic.
Thirteen tools is well within the ideal 3–15 range and each tool maps to a meaningful capability: vault lifecycle, health, search, lessons, session context, worker status, and delegation. Nothing feels redundant enough to cut, and the scope justifies the count.
The vault domain is well covered: list, read, write, patch, search, delete, commit, health, and ask provide a strong lifecycle. Minor gaps exist—no explicit project creation, lesson-specific update/delete, or git history surface—but most can be worked around via existing tools.
Maintenance
Related MCP Connectors
Persistent memory and cross-session learning for AI coding assistants (hosted remote MCP).
shared AI-context layer for teams — persistent memory your agents search and update over MCP
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
Cross-tool persistent memory and context for AI assistants over MCP.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceThis is a connector to allow Claude Desktop (or any MCP client) to read and search any directory containing Markdown notes (such as an Obsidian vault).1,265 npm1,352AGPL 3.0
- AlicenseAqualityFmaintenanceThis project implements a Model Context Protocol (MCP) server for connecting AI models with Obsidian knowledge bases. Through this server, AI models can directly access and manipulate Obsidian notes, including reading, creating, updating, and deleting notes, as well as managing folder structures.1146 npm310MIT
- AlicenseNot gradedqualityDmaintenanceContext Portal (ConPort): A memory bank MCP server building a project-specific knowledge graph to supercharge AI assistants. Enables powerful Retrieval Augmented Generation (RAG) for context-aware development in your IDE.85 PyPI768Apache 2.0
- AlicenseAqualityFmaintenanceA Model Context Protocol (MCP) server that provides AI assistants with secure access to Obsidian vaults. Enables reading, writing, searching, and managing notes without requiring Obsidian to be running.503,893 npmApache 2.0