Skip to main content
Glama
pvliesdonk

markdown-vault-mcp

by pvliesdonk

Reindex Vault

reindex
Idempotent

Update the index when files are modified outside the server, such as by editors or sync tools. Incremental hash-based detection re-parses only changed files, so unchanged files are never reprocessed.

Instructions

Run an incremental reindex on the writer thread.

Only needed when files are modified outside this server — for example, by a text editor, a sync tool, or another process writing directly to the vault directory. Do NOT call this after using 'write', 'edit', 'delete', or 'rename' — those tools queue index updates automatically.

Change detection is hash-based, so an unchanged file is never re-parsed. Use force=True to drop the index and re-parse every file regardless of hashes — the repair for index content that no longer matches what the current server would extract. A version upgrade that changes extraction does this by itself on the next start (#1124), so force=True is a manual escape hatch, not routine maintenance. When semantic search is configured, follow a force=True run with 'build_embeddings' (without force) so the vector index converges to the rebuilt chunk set; an ordinary reindex re-embeds as it goes.

To rebuild all embeddings from scratch (e.g. after changing the embedding model), use 'build_embeddings' with force=True.

A fast reindex (the common case — work scales with the drift, not the vault) returns its result inline. A reindex still running at the server's soft deadline continues in the background and returns {"status": "working", "job_id": ...} immediately — fetch the outcome with get_job_result. get_index_status remains the observability view of the index (it also covers boot-time builds and file-watcher reindexes no client call initiated).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
forceNoWhen True, drop every indexed document and re-parse the whole vault instead of applying the hash-detected delta. The index is not queryable while the rebuild runs, and the cost scales with the vault rather than the drift, so prefer the default.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv4.0.0
    • addedInput schema / properties / force
      Added value: +{
      +  "default": false,
      +  "description": "When True, drop every indexed document and re-parse the\nwhole vault instead of applying the hash-detected delta.\nThe index is not queryable while the rebuild runs, and the\ncost scales with the vault rather than the drift, so prefer\nthe default.",
      +  "type": "boolean"
      +}
  2. First observedv3.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly=false, idempotent=true, destructive=false), the description reveals async continuation: a reindex past the server deadline returns immediately with job_id, and get_index_status covers non-client-initiated builds. It also details hash-based detection, force behavior (index not queryable during rebuild), and interaction with embedding builds. This is rich disclosure out of annotation tags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the function and its trigger, then gives exclusions, force semantics, embedding interactions, and async result handling. Some sentences (e.g., '#1124' reference) are minor extraneous color, but each paragraph advances a decision an agent must make. It's thorough without being bloated for the tool's real complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 1 param, the description references the observability tool get_index_status and attr the async job fetch tool get_job_result; it also pre-charts the build_embeddings path for force=True. The common-case and rare-case flows are both covered, leaving nothing an agent needs to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents force at 100% coverage, but the description gives extra context on when to force (repair for content mismatch, version upgrade auto, manual escape hatch), contrasts it with routine reindexing, and mandates the follow-up build_embeddings. This meaningfully lifts the parameter beyond schema bare definition, justifying a step above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-resource pair ('Run an incremental reindex on the writer thread') and immediately scopes it to files modified outside the server, distinguishing it from write/edit/delete/rename (which queue updates) and from build_embeddings (which rebuilds the vector index). The purpose is unique and not generic.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains when to use: only when files were modified by an external process (text editor, sync tool, direct vault writes). It also says when NOT to use it — never after write/edit/delete/rename — and guides force=True usage, routing to build_embeddings for a full semantic rebuild. This is full when/when-not/alternatives coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.