Skip to main content
Glama

Build the semantic search index

zotero_index
Destructive

Build, update, or repair the local hybrid-search index so Zotero semantic search stays current. Update applies deltas; build resumes or rebuilds; status monitors progress.

Instructions

Manage the local hybrid-search index used by zotero_semantic_search. Every job runs in the background on the server, so this tool returns immediately and never blocks on large libraries. THREE write actions, and picking the right one matters: action: "update" is the cheap one and should be the default for a library that is already indexed; action: "build" and action: "refresh" both rebuild the WHOLE index, which on a large library means many minutes and, with an API embedding provider, real spend (they differ in one thing: build resumes an interrupted build, refresh always starts over). action: "build"/"refresh" pages the library's top-level items (100-at-a-time, stopping at the server's item cap, ZOTEUS_INDEX_MAX_ITEMS, default 5000, or at a smaller limit if one is given), indexes their text (title, abstract, creators, tags) for BM25 keyword search and, if an embedding provider is configured, for vector search, persisting partial progress atomically as it goes; use it for the first build, after changing the embedding model, or to widen a previously capped build. It is ALSO the repair: if the index cannot be read at all, only action:"build" clears it, by deleting the unreadable file and opening a fresh one before rebuilding (nothing repairs it at startup or inside a query). action: "update" instead fetches only the items changed since the version the index recorded (Zotero's ?since=), re-chunks and re-embeds just those, and removes items the library no longer holds (diffed from a cheap keys-only ?format=versions census, since the deletion log is cloud-only); untouched items are never re-embedded, so adding a handful of items costs seconds instead of a full rebuild. Update falls back to a full rebuild by itself, and says so in updateNotice, when a delta would be wrong: no version stamp recorded yet, the library is now served by a different Zotero API (the desktop app and the cloud number their versions independently), or the embedding model changed. An update ALSO asks Zotero's full-text index what it has extracted since the build (that is a separate version sequence from item versions, so a PDF Zotero extracted when it was first opened changes no item version and appears in no delta) and indexes the new body text for items nothing else touched; on a library where nothing was extracted, that costs one request. A build or update interrupted by action:"stop", a crash or a restart leaves a checkpoint, and action: "build" RESUMES from it: the items already committed stay searchable and are never re-fetched or re-embedded, and only work since the last save is redone (resumedFrom on the status reports how many were inherited). action: "refresh" is the one that always starts over. A build also indexes the reader's OWN words by default: every child note, and every PDF annotation (its highlighted passage and its comment), as extra passages carrying the parent item's key — so zotero_annotate writes text that search can then find, an item with forty annotations still takes one result slot, and a hit whose snippet came from one is marked source:"note" or source:"annotation". That corpus is one paged crawl of hand-written text, orders of magnitude smaller than attachment bodies; turn it off with own_words:false or ZOTEUS_INDEX_OWN_WORDS=false. An action:"update" keeps it current for the cost of one request when nothing was written: notes and annotations are ordinary items carrying ordinary versions, so an edit, an addition and a deletion are all found by comparing the library's note/annotation keys against the ones the index holds — which is also how an index built before this existed fills its gap, once, on its first update. Set fulltext:true to ALSO index the body text Zotero extracted from each item's attachments, which is what makes semantic search match a claim buried in a PDF rather than only its title and abstract; it is off by default because it multiplies build time and index size (default cap: 40000 characters per item, tunable with fulltext_max_chars), and only attachments Zotero has already extracted are available. That pass used to be refused inside Claude Desktop, where a build that reached it killed the server process partway through with no error at all (#37); the cause was the on-device embedding model asking Electron's allocator for a block it will not serve, so the server now embeds fewer passages per call there and the build runs to completion. It is somewhat slower inside the app than in a terminal and produces exactly the same index, so a user who wants the fastest possible first build can still run one headlessly against the same ZOTEUS_DATA_DIR and let Desktop read the result. A build runs in TWO passes and reports which one it is on as phase: every item's metadata is indexed first, across the whole library, and only then are attachment bodies crawled (fulltextItemsScanned of fulltextItemsTotal). So the library is fully searchable on titles, abstracts, creators and tags long before a full-text crawl that can run for hours finishes — tell the user they can search already rather than asking them to wait for state:"done". Start a job, then POLL action: "status" every few seconds until state is "done" (or "error"); calling build or update again while one is running just returns current progress. action: "status" reports state (idle|building|done|error), operation (build|update), phase (metadata|fulltext), fetch/embed progress, itemsRemoved, index size, the active embedder, libraryVersion/libraryBackend (the version stamp an update diffs from), fulltextVersion (how far into Zotero's separate full-text sequence the index has read), fulltextPartial (present when the index's body text was gathered over an attachment map that never reached the end of the library, which is usually why that cursor is 0, though a delta can damage coverage an earlier pass had already earned a cursor for: an item holding body passages may still be missing an attachment's text, and the next update asked for full text re-reads every item Zotero's full-text census names, once, which on a large library costs a whole body crawl), resumedFrom (items inherited when a build resumed an interrupted one), itemsTotal/itemsAvailable (which differ, with a warning, when the cap stopped the crawl short of the library), ownWordsItems/ownWordsPassages (the notes and annotations indexed, with ownWordsReason if they could not be read), and (when full text was requested) fulltextItems/fulltextPassages plus fulltextReason if it produced nothing, or if an update could not read part of the body text and therefore withheld its version stamp (or its full-text cursor) so the next update retries. It also reports localApiDegradedAt when the job saturated Zotero's local API and the whole session fell back to the Zotero Web API: that fallback works, so nothing errors, but the Web API is slower and rate-limited and the rest of the build takes far longer than its start suggested, so tell the user rather than letting them watch an unexplained slowdown (the crawl also backs off to one attachment at a time by itself, to let the app recover). It reports where the index is stored (storage: sqlite or memory, set by ZOTEUS_INDEX_BACKEND), storageNotice when opening that store imported or refused an older JSON index, persistError when the index could not be written to disk at all, and how the last semantic query ranked vectors (vectorScan: "codes" for the two-stage path, "exact" for a full scan of every vector, with vectorScanNotice when that needs explaining). When the embedding provider is a paid API with a tokens-per-minute limit (ZOTEUS_EMBEDDINGS=openai or gemini), status also reports embedRate: the batch size, the pause between requests, the estimated tokens per request and the tokens per minute the build is actually sustaining, plus passagesWithoutVectors when the index holds passages nothing has embedded yet. A build whose embedder was rate-limited to a standstill keeps every passage it indexed and stays RESUMABLE: tell the user to run action:"build" again, which embeds only the passages that have no vector and re-fetches nothing, and NOT action:"refresh", which starts the whole crawl over and pays for every vector a second time. A rate-limited request already backs off and retries by itself; if a build reports it is riding the provider's tokens-per-minute limit, the fix is ZOTEUS_EMBED_BATCH_DELAY_MS (with ZOTEUS_EMBED_BATCH_SIZE), not a smaller library. action: "stop" cancels a running job (partial data is kept and stays searchable; a stopped update leaves the version stamp untouched so the next one repeats the delta, and a stopped build leaves a checkpoint the next action:"build" resumes from). stop is a one-shot cancel: the next action:"build" picks the checkpoint straight back up. action: "pause" is the durable form: it stops a running job the same way AND persists a hold that survives restarts, so build, refresh, update and zotero_semantic_search's automatic first build all refuse until action: "resume" clears it (queries keep working on what is indexed). resume clears the hold and starts nothing by itself, so follow it with build to continue a checkpoint or update for a delta; status reports paused. A partially built index is always usable for keyword search. Local embeddings are CPU-bound (see ZOTEUS_EMBEDDINGS), so large builds take a while: poll status rather than retrying build. ONE INDEX FILE HOLDS ONE LIBRARY, and library_type/library_id therefore pick the FILE, not just the crawl: naming a group builds, updates and reports THAT group's own index, kept beside the personal library's, and the two never mix (passage ids carry no library and Zotero item keys repeat across libraries, so one store holding two would alias them). Omitting them means the library this server is configured for. action: "libraries" lists every library that has an index in this data directory with its passage/item/vector counts, its stamp and where its file is; it starts nothing and creates nothing, and neither does action:"status" for a library that has no index yet (it says so instead). Only a bounded number of indexes are held open at once (ZOTEUS_INDEX_MAX_OPEN, default 4); the least recently used is saved and closed to make room, which costs a reopen and nothing else. To search across several libraries at once, use zotero_semantic_search's libraries argument, which fans out over their indexes and labels each hit with the library it came from.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoMax items to index. Lowers the configured cap for this build only; it cannot raise it. The cap defaults to 5000 and is set by ZOTEUS_INDEX_MAX_ITEMS.
actionYesWhat to do. "update" is the cheap delta and the right default for an indexed library; "build" rebuilds (resuming an interrupted build) and "refresh" always starts over; "status" polls progress; "stop" cancels a running job; "pause"/"resume" hold index work across restarts; "libraries" lists which libraries have an index here, with their sizes (it starts nothing, and reports every library rather than the one library_type/library_id would name).
fulltextNoAlso index the full text Zotero extracted from each item's attachments, so searches match the body of a PDF. Resource-intensive (slower build, much larger index); defaults to ZOTEUS_INDEX_FULLTEXT (off unless set).
own_wordsNoAlso index the reader's OWN words — child notes and PDF annotations (highlight text and comments) — as passages carrying the parent item's key. On by default (ZOTEUS_INDEX_OWN_WORDS); the whole corpus is one paged crawl of hand-written text, so it costs a fraction of what fulltext does.
library_idNoNumeric id of the library to address, e.g. 5234875 for a group (zotero_groups lists the ids you can reach). Omit to use the configured default library; an id given without library_type is read as a group id.
library_typeNoWhich library to address: "user" (a personal library) or "group" (a shared group library). Omit to use the library this server is configured for. "group" on its own is refused: pass library_id with it.
fulltext_max_charsNoCap on indexed full-text characters per item; 0 means no cap (default 40000). Only used with fulltext.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
itemsNoLibrary items represented in the index.
phaseNoWhich pass of a build is running: "metadata" or "fulltext".
stateNoLifecycle of the background job: "idle", "building", "done" or "error".
pausedNoWhether index work is held until action:"resume".
libraryNoWhich library's rows this index holds: "user" for the personal library, or "group:<id>" for a group. One index file holds one library, so a build or update for a different one is refused rather than allowed to erase these rows. Absent on an index built before this stamp existed, which guards nothing because there is no way to know whose rows it holds.
storageNoWhere the index lives: "sqlite" or "memory".
vectorsNoPassages that also carry an embedding.
embedderNoThe embedder actually producing vectors, or "none (...)" with the reason.
passagesNoAlias of `documents`.
repairedNoWhat an unreadable index had to have removed before this build could start.
documentsNoPassages held for keyword search.
lastErrorNoSet when state is "error".
librariesNoEvery library with an index in this data directory (action:"libraries").
operationNoWhich job the counters describe: "build" or "update".
itemsTotalNoItems this job expects to index (0 = not yet known).
itemsFetchedNoItems pulled from Zotero so far (on an update: changed items processed).
itemsRemovedNoItems an update dropped because the library no longer holds them.
persistErrorNoLast failure to write the index to disk; the results exist only until restart.
embedderActiveNoTrue only while that provider is genuinely producing vectors.
embedderReasonNoWhy the configured embedder is not active, and what to do about it.
itemsAvailableNoItems the library holds before the build cap is applied.
libraryBackendNoWhich API issued that version: "local" or "cloud" (the two sequences are not comparable).
libraryVersionNoZotero library version this index was last built or updated from.
fulltextEnabledNoWhether attachment body text was indexed.
fulltextVersionNoHow far into Zotero's separate full-text sequence this index has read.
ownWordsEnabledNoWhether the reader's own notes and annotations were indexed.
embedderConfiguredNoThe requested ZOTEUS_EMBEDDINGS value, whether or not it works.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv1.21.0
    • changedInput schema / properties / action / description
      Previous value: -"What to do. \"update\" is the cheap delta and the right default for an indexed library; \"build\" rebuilds (resuming an interrupted build) and \"refresh\" always starts over; \"status\" polls progress; \"stop\" cancels a running job; \"pause\"/\"resume\" hold index work across restarts."New value: +"What to do. \"update\" is the cheap delta and the right default for an indexed library; \"build\" rebuilds (resuming an interrupted build) and \"refresh\" always starts over; \"status\" polls progress; \"stop\" cancels a running job; \"pause\"/\"resume\" hold index work across restarts; \"libraries\" lists which libraries have an index here, with their sizes (it starts nothing, and reports every library rather than the one library_type/library_id would name)."
    • changedInput schema / properties / action / enum
      Previous value: -[
      -  "build",
      -  "refresh",
      -  "update",
      -  "status",
      -  "stop",
      -  "pause",
      -  "resume"
      -]New value: +[
      +  "build",
      +  "refresh",
      +  "update",
      +  "status",
      +  "stop",
      +  "pause",
      +  "resume",
      +  "libraries"
      +]
    • addedInput schema / properties / library_id / exclusiveMinimum
      Added value: +0
    • addedOutput schema / properties / libraries
      Added value: +{
      +  "description": "Every library with an index in this data directory (action:\"libraries\").",
      +  "items": {
      +    "additionalProperties": false,
      +    "properties": {
      +      "documents": {
      +        "description": "Passages held for keyword search.",
      +        "type": "number"
      +      },
      +      "fault": {
      +        "description": "Why this index could not be opened or read at all.",
      +        "type": "string"
      +      },
      +      "items": {
      +        "description": "Library items represented in this index.",
      +        "type": "number"
      +      },
      +      "label": {
      +        "description": "The same library in words: \"the personal library\" or \"group 4523\".",
      +        "type": "string"
      +      },
      +      "library": {
      +        "description": "Canonical token of the library this index holds: \"user\" or \"group:<id>\".",
      +        "type": "string"
      +      },
      +      "libraryVersion": {
      +        "description": "Zotero library version it was last built or updated from (0 = none).",
      +        "type": "number"
      +      },
      +      "path": {
      +        "description": "Absolute path of the index file (the SQLite database sits beside it).",
      +        "type": "string"
      +      },
      +      "primary": {
      +        "description": "True for the default library's index, the one a call that names no library uses.",
      +        "type": "boolean"
      +      },
      +      "stamp": {
      +        "description": "The store's own library stamp. Absent on an index built before the stamp existed, which guards nothing.",
      +        "type": "string"
      +      },
      +      "state": {
      +        "description": "Lifecycle of this index's own job: \"idle\", \"building\", \"done\" or \"error\".",
      +        "type": "string"
      +      },
      +      "vectorEmbedder": {
      +        "description": "Identity of the vectors this index HOLDS, absent when it holds none. Two indexes with different values were embedded by different models, so their scores are not comparable.",
      +        "type": "string"
      +      },
      +      "vectors": {
      +        "description": "Passages that also carry an embedding.",
      +        "type": "number"
      +      }
      +    },
      +    "required": [
      +      "library",
      +      "label",
      +      "path",
      +      "primary",
      +      "documents",
      +      "items",
      +      "vectors",
      +      "state",
      +      "libraryVersion"
      +    ],
      +    "type": "object"
      +  },
      +  "type": "array"
      +}
    • addedOutput schema / properties / library
      Added value: +{
      +  "description": "Which library's rows this index holds: \"user\" for the personal library, or \"group:<id>\" for a group. One index file holds one library, so a build or update for a different one is refused rather than allowed to erase these rows. Absent on an index built before this stamp existed, which guards nothing because there is no way to know whose rows it holds.",
      +  "type": "string"
      +}
  2. Changed2 schema fields changedv1.20.2
    • removedInput schema / $schema
      Removed value: -"http://json-schema.org/draft-07/schema#"
    • removedOutput schema / $schema
      Removed value: -"http://json-schema.org/draft-07/schema#"
  3. Changed4 schema fields changedv1.20.0
    • addedInput schema / properties / action / description
      Added value: +"What to do. \"update\" is the cheap delta and the right default for an indexed library; \"build\" rebuilds (resuming an interrupted build) and \"refresh\" always starts over; \"status\" polls progress; \"stop\" cancels a running job; \"pause\"/\"resume\" hold index work across restarts."
    • addedInput schema / properties / library_id / description
      Added value: +"Numeric id of the library to address, e.g. 5234875 for a group (zotero_groups lists the ids you can reach). Omit to use the configured default library; an id given without library_type is read as a group id."
    • addedInput schema / properties / library_type / description
      Added value: +"Which library to address: \"user\" (a personal library) or \"group\" (a shared group library). Omit to use the library this server is configured for. \"group\" on its own is refused: pass library_id with it."
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "$schema": "http://json-schema.org/draft-07/schema#",
      +  "additionalProperties": true,
      +  "properties": {
      +    "documents": {
      +      "description": "Passages held for keyword search.",
      +      "type": "number"
      +    },
      +    "embedder": {
      +      "description": "The embedder actually producing vectors, or \"none (...)\" with the reason.",
      +      "type": "string"
      +    },
      +    "embedderActive": {
      +      "description": "True only while that provider is genuinely producing vectors.",
      +      "type": "boolean"
      +    },
      +    "embedderConfigured": {
      +      "description": "The requested ZOTEUS_EMBEDDINGS value, whether or not it works.",
      +      "type": "string"
      +    },
      +    "embedderReason": {
      +      "description": "Why the configured embedder is not active, and what to do about it.",
      +      "type": "string"
      +    },
      +    "fulltextEnabled": {
      +      "description": "Whether attachment body text was indexed.",
      +      "type": "boolean"
      +    },
      +    "fulltextVersion": {
      +      "description": "How far into Zotero's separate full-text sequence this index has read.",
      +      "type": "number"
      +    },
      +    "items": {
      +      "description": "Library items represented in the index.",
      +      "type": "number"
      +    },
      +    "itemsAvailable": {
      +      "description": "Items the library holds before the build cap is applied.",
      +      "type": "number"
      +    },
      +    "itemsFetched": {
      +      "description": "Items pulled from Zotero so far (on an update: changed items processed).",
      +      "type": "number"
      +    },
      +    "itemsRemoved": {
      +      "description": "Items an update dropped because the library no longer holds them.",
      +      "type": "number"
      +    },
      +    "itemsTotal": {
      +      "description": "Items this job expects to index (0 = not yet known).",
      +      "type": "number"
      +    },
      +    "lastError": {
      +      "description": "Set when state is \"error\".",
      +      "type": "string"
      +    },
      +    "libraryBackend": {
      +      "description": "Which API issued that version: \"local\" or \"cloud\" (the two sequences are not comparable).",
      +      "type": "string"
      +    },
      +    "libraryVersion": {
      +      "description": "Zotero library version this index was last built or updated from.",
      +      "type": "number"
      +    },
      +    "operation": {
      +      "description": "Which job the counters describe: \"build\" or \"update\".",
      +      "type": "string"
      +    },
      +    "ownWordsEnabled": {
      +      "description": "Whether the reader's own notes and annotations were indexed.",
      +      "type": "boolean"
      +    },
      +    "passages": {
      +      "description": "Alias of `documents`.",
      +      "type": "number"
      +    },
      +    "paused": {
      +      "description": "Whether index work is held until action:\"resume\".",
      +      "type": "boolean"
      +    },
      +    "persistError": {
      +      "description": "Last failure to write the index to disk; the results exist only until restart.",
      +      "type": "string"
      +    },
      +    "phase": {
      +      "description": "Which pass of a build is running: \"metadata\" or \"fulltext\".",
      +      "type": "string"
      +    },
      +    "repaired": {
      +      "description": "What an unreadable index had to have removed before this build could start."
      +    },
      +    "state": {
      +      "description": "Lifecycle of the background job: \"idle\", \"building\", \"done\" or \"error\".",
      +      "type": "string"
      +    },
      +    "storage": {
      +      "description": "Where the index lives: \"sqlite\" or \"memory\".",
      +      "type": "string"
      +    },
      +    "vectors": {
      +      "description": "Passages that also carry an embedding.",
      +      "type": "number"
      +    }
      +  },
      +  "type": "object"
      +}
  4. Changed1 schema field changedv1.16.0
    • changedInput schema / properties / action / enum
      Previous value: -[
      -  "build",
      -  "refresh",
      -  "update",
      -  "status",
      -  "stop"
      -]New value: +[
      +  "build",
      +  "refresh",
      +  "update",
      +  "status",
      +  "stop",
      +  "pause",
      +  "resume"
      +]
  5. Changed1 schema field changedv1.13.0
    • addedInput schema / properties / own_words
      Added value: +{
      +  "description": "Also index the reader's OWN words — child notes and PDF annotations (highlight text and comments) — as passages carrying the parent item's key. On by default (ZOTEUS_INDEX_OWN_WORDS); the whole corpus is one paged crawl of hand-written text, so it costs a fraction of what fulltext does.",
      +  "type": "boolean"
      +}
  6. Changed3 schema fields changedv1.7.1
    • changedInput schema / properties / action / enum
      Previous value: -[
      -  "build",
      -  "refresh",
      -  "status",
      -  "stop"
      -]New value: +[
      +  "build",
      +  "refresh",
      +  "update",
      +  "status",
      +  "stop"
      +]
    • changedInput schema / properties / limit / description
      Previous value: -"Max items to index (default 5000, which is also the hard cap)."New value: +"Max items to index. Lowers the configured cap for this build only; it cannot raise it. The cap defaults to 5000 and is set by ZOTEUS_INDEX_MAX_ITEMS."
    • removedInput schema / properties / limit / maximum
      Removed value: -5000
  7. Changed2 schema fields changedv1.6.0
    • addedInput schema / properties / fulltext
      Added value: +{
      +  "description": "Also index the full text Zotero extracted from each item's attachments, so searches match the body of a PDF. Resource-intensive (slower build, much larger index); defaults to ZOTEUS_INDEX_FULLTEXT (off unless set).",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / fulltext_max_chars
      Added value: +{
      +  "description": "Cap on indexed full-text characters per item; 0 means no cap (default 40000). Only used with fulltext.",
      +  "maximum": 1000000,
      +  "minimum": 0,
      +  "type": "integer"
      +}
  8. Changed2 schema fields changedv1.3.1
    • changedInput schema / properties / action / enum
      Previous value: -[
      -  "build",
      -  "refresh",
      -  "status"
      -]New value: +[
      +  "build",
      +  "refresh",
      +  "status",
      +  "stop"
      +]
    • addedInput schema / properties / limit
      Added value: +{
      +  "description": "Max items to index (default 5000, which is also the hard cap).",
      +  "maximum": 5000,
      +  "minimum": 1,
      +  "type": "integer"
      +}
  9. First observedv1.0.4

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations only say destructiveHint=true and readOnlyHint=false, but the description adds a wealth of behavioral context: background execution, immediate return, checkpoint/resume behavior, fallback from local API to web API, rate-limit handling, one-index-per-library isolation, and the fact that a stopped update leaves the version stamp untouched. It also discloses pitfalls like the Claude Desktop embedding crash and how the tool repairs an unreadable index. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely long and dense, with many nested parentheticals and run-on sentences that make it harder to scan. It is front-loaded with the most important action-selection guidance and organized by action, but the sheer volume of caveats, status fields, and edge cases exceeds what conciseness would demand. Every sentence has content, but several concepts are repeated or buried in complex asides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description is exceptionally complete. It covers all eight actions, the full status payload, crash and restart behavior, interrupted-build resumption, rate limiting, multi-library isolation, environment variables, and even the interaction with zotero_annotate and zotero_semantic_search. Since an output schema exists, the description does not need to document return values, but it still explains the meaning of the key status fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description goes far beyond the schema by explaining each parameter's real-world effect. For example, it clarifies that limit 'cannot raise' the configured cap, that fulltext multiplies build time and index size, that own_words covers child notes and PDF annotations, that library_type/library_id pick the index FILE rather than just the crawl, and that fulltext_max_chars counts per-item body text. This is substantive added meaning, not schema repetition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Manage the local hybrid-search index used by zotero_semantic_search.' It then enumerates eight discrete actions, making the tool's role and scope unmistakable. It also distinguishes itself from related tools by clarifying that this is index management, while zotero_semantic_search performs searches and zotero_annotate writes the text that gets indexed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance for each action: 'update' is 'the cheap one and should be the default for a library that is already indexed'; 'build' is for first builds, model changes, or widening a capped build; 'refresh' 'always starts over.' It also tells the user what NOT to do, e.g., 'NOT action: "refresh"' when resuming a rate-limited build, and explicitly routes cross-library search to zotero_semantic_search's libraries argument.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.