Skip to main content
Glama

repair_vocabulary

Audit and fix sparse vocabulary mismatches in codesearch by comparing Qdrant collections with registered contributions, reconstructing missing vectors, and removing stale registrations.

Instructions

Audit or repair codesearch's global sparse vocabulary.

Qdrant file and chunk points are authoritative. The audit compares their counts with each collection's registered vocabulary contribution and finds registrations whose Qdrant collection no longer exists. With repair enabled, mismatched contributions are reconstructed from stored sparse vectors and stale registrations are removed. Vocabulary reconstruction does not rewrite Qdrant; recovery of a pending interrupted operation may finish its recorded deletion or clear an ambiguously partial path before rebuilding the contribution. full re-registers every live collection and also requires repair; use it to correct token-frequency drift that happens to preserve the count. Collections are processed safely one at a time; if indexing changes the collection set concurrently, run the audit again for a point-in-time report.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fullNo
repairNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. Addedv1.8.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses that repair does not rewrite Qdrant, may complete a pending interrupted deletion, clears ambiguous partial paths before rebuilding, removes stale registrations, and processes collections one at a time. These are concrete, non-obvious behaviors beyond the tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than typical but every sentence adds meaningful detail about modes, side effects, recovery behavior, and concurrency. It is front-loaded with the core purpose and then expands into necessary edge cases without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a maintenance tool with no annotations, the description covers purpose, parameter behavior, side effects, recovery semantics, concurrency implications, and the distinction between repair and full modes. An output schema exists, so return-value documentation is not required here. This is a complete enough description for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It functionally defines both: "repair" enables reconstruction and stale-registration removal, while "full" re-registers every live collection and requires repair. It does not spell out the false/default behavior of each flag, but the meaning is strongly implied and the schema supplies defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Audit or repair codesearch's global sparse vocabulary." This clearly distinguishes the tool from siblings like list_collections, code_search, and cleanup_orphans, which operate on different concerns such as collection listing, search, or orphan cleanup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: audit-only by default, repair mode for reconstructing mismatched contributions and removing stale registrations, and full mode for correcting token-frequency drift. It also advises re-running the audit if indexing changes the collection set concurrently. It does not explicitly name alternatives or exclusions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/michaelkrauty/mcp-codesearch'

If you have feedback or need assistance with the MCP directory API, please join our Discord server