Skip to main content
Glama
jgravelle
by jgravelle

validate_index

Read-only

Verify an indexed dataset's on-disk integrity: run SQLite PRAGMA integrity check, compare row counts and columns against index metadata, hash-verify index.json, and report stale-lock states.

Instructions

Verify an indexed dataset's on-disk integrity. Runs SQLite PRAGMA integrity_check, cross-checks row count and column list against index.json, and verifies index.json content hash. Reports stale-lock state from interrupted index_local runs. Returns overall_status: 'ok' | 'warning' | 'error'.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
datasetYesDataset identifier
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already indicates a safe read operation. The description adds substantial behavioral detail: exactly which integrity checks are performed (SQLite PRAGMA, row/column comparison, content hash), the fact that stale-lock state is reported, and the possible return values. This goes far beyond the annotation and gives the agent a precise model of tool behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long, with the purpose front-loaded in the first sentence. The second sentence details exactly what actions are taken, and the third states the return status. There is no fluff; every sentence adds necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a simple single-parameter schema, no output schema, and one annotation. The description fully covers what the tool does, how it does it, and what it returns. It also references sibling index_local for context, making it complete for selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides a description for the single 'dataset' parameter (100% coverage), so the baseline is 3. The tool description adds meaning by clarifying that the dataset must be an indexed dataset (since it validates 'an indexed dataset's on-disk integrity'), which refines the otherwise generic parameter description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific verb 'Verify' and a clear resource 'an indexed dataset's on-disk integrity.' It then lists concrete operations (PRAGMA integrity_check, row/column cross-check, hash verification) and the output status values, making it fully distinguishable from sibling tools like index_local or get_dataset_health.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool: to validate an indexed dataset, especially after potential interrupted index_local runs. It provides rich context (cross-checking against index.json, stale-lock reporting) but does not explicitly name alternatives or state when not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/jgravelle/jdatamunch-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server