Skip to main content
Glama

document_search

Read-onlyIdempotent

Search PDF, DOCX, MD, or TXT files for exact text or regex matches, returning verbatim results with page, line, offsets, context, and match IDs for follow-up reading.

Instructions

Exact search inside a PDF/DOCX/MD/TXT. Returns each match verbatim with its page/line/paragraph, character offsets, surrounding context and a match_id for document_read_around. Whitespace in the query matches any line break or spacing. Units that cannot be read are listed: no match there is not evidence of absence. [docbridge schema 0.2.4]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryYes
regexNo
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
input_pathYesAbsolute file path.
max_matchesNo
start_matchNo
context_charsNo
case_sensitiveNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
toolYes
errorNo
notesNo
sourceNo
matchesNo
patternNo
invariantNodocbridge may reduce the volume of its own responses, but it never alters, summarizes, paraphrases, semantically filters, or silently truncates source evidence
matches_totalNo
case_sensitiveNo
schema_versionYes
matches_returnedNo
next_start_matchNo
unsearchable_unitsNo
operation_completedYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.4

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent/closed-world annotations by disclosing matching semantics ('Whitespace in the query matches any line break or spacing') and a critical caveat about partial coverage ('Units that cannot be read are listed: no match there is not evidence of absence'). It also describes the shape of each returned match (verbatim text, page/line/paragraph, offsets, context, match_id).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then layers return shape, matching rule, and the unreadable-units caveat in efficient sentences. The trailing '[docbridge schema 0.2.4]' tag is minor noise but does not obscure the content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With rich annotations and an output schema already in place, the description covers the non-obvious essentials an agent needs: file type scope, literal matching semantics, unreadable-unit behavior, and the match_id linkage. The remaining gap is parameter behavior (regex/case/pagination), which is the same deficiency already noted under parameter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (encoding and input_path only), so the description must compensate for six undocumented parameters. It adds query whitespace semantics but says nothing about regex, case_sensitive, max_matches/start_match pagination, or context_chars, leaving those meanings to be inferred from parameter names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Exact search') and resource ('inside a PDF/DOCX/MD/TXT'), and explicitly links its output to the sibling tool document_read_around via match_id. An agent can distinguish it from document_read, document_outline, and the conversion siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies the use case (literal/exact search in a single document) and signals a workflow by mentioning match_id for document_read_around, but gives no explicit when-to-use/when-not or comparison against alternatives like document_read or document_extract. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.