Skip to main content
Glama
Anselmoo

io.github.Anselmoo/mcp-ooxml-ledger

Find text

find_text
Read-onlyIdempotent

Search Office documents for case-insensitive text, returning precise locations such as paragraph, slide, or cell addresses with context snippets to pinpoint each hit.

Instructions

Case-insensitive substring search over every text-bearing part the digest covers, returning the best address this build can give for each hit: paragraph id or index and hash for docx, slide id for pptx, sheet and cell for xlsx. On pptx, real slides (in presentation order) are visited before slide layouts, masters and notes masters, so the default page surfaces slide content first. Each match's text is a bounded window around the hit, not the whole run — text_length and match_offset locate it in the full run, and text_truncated says whether it was cut; a large result set can also stop early on total response size, reported the same way as max_results paging: truncated=True. Results come from the session's working copy as opened; verify is what reports on the file currently on disk.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
partNo
queryYes
session_idYes
max_resultsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
partYes
queryYes
matchesYes
truncatedYes
session_idYes
baseline_digestYesThe canonical digest of this document AS IT WAS WHEN THE SESSION WAS OPENED, and of the working copy these results were read from — not an attestation about the file on disk right now. Call `verify` or `digest` for that.
document_may_have_changed_since_openYesTrue when the document FILE on disk no longer has the size and modification time it had when this session last touched it — so these results may describe a version that no longer exists. It reports writes from OUTSIDE this session: this server re-records both values after each of its own edits, so applying an edit here does not set it. A HINT, not a verification: it is also true after a save that changed nothing, and a rewrite that preserved both would not set it. Call `verify` for an answer about the file as it stands.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv0.3.0
    • changedOutput schema / properties / matches / items / description
      Previous value: -"One hit, with the best address this build can honestly give for its format."New value: +"One hit, with the best address this build can honestly give for its format.\n\n`text` is a bounded window, not the run's full decoded text: `text_length` is the whole\nrun's character count, `match_offset` is the needle's position within it (case-\ninsensitive), and `text_truncated` says whether `text` had to be cut down to\n`SNIPPET_RADIUS` chars either side of that offset. `start`/`end` are unaffected — still\nthe byte span of the WHOLE run in the raw part, because that is what a future edit\nsplices against."
    • addedOutput schema / properties / matches / items / properties / match_offset
      Added value: +{
      +  "type": "integer"
      +}
    • addedOutput schema / properties / matches / items / properties / text_length
      Added value: +{
      +  "type": "integer"
      +}
    • addedOutput schema / properties / matches / items / properties / text_truncated
      Added value: +{
      +  "type": "boolean"
      +}
    • changedOutput schema / properties / matches / items / required
      Previous value: -[
      -  "part",
      -  "text",
      -  "start",
      -  "end"
      -]New value: +[
      +  "part",
      +  "text",
      +  "text_length",
      +  "text_truncated",
      +  "match_offset",
      +  "start",
      +  "end"
      +]
  2. First observedv0.1.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint=false, and idempotentHint. The description adds substantial context beyond this: case-insensitivity, bounded match windows with text_length and match_offset, text_truncated flag, early stop on large result sets, and the pptx traversal order. No contradiction with annotations; it enriches the safety profile with operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence adds distinct value: purpose/output, pptx ordering, window semantics, truncation, and working-copy distinction. It is front-loaded with the main purpose and avoids redundant phrasing. The density is justified by the tool's complexity, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool operating across multiple document formats with an output schema, the description covers the search scope, address resolution, ordering rules, result windowing, truncation signaling, and the distinction from verify. An agent has all necessary behavioral and contextual information to call it correctly, including edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain max_results behavior (paging and truncated flag) and implies the meaning of query and session_id by context. However, it does not explicitly describe the 'part' parameter (e.g., what values it accepts or its default scope), nor does it detail the exact format for session_id. The names are self-explanatory to some degree, but the burden is on the description given zero schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs case-insensitive substring search across text-bearing parts of a digest, and specifies the output address format for each format (docx, pptx, xlsx). It distinguishes itself from siblings like verify (working copy vs disk) and describe_structure (search vs structure). The verb and resource are explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly contrasts with verify, noting results come from the session's working copy and verify reports on disk. It also provides behavioral guidance for pptx (presentation order before layouts/masters) and explains truncation/paging behavior. This gives clear when-to-use and when-not-to-use context relative to at least one sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.