Skip to main content
Glama

pdf_ocr

OCR PDF (Make Searchable) — Make a scanned PDF searchable and selectable by adding an invisible OCR text layer over the page images — the pages look identical, but the text becomes findable, copyable, and indexable. Uses ocrmypdf (Tesseract + Ghostscript); already-searchable pages are skipped, so it is safe to run on mixed documents. This CREATES a text layer — to EXTRACT text that already exists, use pdf_to_text instead. [category: pdf]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fileYesScanned PDF to make searchable.
langNoThe language of the writing in the scan. The wrong language makes the searchable text gibberish.eng

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • changedInput schema / properties / lang / description
      Previous value: -"Tesseract language code(s): three letters, joinable with '+' (e.g. 'eng', 'deu', 'eng+fra'). The pack must be installed on the server."New value: +"The language of the writing in the scan. The wrong language makes the searchable text gibberish."
    • addedInput schema / properties / lang / enum
      Added value: +[
      +  "eng",
      +  "deu",
      +  "fra",
      +  "spa",
      +  "ita",
      +  "por",
      +  "nld",
      +  "rus",
      +  "jpn",
      +  "kor",
      +  "chi_sim",
      +  "chi_tra",
      +  "ara",
      +  "hin",
      +  "eng+fra",
      +  "eng+deu",
      +  "eng+spa"
      +]
    • addedInput schema / properties / lang / x-ui
      Added value: +{
      +  "labels": {
      +    "ara": "Arabic",
      +    "chi_sim": "Chinese (Simplified)",
      +    "chi_tra": "Chinese (Traditional)",
      +    "deu": "German",
      +    "eng": "English",
      +    "eng+deu": "English + German",
      +    "eng+fra": "English + French",
      +    "eng+spa": "English + Spanish",
      +    "fra": "French",
      +    "hin": "Hindi",
      +    "ita": "Italian",
      +    "jpn": "Japanese",
      +    "kor": "Korean",
      +    "nld": "Dutch",
      +    "por": "Portuguese",
      +    "rus": "Russian",
      +    "spa": "Spanish"
      +  }
      +}
  2. First observed

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations confirm it is not read-only but not destructive (it mutates the file by creating a layer). The description adds meaningful context beyond those hints: the pages look identical, already-searchable pages are skipped, and it is safe on mixed documents, plus the underlying ocrmypdf/Tesseract/Ghostscript stack. It stops short of describing output naming or processing-time expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and outcome are front-loaded before the implementation details and the sibling routing note. It is slightly long with multiple em-dash clauses and a redundant title prefix, but every sentence carries information an agent would use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still conveys what the result is (an identical-looking PDF with a searchable layer) and the safe-to-rerun behavior. Only minor gaps remain, such as output file naming and any size/time constraints, for a mutation tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both the binary input and the lang enum (with its warning that the wrong language yields gibberish) are already fully documented in the schema. The description adds no additional parameter-level detail, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Make a scanned PDF searchable... by adding an invisible OCR text layer') and explains the observable outcome ('text becomes findable, copyable, and indexable'). It explicitly distinguishes itself from the sibling pdf_to_text, so an agent can tell them apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (scanned PDFs needing OCR) and names the alternative with a clear routing rule: 'to EXTRACT text that already exists, use pdf_to_text instead'. It also notes the safe-rerun condition on mixed documents, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources