Skip to main content
Glama

pdf_to_excel

PDF to Excel — Extract tables from PDFs into XLSX / CSV / TSV / JSON. Uses tabula-java (lattice + stream modes) with LibreOffice as fallback. Supports page ranges, table selection, sheet strategy (per-table/per-page/single), OCR for scanned PDFs (Starter+), JSON output (Starter+), and a non-destructive inspect endpoint that reports row/col counts plus ragged/sparse confidence flags. [category: pdf]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fileYesInput PDF
pagesNoOptional page range e.g. '1-5,10'. Empty = all pages.
engineNoTable detection engine. auto = tabula lattice → stream → libreoffice fallback.auto
formatNoxlsx/csv/tsv are file downloads; json returns structured data.xlsx
ocrLangNoThe language of the writing in the scan.eng
ocrFirstNoRun ocrmypdf before extraction (beta — scanned PDFs).
sheetModeNoHow the tables are laid out across the workbook's sheets.per-table
tableIndexesNoComma-separated 0-based indexes to keep (e.g. '0,2,3'). Empty = all tables.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • addedInput schema / properties / pages / title
      Added value: +"Pages to read"
    • addedInput schema / properties / pages / x-ui
      Added value: +{
      +  "no_recall": true,
      +  "page_select": {}
      +}
  2. Changed6 schema fields changed
    • changedInput schema / properties / ocrLang / description
      Previous value: -"Any Tesseract code, passed raw to ocrmypdf -l (default eng). Read only when ocrFirst=true; on OCR failure extraction continues un-OCR'd."New value: +"The language of the writing in the scan."
    • addedInput schema / properties / ocrLang / x-show-when
      Added value: +{
      +  "ocrFirst": [
      +    "true"
      +  ]
      +}
    • addedInput schema / properties / ocrLang / x-ui
      Added value: +{
      +  "labels": {
      +    "ara": "Arabic",
      +    "chi_sim": "Chinese (Simplified)",
      +    "deu": "German",
      +    "eng": "English",
      +    "fra": "French",
      +    "hin": "Hindi",
      +    "ita": "Italian",
      +    "jpn": "Japanese",
      +    "kor": "Korean",
      +    "nld": "Dutch",
      +    "pol": "Polish",
      +    "por": "Portuguese",
      +    "rus": "Russian",
      +    "spa": "Spanish"
      +  }
      +}
    • changedInput schema / properties / sheetMode / description
      Previous value: -"XLSX sheet strategy. CSV/TSV/JSON ignore this."New value: +"How the tables are laid out across the workbook's sheets."
    • addedInput schema / properties / sheetMode / x-show-when
      Added value: +{
      +  "format": [
      +    "xlsx"
      +  ]
      +}
    • addedInput schema / properties / sheetMode / x-ui
      Added value: +{
      +  "labels": {
      +    "per-page": "One sheet per page",
      +    "per-table": "One sheet per table",
      +    "single": "Everything on one sheet"
      +  }
      +}
  3. First observed

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate not read-only and not destructive; the description adds meaningful behavioral context beyond that: it discloses the underlying engine stack ('tabula-java with LibreOffice fallback'), plan-tier restrictions (Starter+), and the presence of a non-destructive inspect path with confidence flags. These details help an agent predict side effects and dependencies beyond what annotations cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two information-dense sentences with the primary purpose front-loaded. The trailing '[category: pdf]' and the 'non-destructive inspect endpoint' clause are somewhat tangential and could be removed without losing core value, but the structure is otherwise clean and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers most parameters and engine details, but with no output schema it does not specify the return value format (e.g., file ID vs. inline data for XLSX/CSV/TSV). It also fails to mention the batch sibling or single-file vs. multi-file distinction, which is important for tool selection. The reference to the inspect endpoint slightly muddies the boundary between this tool and pdf_to_excel_inspect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, giving a baseline of 3. The description goes beyond the schema by mapping high-level features ('page ranges, table selection, sheet strategy') to the relevant parameters and even enumerating the sheet strategy options inline ('per-table/per-page/single'). This reinforces the agent's understanding of what values those parameters accept.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource: 'Extract tables from PDFs into XLSX / CSV / TSV / JSON.' This unambiguously distinguishes the core function from siblings like pdf_to_text or pdf_to_images steals no ambiguity. The mention of engine modes and formats further nails down what the tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for single-PDF table extraction but does not explicitly say when to choose this over siblings like pdf_to_excel_batch or pdf_to_excel_inspect. It mentions a 'non-destructive inspect endpoint' without clarifying that this is a separate sibling tool, which could mislead an agent. No explicit exclusions or alternative conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources