Skip to main content
Glama

pdf_to_excel_batch

PDF to Excel (Batch) — Apply the same PDF-to-Excel configuration to up to 20 PDFs. Returns a ZIP with per-file subfolders. [category: pdf]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
filesYesUp to 20 input PDFs
pagesNoOptional page range e.g. '1-5,10'. Empty = all pages.
engineNoTable detection engine. auto = tabula lattice → stream → libreoffice fallback.auto
formatNoxlsx/csv/tsv are file downloads; json returns structured data.xlsx
ocrLangNoAny Tesseract code, passed raw to ocrmypdf -l (default eng). Read only when ocrFirst=true; on OCR failure extraction continues un-OCR'd.eng
ocrFirstNoRun ocrmypdf before extraction (beta — scanned PDFs).
sheetModeNoXLSX sheet strategy. CSV/TSV/JSON ignore this.per-table
tableIndexesNoComma-separated 0-based indexes to keep (e.g. '0,2,3'). Empty = all tables.

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. First observed

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations carry little behavioral signal beyond readOnlyHint=false and destructiveHint=false, so the description shoulders much of the burden. It does add useful context by stating the output is a ZIP with per-file subfolders, but it overgeneralizes: the format parameter allows json, which the schema says returns structured data instead of a file download. The description also does not mention behavior on OCR failure or partial processing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it states the operation, the batch size, and the output format in two sentences. The opening 'PDF to Excel (Batch)' is somewhat redundant with the tool name and title, and the category tag adds little, so it is not a perfect 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a relatively complex tool with 8 parameters, 4 enums, and no output schema, so the description should clarify expected return shapes. It states a ZIP output unconditionally, but the schema's format=json option returns structured data, meaning the description is incomplete and potentially misleading for non-default formats. It also omits guidance on how the batch output maps to individual files beyond 'per-file subfolders.'

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter already has an explicit description, enum, or default in the input schema. The tool description adds no additional parameter-level meaning beyond referring to the 'same PDF-to-Excel configuration,' so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation—applying a PDF-to-Excel configuration to PDFs—along with a clear constraint (up to 20 PDFs) and a concrete output (ZIP with per-file subfolders). The 'Batch' label and 20-file limit distinguish it from sibling tools like pdf_to_excel and pdf_to_excel_inspect without needing to open their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for batch processing multiple PDFs and gives an explicit limit of 20 files, which is useful usage context. However, it does not explicitly state when to prefer the single-file pdf_to_excel tool or the configuration-inspection pdf_to_excel_inspect tool, so the agent must infer routing from sibling names rather than receiving direct guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.2/5.0
Disambiguation2/5

Multiple tool pairs are near-identical: octopus_mkdir/octopus_make_folder and octopus_move/octopus_move_file are literal duplicates, analyze_hash/generate_hash both compute hashes, convert_word_to_pdf overlaps convert_document, and photo_compress/photo_compress_to_size plus pdf_thumbnails/pdf_to_images have fuzzy boundaries. The descriptions are detailed and cross-reference each other helpfully, but at 144 tools an agent will regularly misselect.

Naming Consistency3/5

The dominant {category}_{verb}_{object} snake_case pattern (pdf_*, photo_*, convert_*, analyze_*, media_*) is largely consistent and predictable. However, outliers like chatwithyourpdf and describe_image break the category-prefix convention, and the octopus namespace mixes bare verbs (read, write, mkdir) with verb_noun forms (make_folder, move_file, search_meta) inconsistently.

Tool Count2/5

144 tools is an extreme count for any MCP server. The broad scope (PDF, photo, video, audio, conversion, analysis, generation, file storage, web, e-sign) justifies some volume, but the count is inflated by batch and inspect variants (pdf_to_excel + batch + inspect), duplicate tools, and overlapping converters. An agent faces an overwhelming selection surface.

Completeness4/5

Per-domain coverage is remarkably deep: PDF spans merge/split/compress/protect/unlock/metadata/OCR/watermark and bidirectional conversion; photo covers editing, format conversion, face handling, OCR, and collage; file storage has full CRUD plus search. Minor gaps exist (no audio transcription, no video metadata editing, no deletion of PDF pages is actually covered via pdf_delete_pages) but the surface has no dead ends for its declared domains.

Resources