Skip to main content
Glama

pdf_flatten

Flatten PDF — Flatten PDF forms and annotations into static page content. Supports granular modes (annotations-only, forms-only, all), page ranges, signature-aware handling, link preservation, watermark stamping, image compression, PDF/A archival output, OCR for scanned inputs, and a ZIP bundle that exports form values + annotation metadata alongside the flattened file. [category: pdf]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fileYesInput PDF
modeNoWhich interactive elements to flatten. 'none' runs no flatten (useful for OCR/watermark/PDFA-only pipelines).all
pagesNoOptional page range (e.g. '1-3,5,7-9'). Only listed pages are flattened; others stay interactive.
ocrLangNoWhich language the scanned text is in.eng
ocrFirstNoRun ocrmypdf before flattening (for scanned PDFs). Requires Starter+ tier.
exportDataNoAlso give me the form answers and note text as separate files. You get a ZIP containing the flattened PDF plus form_values.json and annotations.json - not a PDF.
outputFormatNopdfa produces a PDF/A-2b archival output. Requires Starter+ tier.pdf
preserveLinksNoKeep clickable hyperlinks after flattening (uses qpdf --flatten-annotations=print).
signatureModeNoWhat to do if the PDF has been digitally signed. Flattening destroys a signature, so by default we hand the original back untouched.preserve
watermarkFontNoExactly Helvetica, Times-Roman, or Courier (case-sensitive); anything else becomes Helvetica. Read only when watermarkText is set.Helvetica
watermarkTextNoText watermark to stamp before flattening. Leave empty to skip.
compressImagesNoDownsample images after flattening to shrink file size.
compressPresetNoHow hard to squeeze the pictures: screen 72 DPI, ebook 150 DPI, printer and prepress 300 DPI.ebook
outputFilenameNoOptional custom filename for the flattened output (without path).
watermarkColorNoHex color, #rgb or #rrggbb.#808080
watermarkScaleNoAbsolute scale factor; default 1.0. Non-numeric resets to 1.0.
watermarkOpacityNo0 = invisible, 1 = solid; default 0.3. Non-numeric resets to 0.3. Read only when watermarkText is set.
watermarkFontSizeNoPoint size, integer; non-integer input silently resets to 48. Read only when watermarkText is set.
watermarkPositionNoWhere the watermark sits on the page. Only used when there is watermark text.c
watermarkRotationNoDegrees, integer; default 45 = classic diagonal. Non-integer resets to 45. Read only when watermarkText is set.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • addedInput schema / properties / pages / title
      Added value: +"Pages to flatten"
    • addedInput schema / properties / pages / x-ui / no_recall
      Added value: +true
    • addedInput schema / properties / pages / x-ui / page_select
      Added value: +{}
  2. Changed2 schema fields changed
    • addedInput schema / properties / outputFilename / x-ui
      Added value: +{
      +  "unset_label": "Named after your file"
      +}
    • addedInput schema / properties / pages / x-ui
      Added value: +{
      +  "unset_label": "All pages"
      +}
  3. Changed3 schema fields changed
    • addedInput schema / properties / watermarkFontSize / x-ui
      Added value: +{
      +  "unit": "pt"
      +}
    • addedInput schema / properties / watermarkRotation / x-ui
      Added value: +{
      +  "unit": "deg"
      +}
    • addedInput schema / properties / watermarkScale / x-ui
      Added value: +{
      +  "unit": "x"
      +}
  4. Changed11 schema fields changed
    • changedInput schema / properties / compressPreset / description
      Previous value: -"screen|ebook|printer|prepress (Ghostscript). Read only when compressImages=true; unknown → ebook. screen=72dpi, printer/prepress=300dpi."New value: +"How hard to squeeze the pictures: screen 72 DPI, ebook 150 DPI, printer and prepress 300 DPI."
    • addedInput schema / properties / compressPreset / x-show-when
      Added value: +{
      +  "compressImages": [
      +    "true"
      +  ]
      +}
    • changedInput schema / properties / exportData / description
      Previous value: -"Return a ZIP containing the flattened PDF plus form_values.json and annotations.json side files."New value: +"Also give me the form answers and note text as separate files. You get a ZIP containing the flattened PDF plus form_values.json and annotations.json - not a PDF."
    • changedInput schema / properties / ocrLang / description
      Previous value: -"Allowlist: eng fra spa deu ita por nld pol chi_sim jpn kor ara rus hin; unknown → eng. Read only when ocrFirst=true (paid OCR tier)."New value: +"Which language the scanned text is in."
    • addedInput schema / properties / ocrLang / x-show-when
      Added value: +{
      +  "ocrFirst": [
      +    "true"
      +  ]
      +}
    • addedInput schema / properties / ocrLang / x-ui
      Added value: +{
      +  "labels": {
      +    "ara": "Arabic",
      +    "chi_sim": "Chinese (Simplified)",
      +    "deu": "German",
      +    "eng": "English",
      +    "fra": "French",
      +    "hin": "Hindi",
      +    "ita": "Italian",
      +    "jpn": "Japanese",
      +    "kor": "Korean",
      +    "nld": "Dutch",
      +    "pol": "Polish",
      +    "por": "Portuguese",
      +    "rus": "Russian",
      +    "spa": "Spanish"
      +  }
      +}
    • changedInput schema / properties / signatureMode / description
      Previous value: -"preserve = return original when signatures detected; ignore = flatten anyway (invalidates sigs); block = 409 error."New value: +"What to do if the PDF has been digitally signed. Flattening destroys a signature, so by default we hand the original back untouched."
    • addedInput schema / properties / signatureMode / x-ui
      Added value: +{
      +  "labels": {
      +    "block": "Stop and tell me",
      +    "ignore": "Flatten anyway (breaks the signature)",
      +    "preserve": "Leave signed files untouched"
      +  }
      +}
    • changedInput schema / properties / watermarkPosition / description
      Previous value: -"pdfcpu anchor: c tl tc tr ml mr bl bc br (ml/mr are folded to the engine's l/r); unknown → c (center). Read only when watermarkText is set."New value: +"Where the watermark sits on the page. Only used when there is watermark text."
    • addedInput schema / properties / watermarkPosition / x-ui
      Added value: +{
      +  "labels": {
      +    "bc": "Bottom centre",
      +    "bl": "Bottom left",
      +    "br": "Bottom right",
      +    "c": "Centre",
      +    "ml": "Middle left",
      +    "mr": "Middle right",
      +    "tc": "Top centre",
      +    "tl": "Top left",
      +    "tr": "Top right"
      +  }
      +}
    • removedInput schema / properties / watermarkTile
      Removed value: -{
      -  "default": false,
      -  "description": "NOT available here — true returns a clear error (the flatten tile path is broken in the pinned engine; run pdf_watermark, which tiles, before flattening).",
      -  "type": "boolean"
      -}
  5. First observed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint/destructiveHint false), so the description carries the transparency burden and does so well: it discloses signature-aware handling ('Flattening destroys a signature' — reinforced in the schema), the ZIP bundle that returns JSON sidecar files instead of a PDF for exportData, and the Starter+ tier requirement for OCR. No contradiction with the annotations; destructiveHint:false is consistent since flattening emits a new file and returns signed originals untouched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first clause, which is good, but the rest is a single dense run-on sentence cramming ~10 features into a comma-separated list. It is informative but difficult to scan quickly; it could be broken into a purpose statement plus a short capabilities list. The trailing [category: pdf] tag is useful for discovery.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 20-parameter tool with no output schema, the description covers the major behaviors: flatten modes, page ranges, signature policy, link preservation, watermarking, compression, PDF/A archival, OCR, and the ZIP sidecar export. It does not explicitly state that the default output is a flattened PDF, though this is implied by 'alongside the flattened file.' Given the tool's complexity and absent output schema, coverage is strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description enumerates capabilities that align with parameters (mode, pages, signatureMode, watermarkText, compressImages, outputFormat, ocrFirst) but adds no new semantic detail beyond what the schema already richly documents — e.g., signatureMode's default-preserve behavior and the enum labels are already in the schema. Description adds marginal value here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause 'Flatten PDF forms and annotations into static page content' is a specific verb+resource+outcome statement. The subsequent feature list (modes, page ranges, signature handling, watermark, compression, PDF/A, OCR, ZIP) maps directly onto the 20-parameter schema and differentiates it from siblings like pdf_watermark, pdf_compress, pdf_ocr, and pdf_to_pdfa.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied through the capability enumeration, but there is no explicit guidance on when to prefer this tool over its siblings. Notably, pdf_flatten_batch exists as a sibling yet the description never names it or states when the single-file version is appropriate, nor does it point to pdf_watermark for watermark-only jobs. No when/when-not routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources