Extract tables from a PDF
documents_tablesExtracts tables from a text-based PDF (price list, data sheet, report; up to 20 MiB, 60 pages) with page, bounding box and a quality level per table. Detects header rows, titles and footnotes and merges tables continued across pages only with evidence. Scanned pages are reported as OCR_REQUIRED, nothing is guessed. Works with any language embedded in the PDF. The PDF can be passed as upload or URL.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | The PDF. A file uploaded in ChatGPT (object with download_url and file_id). | |
| options | No | Optional TablesOptions, e.g. {"pages": "1-3,7", "strategy": "auto", "include_cells": false}. Schema: /v1/schemas/documents-tables-options. | |
| file_url | No | The PDF. Public URL of the file (http/https). Google Drive, Google Sheets and Dropbox share links are converted to direct downloads. Max 20 MiB. | |
| filename | No | Original filename with extension (e.g. prices.xlsx); used to detect the format. | |
| max_rows | No | Rows returned per table (counts stay complete). | |
| file_path | No | Local file path; only allowed when the server runs locally over stdio. | |
| file_base64 | No | The PDF. File content as Base64 (standard alphabet). Max 20 MiB decoded. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||