pdf_ocr
Extract text from scanned PDFs using OCR, and optionally create a searchable PDF with a text layer. Get confidence scores to verify accuracy.
Instructions
Read a scanned PDF with OCR, and optionally add a real text layer to it.
The way to read a page with no text: a scan, or an export that lost its text. Prefer it to pdf_render_pages for reading; run pdf_check_text when unsure. Tesseract is a system install, and a missing binary or language errors with how to install it.
OCR can be confidently wrong, so pages report mean_confidence and
low_confidence_words: say when confidence is low rather than presenting text as
certain, and look at a bad page with pdf_render_pages.
Text pages are skipped unless force. A long document may stop early; when
truncated, do exactly what the summary says, artifact and range included.
Args: ref: PDF file path, or a workspace artifact id. pages: Which pages: "1-10", "3", "1,5,9-12" or "all". Defaults to all. lang: Tesseract language code, or several as "eng+deu". dpi: Resolution, 72 to 600. 200 suits printed text; higher is slower, not better. output: "text", "pdf" for a searchable copy as an artifact, or "both". force: OCR pages that already have text instead of skipping them.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| ref | Yes | ||
| lang | No | eng | |
| force | No | ||
| pages | No | ||
| output | No | text |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |