photo_to_text
Photo to Text (OCR) — Extract text from an image via OCR. Language selection takes two-letter codes (en, fr, de, ...) separated by commas; Tesseract codes such as eng or chi_sim also work. [category: photo]
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | JPG, PNG, WebP, BMP, TIFF (max 15MB) | |
| output | No | json returns structured results; text returns plain text. | json |
| binarize | No | Force the picture to pure black and white before reading it. Off by default because it destroys text in uneven light; try it on faint or washed-out scans. | |
| languages | No | Which language or languages the writing is in. One code, or several separated by commas (en,fr); a plus sign works too, as Tesseract writes it (eng+fra). Two-letter codes are the usual form: en, fr, de, es, pt, it, nl, ru, ar, zh, ja, ko. Tesseract's own codes (eng, fra, deu, chi_sim, chi_tra, ...) are also accepted; each of the two reading engines is handed its own spelling of the language. A language whose reading pack is not installed on our server is refused with a message that says so, rather than reported as an engine fault. Letters and underscores only, not case-sensitive; a value with no code in it, such as a lone comma, is refused with a 400. Field name is languages, not language. | en |
| preprocess | No | Apply image preprocessing before OCR. | |
| binarize_threshold | No | The cut-off between black and white, as a percent. Lower keeps more of the picture black. Only used when black and white is forced on; anything outside 1-99 quietly reverts to 60. |