transcribe
Extract verbatim text from images, documents, screenshots, and video subtitles. Returns JSON output with frame timestamps for video sources.
Instructions
Verbatim transcription of text in images/documents/screenshots/video subtitles (no summary, judgment, or interpretation). Returns JSON with text and, for videos, frame timestamps.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| media | Yes | Image or video path / URL | |
| max_frames | No | Frame cap for video transcription (default 48) |