extract_content
Extract text from web pages, PDFs, documents, and YouTube videos. Optionally enable formula, image, or OCR extraction with engine selection.
Instructions
Extract content from a URL or file. Does not require an API key for most sources (web pages, PDFs, documents, YouTube transcripts). API key is only needed for audio/video transcription.
Args:
url: URL to extract content from (web page, YouTube video, PDF link, etc.)
file_path: Local file path to extract content from
engine: Optional extraction engine override, routed by input type.
With url: auto, simple, firecrawl, jina, crawl4ai.
With file_path: auto, simple, docling — docling requires
pip install "content-core[docling]" and fails with a
configuration error when the extra is missing, in which case use
auto or simple.
Any other value is rejected with an error naming the accepted ones.
formulas: Enable formula extraction via Docling (requires engine=docling)
pictures: Enable image description + chart data extraction via Docling (requires engine=docling)
no_ocr: Disable OCR in Docling (requires engine=docling)
Returns: Extracted text content
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| engine | No | ||
| no_ocr | No | ||
| formulas | No | ||
| pictures | No | ||
| file_path | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |