article_extractor
Extract clean article content, metadata, and images from any news or blog URL. Returns one row per URL in Markdown, plain text, or HTML for RAG and LLM ingestion.
Instructions
Article & News Extractor returns clean article text, title, author(s), publish/modified date, tags and images from any news or blog URL — one row per URL, as Markdown, plain text or HTML. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | Article URLs — Enter the article or blog post URLs to extract, one row is returned per URL, e.g. https://blog.apify.com/best-web-scraping-tools/. Works on news sites, blogs and any page that publishes a Schema.org Article/NewsArticle JSON-LD block, Open Graph tags, or a plain readable body. Example: ["https://blog.apify.com/best-web-scraping-tools/"]. | |
| outputFormat | No | Output format — Choose which body format(s) to include in each row. "Markdown" is the smallest and best for LLM/RAG ingestion; "all" returns markdown, text and html together for debugging or comparison. Options: markdown = Markdown; text = Plain text; html = HTML; all = All (markdown + text + html). | markdown |
| includeImages | No | Include images — Keep this on to return the mainImage and images fields and keep image references in the markdown/html body. Turn it off for a smaller, text-only dataset. |