structured_data_extractor
Extract Schema.org JSON-LD and Open Graph tags from any URL, returning one clean row per page for RAG pipelines and SEO audits.
Instructions
Structured Data & JSON-LD Extractor reads every Schema.org JSON-LD block and Open Graph tag on a page and returns one clean row per URL — product, price, rating, article, job posting, event, recipe, FAQ and breadcrumb data, ready for RAG pipelines and SEO rich-result audits. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs — Enter the page URLs to extract structured data from, one row is returned per URL, e.g. https://www.allbirds.com/products/mens-tree-dashers. Works on product, article, job, event, recipe and FAQ pages — anything that publishes Schema.org JSON-LD or Open Graph tags. Example: ["https://www.apify.com"]. | |
| followRedirects | No | Follow redirects — Keep this on to follow HTTP redirects to the final page, e.g. a shortened or tracking URL. Turn it off to get an error row instead when a URL 301/302s, useful for auditing which URLs redirect. | |
| includeOpenGraph | No | Include Open Graph / Twitter Card data — Keep this on to return the openGraph field (og:title, og:description, og:image, og:type, og:site_name, twitter:card and related tags). Most pages publish these even without JSON-LD. | |
| includeRawJsonLd | No | Include raw JSON-LD — Keep this on to also return the page's raw parsed JSON-LD documents in the jsonLd field (capped at 400 KB per row), useful when you need a schema type the mapped fields do not cover. Turn it off for a smaller, cheaper-to-store dataset. |