Get Structured Site Data
webscraping_ai_dataExtract structured JSON from public YouTube, TikTok, X, LinkedIn, Instagram, or Reddit URLs by auto-detecting the site and page type for parsed data.
Instructions
Get structured JSON for a public page on a supported site from its normal URL, e.g. a YouTube video, channel or playlist, a TikTok video or profile, an X post or profile, a LinkedIn company, job or profile, an Instagram post, reel or profile, or a Reddit post, subreddit or user. Returns {request_parameters: {url, provider, type}, parse_status, data}: provider (site) and type (page kind) are detected from the URL, parse_status is ok, parse_failed or not_found, and data holds snake_case fields whose shape depends on provider and type (null fields for values the page doesn't expose; data itself can be null when parsing fails). More sites are added on the server over time: an unsupported URL or page type returns a 400 error, not charged, whose message lists what is supported. For other sites, use webscraping_ai_fields. Priced per site (see https://webscraping.ai/docs#data), including parse_failed and not_found results; failed fetches are not charged.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Normal URL of a public page on a supported site, e.g. https://www.youtube.com/watch?v=dQw4w9WgXcQ. Sent as-is; the site and page type are detected from it. | |
| params | No | Extra site-specific query parameters sent to the API as-is, as an object of string, number or boolean values, for parameters added after this tool was released. Must not contain url, api_key or the parameters above. | |
| country | No | Two-letter country code of the proxy used to fetch the page, e.g. us, gb, de (us by default). | |
| transcript | No | YouTube videos only. Also fetch the video's transcript into data.transcript (null when no matching captions are available; false by default). If the transcript fetch fails, the whole request fails with a 500 and is not charged. | |
| transcript_language | No | YouTube videos only, with transcript: true. Caption language to pick, e.g. en, de. Without it, English is preferred, then the first available track; if the video has no captions in that language, data.transcript is null. | |
| disable_content_sandboxing | No | Return the raw result without the external-content security boundaries that guard against prompt injection (false by default). |