extract_from_browser
Extract structured data from a web page using Zyte API automatic (AI) extraction, fetching the page with a browser: give it a URL and an extraction type, get typed JSON back (product, article, job posting, ...). Highest quality and the default choice; the only extract tool with browser actions and viewport. Omit extractFrom to let Zyte API choose the browser source (currently 'browserHtml'); 'browserHtmlOnly' skips screenshot signals. For server-rendered pages extract_from_http is faster and cheaper. Choose the type by page kind — detail types for a single entity's page, list types for per-item summaries from one page, navigation types when crawling, pageContent as the generic fallback, webPageInfo for language only. One type per call. product, article, jobPosting and pageContent results carry metadata.probability — the confidence the page matches the type (below ~0.5 means it might not; a signal, not an error). If unsure between two of those types, make one call per type and keep the higher-probability result; list/navigation/forumThread results have no page-level probability. customAttributes adds caller-defined LLM-extracted fields on top of the requested type; for custom attributes alone use type "pageContent". Supports sessions, cookies, geolocation, and IP type. For raw pages (HTML, screenshots, files) use fetch_page or fetch_http instead. Returns a JSON metadata block (final URL, target HTTP status, optional action results/session), then one JSON block with the extracted data keyed by type.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute http(s) URL of the page to extract from (max 8192 chars). The host must be a domain name, not an IP address. | |
| type | Yes | The extraction type to run — one per call. Detail types (product, article, jobPosting, forumThread) suit a page about a single entity; list types (productList, articleList) return per-item summaries from a listing page; navigation types (productNavigation, articleNavigation, jobPostingNavigation) return item/next-page links for crawling; pageContent separates main content from site boilerplate on any page; webPageInfo returns the page language. | |
| ipType | No | Type of IP address to send the request from. Default: Zyte API picks the type that avoids bans for the target site. | |
| actions | No | Browser actions executed after the page loads and before extraction: click, type, scrollBottom, waitForSelector, evaluate, etc. Per-action results are returned in the metadata block. See the Zyte API actions reference (https://docs.zyte.com/zyte-api/usage/reference.html) for behavior details. | |
| viewport | No | Browser viewport size; changes what the extractor sees. | |
| sessionId | No | Client-managed session ID — a version 4 UUID you generate. Requests with the same ID reuse the same session (IP, cookies). Sessions expire 15 minutes after creation, after 2 idle minutes, or after 3 bans. | |
| extractFrom | No | Which browser source the extractor reads. 'browserHtml': rendered HTML plus screenshot signals, highest quality but less robust to rendering issues. 'browserHtmlOnly': rendered HTML without screenshot signals. Omit the field entirely to let Zyte API choose (currently 'browserHtml'). For a raw HTTP fetch use extract_from_http. | |
| geolocation | No | ISO 3166-1 alpha-2 country code to route the request from (e.g. 'US', 'DE'); must be a real country code. Default: Zyte API picks a geolocation that avoids bans and locale surprises for the target site. | |
| organizationId | Yes | Required. The Zyte organization to attribute this call to (max 100 characters, printable ASCII without spaces). If you do not already have an id, call the user_info tool: it lists the organizations your credential belongs to. IMPORTANT: if it lists more than one, ask the user which to use and wait for their answer — this call is billed to whichever organization you name here, so it is the user's choice to make, not yours. Never guess an id, and never fall back to a default. Once the user has chosen, reuse that id across the session unless they ask for a different organization. | |
| requestCookies | No | Cookies to send with the request (max 100). The responseCookies output of a previous fetch_page/fetch_http call can be passed here verbatim. | |
| sessionContext | No | Server-managed session context: up to 10 name/value pairs. Zyte API reuses or creates a session per distinct context. Sessions expire after 4 hours or 3 bans. | |
| enableZeroTrace | No | Keep the URL and other potentially sensitive request data out of Zyte API's request logs, metrics and stats records for this request. Default false. Use it for sensitive targets; it also leaves Zyte support less to go on when investigating the request. | |
| cookieManagement | No | How cookies are handled: 'auto' (default) uses requestCookies if given, otherwise Zyte API's automatic cookies; 'discard' uses requestCookies if given, otherwise no cookies. | |
| customAttributes | No | Ad-hoc fields extracted by a Zyte-operated LLM on top of the requested type (max 20 attributes): attribute name to attribute schema. The requested type scopes the page region fed to the LLM. If you only want custom attributes, use type "pageContent". | |
| sessionContextActions | No | Browser actions run once to initialize a server-managed session for the given sessionContext (e.g. login steps). |