Extract one page in full
extract_urlFetch one URL exactly as a project with the given settings would, and return the whole page: the bodies config's formats ask for plus the raw html (each capped at 12,000 characters), head and extracted fields, its images, and which engine read it -- without its link list. Use it for a single page that needs browser steps, structured fields or a settings trial before create_project; for plain content of one or many known URLs use scrape_urls, and get_page reads a page a project already stored without fetching. It climbs the fetch ladder -- plain http first, a browser only when the page needs one or render_js is 'always' -- and costs the credits of the rung that read it (1 to 4 for most pages, +1 when formats ask for a screenshot). Browser actions in config (click, type, select, press, wait, scroll; repeat 'until_gone' for Load-more buttons; each to act on every match) run before the page is read.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The absolute http(s) URL of the page to read. | |
| config | No | Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults. | |
| parse_documents | No | Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins. |