Extract web page content from specified URLs using Tavily Extract.
post_tavily_extractFetch clean, parsed content for URLs you already have — from a search result, a sitemap, or a user. urls is required and takes several at once. Returns results[] with url, title, raw_content and images, plus a failed_results[] array — read that one, because a page that could not be fetched is reported there rather than raising an error. Choose format (markdown or text) and extract_depth. Measured at about 1 second for one page. Use this instead of post_tavily_search whenever you can already name the pages; searching for pages you can name costs more and may not return them. For a long list that can wait, post_firecrawl_batch_scrape runs it as a background job. To discover the URLs of a whole site first, use post_tavily_map.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | The URL or URLs to extract content from. A single URL string and an array of URL strings are both accepted. | |
| query | No | User intent for reranking extracted content chunks. | |
| format | No | Format of the extracted web page content. | markdown |
| timeout | No | Maximum time in seconds to wait for URL extraction. If omitted, default timeouts depend on extract_depth: 10 seconds for basic and 30 seconds for advanced. | |
| extract_depth | No | Depth of the extraction process. | basic |
| include_usage | No | Include credit usage information in the response. | |
| include_images | No | Include a list of images extracted from the URLs. | |
| include_favicon | No | Include the favicon URL for each result. | |
| chunks_per_source | No | Maximum number of relevant chunks returned per source. Available only when query is provided. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||