visit_with_browser
Visit any website with browser automation to extract clean page content as Markdown, HTML, Readability, or cleaned HTML, plus screenshots and PDFs; handles JavaScript.
Instructions
Visit any website using full browser automation (stealth mode, anti-detection). Returns page content in your chosen format: 'html' for raw HTML source, 'markdown' for clean formatted text (recommended for reading), 'readability' for Mozilla Readability format, or 'cleaned_html' for cleaned HTML. Supports screenshot and PDF generation. Automatically handles JavaScript rendering and provides clean output by default.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| No | Generate a PDF of the page (returns base64 encoded PDF) | ||
| url | Yes | The complete URL to scrape (must include http:// or https://) | |
| delay | No | Delay in seconds to wait after page load before scraping | |
| format | No | Content formats to extract: 'html'=raw HTML source (may be very large), 'markdown'=clean formatted text converted from HTML (recommended for reading), 'readability'=Mozilla Readability format, 'cleaned_html'=cleaned HTML. You can request multiple formats. | |
| logUrl | No | URL to send logs to for debugging purposes | |
| proxyUrl | No | Proxy URL to use for the request (e.g., 'http://proxy:port') | |
| maxLength | No | Maximum characters to return (optional). Smart defaults: markdown=8000, readability=10000, html=15000, cleaned_html=12000. For markdown, automatically reserves space for metadata. | |
| screenshot | No | Take a screenshot of the page (returns base64 encoded image) | |
| verboseMode | No | Return full metadata instead of clean content-focused output (optional, default: false). Use when you need detailed scraping information. |