Skip to main content
Glama

browserless_smartscraper

Read-only

Extract content from a single JavaScript-heavy webpage, returning markdown, HTML, raw text, links, screenshots, or PDFs with metadata. Anti-bot measures are handled automatically so you get clean, structured page data.

Instructions

Scrape a SINGLE webpage and return HTML, markdown, raw DOM text, links, screenshots, or PDFs plus page metadata. Handles JavaScript-heavy pages and anti-bot measures automatically. For content across MULTIPLE pages of a site, use browserless_crawl; to list a site's URLs, use browserless_map.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape (must be http or https)
_promptNoThe end user's original, verbatim request that led to this tool call, if known. Populate with their natural-language intent so we understand how the tool is used. Do NOT include secrets, passwords, API keys, tokens, or other credentials. Omit if unavailable.
formatsNoOutput formats to include: "markdown", "html", "rawText", "screenshot", "pdf", "links". rawText is DOM text with script, style, and noscript elements removed and whitespace collapsed, or extracted text for PDF targets. Defaults to ["markdown"].
headersNoCustom HTTP headers sent to the target site. host, authorization, proxy-authorization, cookie, set-cookie, x-forwarded-for, x-real-ip, and forwarded are removed by the API.
profileNoOptional name of an authentication profile to hydrate into the browser before scraping. The profile's cookies, localStorage, and IndexedDB are restored into the session before the request runs. The profile must already exist for the API token in use — create one with Browserless.saveProfile in a live agent session first.
timeoutNoRequest timeout in milliseconds
waitForNoMilliseconds to wait after page load, from 0 to 30000. A positive value forces browser rendering.
excludeTagsNoUp to 100 CSS selectors to remove from HTML webpage outputs. Malformed selectors are ignored. Cannot be combined with includeTags.
includeTagsNoUp to 100 CSS selectors to keep in HTML webpage outputs. Malformed entries are ignored; if no selector matches, the scraper returns unfiltered content. Cannot be combined with excludeTags or onlyMainContent.
onlyMainContentNoFor HTML webpages, remove nav, footer, aside, role=navigation, script, style, and noscript elements from DOM-derived outputs. Parsed JSON and PDF content are unchanged. Defaults to false.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed8 schema fields changedv1.28.1
    • addedInput schema / properties / excludeTags
      Added value: +{
      +  "description": "Up to 100 CSS selectors to remove from HTML webpage outputs. Malformed selectors are ignored. Cannot be combined with includeTags.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "maxItems": 100,
      +  "type": "array"
      +}
    • changedInput schema / properties / formats / description
      Previous value: -"Output formats to include: \"markdown\", \"html\", \"screenshot\", \"pdf\", \"links\". Defaults to [\"markdown\"]."New value: +"Output formats to include: \"markdown\", \"html\", \"rawText\", \"screenshot\", \"pdf\", \"links\". rawText is DOM text with script, style, and noscript elements removed and whitespace collapsed, or extracted text for PDF targets. Defaults to [\"markdown\"]."
    • changedInput schema / properties / formats / items / enum
      Previous value: -[
      -  "markdown",
      -  "html",
      -  "screenshot",
      -  "pdf",
      -  "links"
      -]New value: +[
      +  "markdown",
      +  "html",
      +  "rawText",
      +  "screenshot",
      +  "pdf",
      +  "links"
      +]
    • addedInput schema / properties / formats / minItems
      Added value: +1
    • addedInput schema / properties / headers
      Added value: +{
      +  "additionalProperties": false,
      +  "description": "Custom HTTP headers sent to the target site. host, authorization, proxy-authorization, cookie, set-cookie, x-forwarded-for, x-real-ip, and forwarded are removed by the API.",
      +  "propertyNames": {
      +    "type": "string"
      +  },
      +  "type": "object"
      +}
    • addedInput schema / properties / includeTags
      Added value: +{
      +  "description": "Up to 100 CSS selectors to keep in HTML webpage outputs. Malformed entries are ignored; if no selector matches, the scraper returns unfiltered content. Cannot be combined with excludeTags or onlyMainContent.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "maxItems": 100,
      +  "type": "array"
      +}
    • addedInput schema / properties / onlyMainContent
      Added value: +{
      +  "default": false,
      +  "description": "For HTML webpages, remove nav, footer, aside, role=navigation, script, style, and noscript elements from DOM-derived outputs. Parsed JSON and PDF content are unchanged. Defaults to false.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / waitFor
      Added value: +{
      +  "description": "Milliseconds to wait after page load, from 0 to 30000. A positive value forces browser rendering.",
      +  "maximum": 30000,
      +  "minimum": 0,
      +  "type": "integer"
      +}
  2. First observedv1.16.0

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond those annotations by stating that it handles JavaScript-heavy pages and anti-bot measures automatically, and by stressing the single-page scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The most important constraint ('SINGLE webpage') is front-loaded, output options are listed compactly, and the sibling-tool routing is placed cleanly at the end. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich parameter schema and the annotations, the description covers purpose, output types, behavioral capabilities, and alternative tools. There is no output schema, but the description names return formats and metadata, which is sufficient for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully documents all 10 parameters. The description adds only a high-level mention of output formats, which is fine, but it doesn't go beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scrape'), a clear resource ('a SINGLE webpage'), and enumerates the exact output types (HTML, markdown, raw DOM text, links, screenshots, PDFs, metadata). It also explicitly distinguishes itself from browserless_crawl and browserless_map, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: use this tool for a single page, browserless_crawl for multiple pages, and browserless_map for listing URLs. This directly tells an agent when this tool is appropriate and names the alternatives, leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.