Skip to main content
Glama
crawlbrulee

@crawlbrulee/mcp

Official
by crawlbrulee

Scrape a URL

scrape

Fetches a single URL and returns page content as markdown, cleaned HTML, raw HTML, links, images, screenshot, or metadata.

Instructions

Fetch a single URL via the crawlbrulee scraping API and return the requested content (markdown, cleaned HTML, raw HTML, links, images, screenshot, page metadata in metadata). Use this for one-shot page extraction. For full-site discovery use the map tool first. Screenshot URLs in the response are signed download links — the agent can fetch them when needed. The response also carries response_meta.usage = { credits, engine, proxy (the resolved tier — never auto), screenshot_slices }. engine is the billed engine (http, browser, screenshot, cache); a cache hit is represented by engine: "cache", so you can see what the request cost. Extraction is capped per page: 30,000 links, 10,000 inline images and 10,000,000 characters of HTML. A page past a cap is truncated rather than refused, and the response warnings array names which one (links_truncated, inline_images_truncated, raw_html_truncated) — so treat that output as incomplete. Use cleanup to control what is removed before any output is built: ads_and_popups (on by default) drops ads, cookie banners and chat widgets, and exclude_selectors removes anything else. It shapes markdown, cleaned_html, links, images and the screenshot, and NEVER raw_html — so request raw_html when you need the page exactly as it arrived. warnings also reports a section whose extraction failed outright (links_unavailable, inline_images_unavailable, metadata_unavailable): that field comes back omitted or empty while the rest of the scrape succeeded, so do NOT conclude the page had no links/images/metadata — re-run the scrape instead. An empty field with no such warning does mean the page had none.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape. Known tracking parameters are removed before the page is fetched, so they are neither sent to the target site nor part of the cache key. Every other query parameter is kept verbatim and is part of the cache key.
cacheNoCache settings for this request
proxyNoProxy tier to use for fetchingauto
cleanupNoWhat is removed from the page before any output is built. Applies to markdown, cleaned_html, links and images on every engine, and to the screenshot. Never applies to raw_html, which is always the page before we removed anything.
extractNoWhich content formats to extract. Defaults to metadata + cleaned_html.
locationNoOptional locale + country emulation for the scrape
require_jsNoUse a headless browser to render JavaScript before scraping

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL that was actually scraped, after any redirects, in normalized form — the base that links, images, and internal labels are computed against
linksNoLinks found on the page
imagesNoInline images found on the page
markdownNoPage content converted to clean Markdown
metadataNoExtracted page metadata (title, OG tags, etc.)
raw_htmlNoRaw, unprocessed HTML of the page
warningsNoNon-error notices about the scrape. Truncation codes — `screenshot_truncated` (long page exceeded the scrolling-screenshot height cap), `links_truncated` / `inline_images_truncated` (page had more links/images than the per-page extraction caps), `raw_html_truncated` / `metadata_truncated` (rendered HTML exceeded the per-page size budget) — mean the field is present but capped. Unavailability codes — `links_unavailable` / `inline_images_unavailable` / `metadata_unavailable` — mean that optional field could not be extracted and was omitted (null/empty) while the rest of the scrape succeeded, so an empty field carrying one of these does NOT mean the page had none. Stable string codes — clients can switch on them. Warnings are stored with the result: async result fetches and cache hits carry them too, filtered to the fields the request asked for.
screenshotNoScreenshot of the page, if requested
cleaned_htmlNoCleaned HTML of the main page content
content_typeNoContent-Type header returned by the server
requested_urlYesThe URL you requested, echoed verbatim — before any redirects
response_metaYesRequest-level metadata. `response_meta.usage` reports credits charged, the billed engine, the resolved proxy tier, and any screenshot-slice add-on.
unsupported_fieldsNoExtract fields that were requested but are not supported for this content type

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.3

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description fully carries the behavioral burden and does so exceptionally: it discloses cache-hit behavior via engine:'cache', signed screenshot URLs, per-page truncation caps and their warnings, partial-failure semantics (links_unavailable etc.), and the meaning of empty fields. This is far beyond any structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool is complex (7 parameters, nested screenshot/cleanup/cache objects) and the prose is dense with non-redundant operational details. It is front-loaded with the purpose statement and only sacrifices brevity where the behavior of the API genuinely needs explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and the complexity of the tool, the description is complete: it covers selection, output formats, cleanup effects, limits, warnings, cache semantics, and response metadata. An agent has enough context to invoke the tool and interpret partial results correctly; the output schema covers the rest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema texts are already detailed, so the baseline is 3. The description adds genuine value by clarifying that cleanup shapes all outputs except raw_html, that a cache hit is distinguishable in response_meta, and that truncated output must be treated as incomplete (warnings list). It does not need to re-document every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Fetch a single URL') and lists every output format it can return. It explicitly distinguishes itself from the `map` sibling ('For full-site discovery use the map tool first') and frames itself as one-shot page extraction, so an agent cannot confuse it with the crawl tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states exactly when to use this tool ('one-shot page extraction') and steers the agent to `map` for full-site discovery. Additional operational guidance (request raw_html when the untouched page is needed, use cleanup for ad/popup removal) helps the agent choose output modes correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.