Skip to main content
Glama
Prog-up

Web Scraper MCP

by Prog-up

Web Scraper MCP

A self-hosted Model Context Protocol server for reading public web pages, bounded crawling, search, model-backed extraction and browser interaction. Supports Python 3.11–3.13 and stdio or authenticated HTTP.

Run it with your own models and infrastructure. Fetching is static by default; browser rendering is optional and uses sandboxed Chromium with guarded egress.

Tools

Tool

Behavior

scrape

Fetch a public URL and return its title, Markdown and optional capped links or HTML.

crawl

Start a background crawl with page/depth limits and an optional same-host policy.

check_crawl_status

Read progress/results or cancel a job while retaining completed pages.

map

Return canonical, deduplicated links; defaults to the source host.

browser_navigate

Create or reuse a browser session and return an ARIA snapshot.

browser_act

Click, fill, press, select or wait using a selector. Explicit actions may submit same-origin forms.

browser_close

Close a session and release its page capacity.

search

Use Tavily, Brave, SearXNG or DuckDuckGo, in that priority order.

extract

Ask Ollama or Anthropic for text or data matching a locally validated JSON Schema.

deep_research

Search, read bounded sources and synthesize a report with checked citation indices and source failures.

Related MCP server: Universal Web Data Extraction Platform

Quick start

Install Python 3.11–3.13 and uv, then:

git clone https://github.com/Prog-up/web-scraper-mcp.git
cd web-scraper-mcp
uv sync --frozen
cp .env.example .env

For a local MCP client using stdio:

SCRAPER_TRANSPORT=stdio uv run --frozen web-scraper-mcp

Stdio does not require a token. Clients with an mcpServers configuration can use:

{
  "mcpServers": {
    "web-scraper": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/web-scraper-mcp", "run", "--frozen", "web-scraper-mcp"],
      "env": {"SCRAPER_TRANSPORT": "stdio"}
    }
  }
}

For HTTP:

export SCRAPER_AUTH_TOKEN=$(openssl rand -hex 32)
./start.sh

Connect to http://127.0.0.1:8000/mcp with Authorization: Bearer YOUR_TOKEN. Keep the token in your client's secret configuration. HTTP startup rejects missing or blank tokens. The launcher stays in the foreground; use a service manager for supervision.

Scraping and browser tools do not require an LLM. For extract and deep_research, configure an installed Ollama model and reachable endpoint in .env, or select Anthropic and supply its API key. Search defaults to DuckDuckGo; other providers are optional. See configuration and the environment template.

Docker and browser tools

For sandboxed rendering and browser sessions:

export SCRAPER_AUTH_TOKEN=$(openssl rand -hex 32)
docker compose up --build

Compose runs the app on an internal network with a separate guarded egress proxy, a read-only filesystem, resource limits and Chromium's sandbox enabled. Use Docker Engine 28 or newer and a host that supports unprivileged user namespaces. Sandbox startup failures stop rendering; there is no automatic fallback that disables the sandbox.

Compose does not start Ollama. Its model endpoint must be reachable from the isolated app network; an external endpoint needs a trusted route or relay. See deployment for static-only containers, networking, model connectivity and browser restrictions, and seccomp provenance.

Scope and limitations

  • Intended for public HTTP(S) pages. Private destinations are denied by default. Error status pages and unsupported content types are rejected.

  • Protected websites, CAPTCHAs and some JavaScript applications may fail. There is no managed proxy pool or guaranteed challenge bypass.

  • Browser sessions block downloads, WebSockets, service workers and background writes. Explicit actions allow writes only to the current page origin.

  • Crawl jobs and browser sessions live in memory. They expire and are lost on restart; this is not a durable distributed job system.

  • One HTTP bearer token grants all tools. There are no per-user quotas or tenant isolation; use TLS and trusted access boundaries beyond localhost.

  • JSON Schema validation checks output structure. Citation validation checks source indices, not whether a source supports every claim. Review important extracted facts against the original sources.

  • Self-hosting still requires maintenance and compute. Configured search/model providers receive their requests; provider fees and policies apply.

Development and support

See CONTRIBUTING.md for setup and checks, benchmarks for reproducible measurements, and SECURITY.md for vulnerability reporting and security boundaries.

CI checks Python 3.11–3.13, lint, formatting, typing, tests, secrets and locked dependencies. Container releases also test the exact sandboxed browser artifact, scan vulnerabilities, and sign the published digest with an SPDX SBOM. See release procedures, release history and the known security exceptions for 1.1.0 and 1.1.1.

Report bugs through GitHub issues. Support is community-based, with no service-level agreement.

License

See LICENSE for the existing GNU GPL version 3 terms.

Available Tools

4 tools
deep_researchA

Search the web, read the top sources, and return a cited synthesis report.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe research question.
max_sourcesNoHow many top results to read.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses the general workflow (searching the web, reading top sources, producing a cited report) but does not mention read-only behavior, potential latency, rate limits, or any limitations. The description is truthful and somewhat informative, but leaves important operational behavior unaddressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with a clear chronological structure, no filler, and immediate front-loading of the core action ('Search the web'). Every phrase earns its place: 'read the top sources' and 'cited synthesis report' are distinct and necessary components. This is appropriately concise for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two fully documented parameters and the presence of an output schema (likely covering the report structure), the description covers the essential process and outcome well. What is missing is more explicit context about when to use the tool versus siblings and a few behavioral notes, but the combination of description and schema is largely sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: both 'query' and 'max_sources' have descriptions in the schema. The tool description's phrasing ('research question', 'top results') mirrors the schema descriptions without adding new meaning about formats, constraints, or interpretation. This meets the baseline of 3 but does not exceed what structured data already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb sequence ('Search', 'read', 'return') with a clear deliverable ('a cited synthesis report'). It unambiguously identifies the tool's function and naturally distinguishes deep_research from siblings like search (just web results) or scrape/extract (page-level content gathering). The multi-step process is explicit and not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for comprehensive research with cited synthesis, but it never explicitly states when to prefer this over search, scrape, or extract, nor does it give any 'instead of' guidance. Context makes the intended use reasonably clear, yet there are no exclusions or alternative routing as seen in higher-quality definitions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extractB

Extract structured data (JSON matching json_schema) or a text answer from a page.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to extract from.
promptNoNatural-language extraction instruction.
renderNoForce a browser render.
json_schemaNoJSON Schema describing the fields to extract (recommended).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It does disclose the two operating modes (structured JSON versus text answer), which is meaningful. However, it does not mention rendering defaults, page-fetch behavior, rate limits, or side effects. The render parameter is defined in the schema, so some of that context is available, but the description itself stays minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It could be slightly better structured by explicitly separating the two modes, but as written it is efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return-value details are covered elsewhere, and the input schema is fully documented. The main gaps are the lack of choice guidance between prompt and json_schema, and no indication of when to use extract instead of scrape. These gaps make the description adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already defines url, prompt, render, and json_schema. The description adds a high-level mapping by linking json_schema to structured JSON and implying the text answer comes from the prompt, but it does not add much beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('extract'), the target resource ('a page'), and the two possible outputs: JSON matching json_schema or a text answer. It does not explicitly distinguish itself from sibling tools like scrape or search, though the verb and output specification make the core purpose clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use extract versus scrape, search, or deep_research, and no exclusions or conditions. The only context is 'from a page,' which is too thin to help an agent choose between this and its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrapeA

Scrape a single URL into clean markdown (boilerplate/ads stripped).

Tries a fast static fetch first and falls back to a stealth headless browser automatically when the page looks JS-gated.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape (http/https).
renderNoForce a headless browser render (for JS-heavy pages).
include_linksNoAlso return all links found on the page.
include_raw_htmlNoAlso return the raw HTML (large).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It clearly reveals the two-phase fetch strategy, the automatic fallback to headless browsing, and the output transformation to clean markdown. It does not discuss rate limits, authentication, or failure behavior, but the core operational traits are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no fluff, front-loading the primary purpose before the fallback behavior. Every sentence contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema means return values do not need to be explained in the description. Combined with fully documented parameters and a clear explanation of the scraping strategy, the definition is largely complete, though it omits failure semantics and any rate-limit or authentication caveats.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the input schema. The tool description adds no additional parameter-level meaning, which matches the baseline for fully-covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scrape a single URL') and a concrete outcome ('into clean markdown'), making the tool's purpose immediately clear. The boilerplate/ads-stripping detail further distinguishes it from general fetch or research siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains that a fast static fetch is attempted first and that a stealth headless browser is used automatically for JS-gated pages, giving clear context on when the tool adapts. It does not explicitly name alternatives or say when to prefer sibling tools, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observeddeep_research
    • First observedextract
    • First observedscrape
    • First observedsearch

TDQS

A3.8/5.0

Scored across 4 tools

Disambiguation4/5

scrape and extract both operate on a page but are clearly distinguished by output type: markdown document vs structured JSON/answer. search and deep_research are also distinct, though deep_research could be seen as a superset of search.

Naming Consistency5/5

All tools use lowercase snake_case imperative verbs: scrape, extract, search, deep_research. The naming pattern is consistent and predictable.

Tool Count5/5

Four tools is well-scoped for a web scraper/research server: one for raw page content, one for structured extraction, one for search, and one for synthesis. Each tool has a clear purpose with no redundant bloat.

Completeness4/5

The core web research workflow is covered: search, fetch, extract, and synthesize. A minor gap is the lack of batch crawling or link-following features, but most common agent tasks can be completed with these tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables web scraping and crawling capabilities for LLM clients, supporting single-page scraping, multi-page website crawling, and web search with multiple engines (Playwright, Cheerio, Puppeteer) and flexible output formats including markdown, HTML, text, and screenshots.
    9 npm
    6
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables LLMs to extract content from websites using automated static and dynamic scraping engines with built-in anti-bot protections. It provides tools for web data retrieval and stores results in MongoDB with support for JSON and CSV exports.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Web scraping, crawling, and structured data extraction for AI agents. 5 tools: scrape (clean markdown from any URL), crawl (entire sites), map (discover URLs), extract (structured JSON), and search. 833ms avg latency, single binary, self-hostable.
    8
    1,105
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables LLMs to fetch and extract web content using browser automation, OCR, and multiple extraction methods, handling JavaScript rendering and anti-scraping techniques.
    17
    MIT