Skip to main content
Glama
Prog-up

Web Scraper MCP

by Prog-up

Web Scraper MCP

A self-hosted Model Context Protocol server that gives an LLM client (Claude Code, Cursor, ChatGPT) the same tool surface as paid scraping services — scrape, crawl, map, search, extract, interact, deep_research — running entirely on your own hardware.

No paid proxy/CAPTCHA services: anti-bot is self-hosted (headless Chromium + playwright-stealth, robots.txt, polite rate limiting). Hardened sites may still block; see Limitations.

Tools

Tool

What it does

scrape

One URL → clean markdown (boilerplate stripped). Static-first, auto browser fallback for JS pages.

crawl / check_crawl_status

Background BFS crawl job (dedup, depth/page caps); poll for results.

map

List the links on a page (optionally same-domain) — decide what to crawl.

search

Web search. Pluggable backend: DuckDuckGo (default), SearXNG, Brave, or Tavily.

extract

Fetch a page and pull structured JSON matching your schema, via an LLM.

browser_navigate / browser_act / browser_close

Drive a persistent browser session (click/fill/press) using token-cheap ARIA snapshots.

deep_research

Search → read top sources → return a cited synthesis report.

extract and deep_research require ANTHROPIC_API_KEY.

Related MCP server: Universal Web Data Extraction Platform

Quick start

uv sync                      # install deps (uses the pinned uv.lock)
uv run playwright install chromium
export SCRAPER_AUTH_TOKEN=$(openssl rand -hex 32)
uv run web-scraper-mcp       # HTTP server on http://127.0.0.1:8000/mcp

Stdio (local, for a desktop client): SCRAPER_TRANSPORT=stdio uv run web-scraper-mcp.

Docker

The fastest way to get started is by pulling the pre-built image from DockerHub.

# Pull the latest image
docker pull PROG_UP_USERNAME/web-scraper-mcp:latest

# Run the container (with Anthropic / Claude)
docker run -p 8000:8000 \
  -e SCRAPER_AUTH_TOKEN=your_secure_token_here \
  -e ANTHROPIC_API_KEY=sk-ant-api03-... \
  PROG_UP_USERNAME/web-scraper-mcp:latest

# Or run the container over stdio (useful for local MCP clients)
docker run -i --rm \
  -e SCRAPER_TRANSPORT=stdio \
  -e SCRAPER_AUTH_TOKEN=your_secure_token_here \
  PROG_UP_USERNAME/web-scraper-mcp:latest

(Make sure to replace PROG_UP_USERNAME with your actual DockerHub username).

Using Local Models (Ollama) & Context Windows

If you prefer to run models locally instead of using Anthropic's API, the server fully supports Ollama as an alternative backend for the extract and deep_research tools.

docker run -p 8000:8000 \
  -e SCRAPER_AUTH_TOKEN=your_secure_token_here \
  -e SCRAPER_OLLAMA_HOST=http://host.docker.internal:11434 \
  -e SCRAPER_EXTRACT_MODEL=qwen3.5:2b \
  -e SCRAPER_RESEARCH_MODEL=qwen3.5:2b \
  PROG_UP_USERNAME/web-scraper-mcp:latest
WARNING

Context Windows are Critical! Web scraping produces a massive amount of Markdown. extract can send up to 100,000 characters and deep_research can send up to 16,000 characters to the LLM.

By default, Claude handles massive contexts natively. However, Ollama's default context window (num_ctx) is often configured to just 2,048 tokens. If you pass an enormous Wikipedia page to a local model, it will silently truncate the prompt (dropping your instructions) and return empty strings!

We automatically pass "num_ctx": 32768 to Ollama in the API payloads to prevent this truncation. Ensure that your local machine has enough RAM/VRAM to support a 32K context window when using Ollama!

Register in a client (mcp.json)

If running via HTTP:

{
  "mcpServers": {
    "web-scraper": {
      "url": "http://127.0.0.1:8000/mcp",
      "headers": { "Authorization": "Bearer your_secure_token_here" }
    }
  }
}

If running via stdio (Docker):

{
  "mcpServers": {
    "web-scraper": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "SCRAPER_TRANSPORT=stdio",
        "-e", "SCRAPER_OLLAMA_HOST=http://host.docker.internal:11434",
        "-e", "SCRAPER_EXTRACT_MODEL=qwen3.5:2b",
        "-e", "SCRAPER_RESEARCH_MODEL=qwen3.5:2b",
        "PROG_UP_USERNAME/web-scraper-mcp:latest"
      ]
    }
  }
}

Configuration

All settings are env vars (prefix SCRAPER_), or a .env file — see .env.example.

Var

Default

Notes

SCRAPER_AUTH_TOKEN

(unset)

Bearer token for the HTTP endpoint. Required for any networked deploy; unset = unauthenticated (warned).

SCRAPER_HOST / SCRAPER_PORT

127.0.0.1 / 8000

Bind address. Docker image sets host 0.0.0.0.

SCRAPER_MAX_CONCURRENT_PAGES

8

Headless-page concurrency cap (RAM/CPU bound).

SCRAPER_MAX_CRAWL_PAGES / _DEPTH

100 / 3

Hard ceilings for crawl jobs.

SCRAPER_PER_DOMAIN_DELAY_S

1.0

Polite per-domain rate limit.

SCRAPER_RESPECT_ROBOTS

true

Honour robots.txt.

SCRAPER_ALLOW_PRIVATE_NETWORKS

false

Keep false — disables the SSRF guard if true.

ANTHROPIC_API_KEY

(unset)

Enables extract / deep_research.

SCRAPER_SEARXNG_URL, BRAVE_API_KEY, TAVILY_API_KEY

(unset)

Optional search backends (first set wins, else DuckDuckGo).

Security

  • SSRF guard — every fetched URL (and each redirect hop) is DNS-resolved and rejected if it points at a private / loopback / link-local / cloud-metadata address. The browser also aborts subresource requests to private IPs.

  • Auth — bearer token on the HTTP transport; bind localhost by default.

  • Resource caps — response-size, timeout, page-concurrency, crawl page/depth limits to protect the host.

  • robots.txt + rate limiting on by default.

  • Container — runs as a non-root user; secrets via env only.

Supply chain — verifying the image

CI signs the image keylessly with cosign (Sigstore) using GitLab's OIDC identity, and attaches an SPDX SBOM attestation. Verify before running:

cosign verify \
  --certificate-oidc-issuer https://gitlab.cri.epita.fr \
  --certificate-identity-regexp 'https://gitlab.cri.epita.fr/enzo.juhel/web-scraper//.*' \
  registry.gitlab.cri.epita.fr/enzo.juhel/web-scraper@sha256:...

Benchmark

benchmarks/run.py scores our scrape/extract against public datasets and the Crawl4AI baseline, emitting a scorecard (per page-type F1/accuracy + a Limitations section). Run locally or via the manual CI benchmark job:

uv run python benchmarks/run.py --output scorecard.md

Limitations

  • No paid proxies/CAPTCHA: heavily-defended sites (LinkedIn, Amazon, Cloudflare challenges) will sometimes block us. The benchmark scorecard quantifies where.

  • Main-content extraction is strong on articles, weaker on forums / product / listing pages (a known property of all extractors).

  • In-memory crawl/session state — single-process, single-user by design.

Development

uv run pre-commit install        # local lint/secret hooks
uv run ruff check . && uv run ruff format --check .
uv run mypy src
uv run pytest -q
A
license - permissive license
A
quality
B
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables web scraping and crawling capabilities for LLM clients, supporting single-page scraping, multi-page website crawling, and web search with multiple engines (Playwright, Cheerio, Puppeteer) and flexible output formats including markdown, HTML, text, and screenshots.
    18
    6
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables LLMs to extract content from websites using automated static and dynamic scraping engines with built-in anti-bot protections. It provides tools for web data retrieval and stores results in MongoDB with support for JSON and CSV exports.
  • A
    license
    A
    quality
    A
    maintenance
    Web scraping, crawling, and structured data extraction for AI agents. 5 tools: scrape (clean markdown from any URL), crawl (entire sites), map (discover URLs), extract (structured JSON), and search. 833ms avg latency, single binary, self-hostable.
    8
    852
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables LLMs to fetch and extract web content using browser automation, OCR, and multiple extraction methods, handling JavaScript rendering and anti-scraping techniques.
    17
    MIT

View all related MCP servers

Related MCP Connectors

  • Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabili…

  • Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents.

  • Live web access for agents: scrape, SERP search, crawl/map, 74 collectors, datasets, proxies.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Prog-up/web-scraper-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server