Web Scraper MCP
Optional web search backend for the search tool, configured via [BRAVE](/mcp/servers/integrations/brave)_API_KEY to use Brave Search.
Default web search backend for the search tool, providing privacy-focused search results without requiring additional configuration.
Alternative local LLM backend for the extract and deep_research tools, allowing use of self-hosted models (e.g., Qwen) via SCRAPER_[OLLAMA](/mcp/servers/integrations/ollama)_HOST and model environment variables.
Optional self-hosted web search backend for the search tool, configurable via SCRAPER_[SEARXNG](/mcp/servers/integrations/searxng)_URL to use a SearXNG instance.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Web Scraper MCPScrape https://example.com and summarize it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Web Scraper MCP
A self-hosted Model Context Protocol server that gives an LLM client (Claude Code, Cursor, ChatGPT) the same tool surface as paid scraping services — scrape, crawl, map, search, extract, interact, deep_research — running entirely on your own hardware.
No paid proxy/CAPTCHA services: anti-bot is self-hosted (headless Chromium + playwright-stealth, robots.txt, polite rate limiting). Hardened sites may still block; see Limitations.
Tools
Tool | What it does |
| One URL → clean markdown (boilerplate stripped). Static-first, auto browser fallback for JS pages. |
| Background BFS crawl job (dedup, depth/page caps); poll for results. |
| List the links on a page (optionally same-domain) — decide what to crawl. |
| Web search. Pluggable backend: DuckDuckGo (default), SearXNG, Brave, or Tavily. |
| Fetch a page and pull structured JSON matching your schema, via an LLM. |
| Drive a persistent browser session (click/fill/press) using token-cheap ARIA snapshots. |
| Search → read top sources → return a cited synthesis report. |
extract and deep_research require ANTHROPIC_API_KEY.
Related MCP server: Universal Web Data Extraction Platform
Quick start
uv sync # install deps (uses the pinned uv.lock)
uv run playwright install chromium
export SCRAPER_AUTH_TOKEN=$(openssl rand -hex 32)
uv run web-scraper-mcp # HTTP server on http://127.0.0.1:8000/mcpStdio (local, for a desktop client): SCRAPER_TRANSPORT=stdio uv run web-scraper-mcp.
Docker
The fastest way to get started is by pulling the pre-built image from DockerHub.
# Pull the latest image
docker pull PROG_UP_USERNAME/web-scraper-mcp:latest
# Run the container (with Anthropic / Claude)
docker run -p 8000:8000 \
-e SCRAPER_AUTH_TOKEN=your_secure_token_here \
-e ANTHROPIC_API_KEY=sk-ant-api03-... \
PROG_UP_USERNAME/web-scraper-mcp:latest
# Or run the container over stdio (useful for local MCP clients)
docker run -i --rm \
-e SCRAPER_TRANSPORT=stdio \
-e SCRAPER_AUTH_TOKEN=your_secure_token_here \
PROG_UP_USERNAME/web-scraper-mcp:latest(Make sure to replace PROG_UP_USERNAME with your actual DockerHub username).
Using Local Models (Ollama) & Context Windows
If you prefer to run models locally instead of using Anthropic's API, the server fully supports Ollama as an alternative backend for the extract and deep_research tools.
docker run -p 8000:8000 \
-e SCRAPER_AUTH_TOKEN=your_secure_token_here \
-e SCRAPER_OLLAMA_HOST=http://host.docker.internal:11434 \
-e SCRAPER_EXTRACT_MODEL=qwen3.5:2b \
-e SCRAPER_RESEARCH_MODEL=qwen3.5:2b \
PROG_UP_USERNAME/web-scraper-mcp:latestContext Windows are Critical!
Web scraping produces a massive amount of Markdown. extract can send up to 100,000 characters and deep_research can send up to 16,000 characters to the LLM.
By default, Claude handles massive contexts natively. However, Ollama's default context window (num_ctx) is often configured to just 2,048 tokens. If you pass an enormous Wikipedia page to a local model, it will silently truncate the prompt (dropping your instructions) and return empty strings!
We automatically pass "num_ctx": 32768 to Ollama in the API payloads to prevent this truncation. Ensure that your local machine has enough RAM/VRAM to support a 32K context window when using Ollama!
Register in a client (mcp.json)
If running via HTTP:
{
"mcpServers": {
"web-scraper": {
"url": "http://127.0.0.1:8000/mcp",
"headers": { "Authorization": "Bearer your_secure_token_here" }
}
}
}If running via stdio (Docker):
{
"mcpServers": {
"web-scraper": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "SCRAPER_TRANSPORT=stdio",
"-e", "SCRAPER_OLLAMA_HOST=http://host.docker.internal:11434",
"-e", "SCRAPER_EXTRACT_MODEL=qwen3.5:2b",
"-e", "SCRAPER_RESEARCH_MODEL=qwen3.5:2b",
"PROG_UP_USERNAME/web-scraper-mcp:latest"
]
}
}
}Configuration
All settings are env vars (prefix SCRAPER_), or a .env file — see
.env.example.
Var | Default | Notes |
| (unset) | Bearer token for the HTTP endpoint. Required for any networked deploy; unset = unauthenticated (warned). |
|
| Bind address. Docker image sets host |
|
| Headless-page concurrency cap (RAM/CPU bound). |
|
| Hard ceilings for crawl jobs. |
|
| Polite per-domain rate limit. |
|
| Honour robots.txt. |
|
| Keep false — disables the SSRF guard if true. |
| (unset) | Enables |
| (unset) | Optional search backends (first set wins, else DuckDuckGo). |
Security
SSRF guard — every fetched URL (and each redirect hop) is DNS-resolved and rejected if it points at a private / loopback / link-local / cloud-metadata address. The browser also aborts subresource requests to private IPs.
Auth — bearer token on the HTTP transport; bind localhost by default.
Resource caps — response-size, timeout, page-concurrency, crawl page/depth limits to protect the host.
robots.txt + rate limiting on by default.
Container — runs as a non-root user; secrets via env only.
Supply chain — verifying the image
CI signs the image keylessly with cosign (Sigstore) using GitLab's OIDC identity, and attaches an SPDX SBOM attestation. Verify before running:
cosign verify \
--certificate-oidc-issuer https://gitlab.cri.epita.fr \
--certificate-identity-regexp 'https://gitlab.cri.epita.fr/enzo.juhel/web-scraper//.*' \
registry.gitlab.cri.epita.fr/enzo.juhel/web-scraper@sha256:...Benchmark
benchmarks/run.py scores our scrape/extract against public datasets and
the Crawl4AI baseline, emitting a scorecard (per page-type F1/accuracy + a
Limitations section). Run locally or via the manual CI benchmark job:
uv run python benchmarks/run.py --output scorecard.mdLimitations
No paid proxies/CAPTCHA: heavily-defended sites (LinkedIn, Amazon, Cloudflare challenges) will sometimes block us. The benchmark scorecard quantifies where.
Main-content extraction is strong on articles, weaker on forums / product / listing pages (a known property of all extractors).
In-memory crawl/session state — single-process, single-user by design.
Development
uv run pre-commit install # local lint/secret hooks
uv run ruff check . && uv run ruff format --check .
uv run mypy src
uv run pytest -qMaintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables web scraping and crawling capabilities for LLM clients, supporting single-page scraping, multi-page website crawling, and web search with multiple engines (Playwright, Cheerio, Puppeteer) and flexible output formats including markdown, HTML, text, and screenshots.186MIT
- FlicenseNot gradedqualityDmaintenanceEnables LLMs to extract content from websites using automated static and dynamic scraping engines with built-in anti-bot protections. It provides tools for web data retrieval and stores results in MongoDB with support for JSON and CSV exports.
- AlicenseAqualityAmaintenanceWeb scraping, crawling, and structured data extraction for AI agents. 5 tools: scrape (clean markdown from any URL), crawl (entire sites), map (discover URLs), extract (structured JSON), and search. 833ms avg latency, single binary, self-hostable.8852AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceEnables LLMs to fetch and extract web content using browser automation, OCR, and multiple extraction methods, handling JavaScript rendering and anti-scraping techniques.17MIT
Related MCP Connectors
Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabili…
Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents.
Live web access for agents: scrape, SERP search, crawl/map, 74 collectors, datasets, proxies.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Prog-up/web-scraper-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server