web-scrape
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@web-scrapefetch the content of example.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
web-scrape
CLI-first website scrape + anonymous browser fetch.
SSRF-safe public
http(s)onlyrobots.txt honored (fail-open on missing/unreachable robots)
Protection taxonomy — blocked / captcha / Cloudflare / login / rate-limit statuses instead of garbage text
Optional CapSolver — Turnstile + Cloudflare “Just a moment…” clears (sticky residential proxy; CapSolver needs socks5 for Webshare)
Thin MCP adapter — same library over JSON-RPC HTTP
Does not bypass protections by default. CapSolver is opt-in via env. No social-login profiles or cookies are reused.
Install
git clone https://github.com/teslashibe/web-scrape.git
cd web-scrape
npm install
npx playwright install chromium # for `fetch` / browser pathNode ≥ 20. Uses undici v7 (compatible with Node 20).
Related MCP server: scout-mcp-server
CLI (preferred for agents)
# Readable browser fetch (Playwright)
./bin/web-scrape.mjs fetch https://example.com/ --json
# Structured single-page scrape (HTTP + HTML extract)
./bin/web-scrape.mjs scrape https://example.com/ --json
# MCP HTTP server (default :8091)
./bin/web-scrape.mjs serve --port 8091Exit codes: 0 ok · 2 structured non-ok status · 1 usage/error.
Library
import { scrape, browserFetchURL } from "web-scrape";
const page = await browserFetchURL({ url: "https://example.com/" });
const brief = await scrape({ url: "https://example.com/" });Docker / Kubernetes
docker build -t web-scrape .
docker run --rm -p 8091:8091 \
-e WEB_FETCH_TURNSTILE_PROVIDER=capsolver \
-e WEB_FETCH_TURNSTILE_API_KEY=CAP-... \
-e WEBSHARE_RESIDENTIAL_USER=... \
-e WEBSHARE_RESIDENTIAL_PASS=... \
-e WEBSHARE_USERNAME_TEMPLATE='{user}-{country}-1' \
web-scrapeExpose Service port 8091. Probe GET /mcp/v1/ready (200 ready / 503 not_ready).
MCP
POST /mcp/v1 JSON-RPC 2.0 (initialize, tools/list, tools/call)
GET /mcp/v1/health liveness + limits + metrics
GET /mcp/v1/ready readinessTools:
Tool | Purpose |
| Ephemeral Playwright fetch → readable text or structured status |
| Single-page HTTP scrape → title/summary/product/audience fields |
CapSolver + proxy (optional)
Env | Purpose |
| Enable provider |
| CapSolver key |
| Residential proxy |
| Prefer sticky |
| Full proxy URL override |
| Playwright deadline (default rises when CapSolver+proxy set; max 180s) |
CapSolver AntiCloudflareTask is sent as socks5:host:port:user:pass (no page html — CapSolver rejects it as invalid html). Playwright egress stays HTTP proxy. Managed “Just a moment…” pages use AntiCloudflareTask even when a Turnstile iframe is present. Never log proxy credentials or API keys.
Tests
npm test # unit fixtures (no live network)
npm run validate # loopback MCP discovery/call
npm run smoke # status-matrix smokeLicense
Apache-2.0. CapSolver / Webshare are optional third-party paid services — you bring your own keys.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseCqualityBmaintenanceProvides browser automation and web scraping as MCP tools, enabling autonomous URL ingestion, crawling, extraction, and anti-bot handling with interactive browser control.625MIT
- AlicenseAqualityBmaintenanceMCP server for browser automation with anti-detection. Scout pages, find elements, interact with websites, and monitor network traffic from any AI client that supports the Model Context Protocol.211MIT
- Alicense-qualityAmaintenanceRemote MCP server for web scraping with anti-bot evasion. Provides stealth HTTP fetching, headless browser with Cloudflare bypass, CSS selectors, YouTube transcripts, and Markdown conversion.MIT
- Alicense-qualityDmaintenanceProvides a real browser that bypasses bot detection (Cloudflare, Turnstile) for AI agents, enabling navigation, clicking, typing, screenshots, and data collection through MCP tools.98MIT
Related MCP Connectors
Free remote MCP server for fetching public web pages through a rotating proxy pool.
One MCP for 160+ live web-data APIs — clean JSON from sites that block scrapers.
The most accurate web access API. Stop getting blocked.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/teslashibe/web-scrape'
If you have feedback or need assistance with the MCP directory API, please join our Discord server