Skip to main content
Glama
teslashibe

web-scrape

by teslashibe

web-scrape

CLI-first website scrape + anonymous browser fetch.

  • SSRF-safe public http(s) only

  • robots.txt honored (fail-open on missing/unreachable robots)

  • Protection taxonomy — blocked / captcha / Cloudflare / login / rate-limit statuses instead of garbage text

  • Optional CapSolver — Turnstile + Cloudflare “Just a moment…” clears (sticky residential proxy; CapSolver needs socks5 for Webshare)

  • Thin MCP adapter — same library over JSON-RPC HTTP

Does not bypass protections by default. CapSolver is opt-in via env. No social-login profiles or cookies are reused.

Install

git clone https://github.com/teslashibe/web-scrape.git
cd web-scrape
npm install
npx playwright install chromium   # for `fetch` / browser path

Node ≥ 20. Uses undici v7 (compatible with Node 20).

Related MCP server: scout-mcp-server

CLI (preferred for agents)

# Readable browser fetch (Playwright)
./bin/web-scrape.mjs fetch https://example.com/ --json

# Structured single-page scrape (HTTP + HTML extract)
./bin/web-scrape.mjs scrape https://example.com/ --json

# MCP HTTP server (default :8091)
./bin/web-scrape.mjs serve --port 8091

Exit codes: 0 ok · 2 structured non-ok status · 1 usage/error.

Library

import { scrape, browserFetchURL } from "web-scrape";

const page = await browserFetchURL({ url: "https://example.com/" });
const brief = await scrape({ url: "https://example.com/" });

Docker / Kubernetes

docker build -t web-scrape .
docker run --rm -p 8091:8091 \
  -e WEB_FETCH_TURNSTILE_PROVIDER=capsolver \
  -e WEB_FETCH_TURNSTILE_API_KEY=CAP-... \
  -e WEBSHARE_RESIDENTIAL_USER=... \
  -e WEBSHARE_RESIDENTIAL_PASS=... \
  -e WEBSHARE_USERNAME_TEMPLATE='{user}-{country}-1' \
  web-scrape

Expose Service port 8091. Probe GET /mcp/v1/ready (200 ready / 503 not_ready).

MCP

POST /mcp/v1          JSON-RPC 2.0 (initialize, tools/list, tools/call)
GET  /mcp/v1/health   liveness + limits + metrics
GET  /mcp/v1/ready    readiness

Tools:

Tool

Purpose

browser_fetch_url

Ephemeral Playwright fetch → readable text or structured status

scrape_website_context

Single-page HTTP scrape → title/summary/product/audience fields

CapSolver + proxy (optional)

Env

Purpose

WEB_FETCH_TURNSTILE_PROVIDER=capsolver

Enable provider

WEB_FETCH_TURNSTILE_API_KEY

CapSolver key

WEBSHARE_RESIDENTIAL_USER / PASS

Residential proxy

WEBSHARE_USERNAME_TEMPLATE

Prefer sticky {user}-{country}-1 (required for cf_clearance IP affinity)

WEB_SCRAPER_PROXY_URL

Full proxy URL override

BROWSER_FETCH_TIMEOUT_MS

Playwright deadline (default rises when CapSolver+proxy set; max 180s)

CapSolver AntiCloudflareTask is sent as socks5:host:port:user:pass (no page html — CapSolver rejects it as invalid html). Playwright egress stays HTTP proxy. Managed “Just a moment…” pages use AntiCloudflareTask even when a Turnstile iframe is present. Never log proxy credentials or API keys.

Tests

npm test          # unit fixtures (no live network)
npm run validate  # loopback MCP discovery/call
npm run smoke     # status-matrix smoke

License

Apache-2.0. CapSolver / Webshare are optional third-party paid services — you bring your own keys.

A
license - permissive license
-
quality - not tested
B
maintenance

Maintenance

Maintainers
Response time
0dRelease cycle
2Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    C
    quality
    B
    maintenance
    Provides browser automation and web scraping as MCP tools, enabling autonomous URL ingestion, crawling, extraction, and anti-bot handling with interactive browser control.
    62
    5
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    MCP server for browser automation with anti-detection. Scout pages, find elements, interact with websites, and monitor network traffic from any AI client that supports the Model Context Protocol.
    21
    1
    MIT
  • A
    license
    -
    quality
    A
    maintenance
    Remote MCP server for web scraping with anti-bot evasion. Provides stealth HTTP fetching, headless browser with Cloudflare bypass, CSS selectors, YouTube transcripts, and Markdown conversion.
    MIT
  • A
    license
    -
    quality
    D
    maintenance
    Provides a real browser that bypasses bot detection (Cloudflare, Turnstile) for AI agents, enabling navigation, clicking, typing, screenshots, and data collection through MCP tools.
    98
    MIT

View all related MCP servers

Related MCP Connectors

  • Free remote MCP server for fetching public web pages through a rotating proxy pool.

  • One MCP for 160+ live web-data APIs — clean JSON from sites that block scrapers.

  • The most accurate web access API. Stop getting blocked.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/teslashibe/web-scrape'

If you have feedback or need assistance with the MCP directory API, please join our Discord server