Skip to main content
Glama
teslashibe

web-scrape

by teslashibe
README.md
# web-scrape

CLI-first website scrape + anonymous browser fetch.

- **SSRF-safe** public `http(s)` only
- **robots.txt** honored (fail-open on missing/unreachable robots)
- **Protection taxonomy** — blocked / captcha / Cloudflare / login / rate-limit statuses instead of garbage text
- **Optional CapSolver** — Turnstile + Cloudflare “Just a moment…” clears (sticky residential proxy; CapSolver needs **socks5** for Webshare)
- **Thin MCP adapter** — same library over JSON-RPC HTTP

> Does **not** bypass protections by default. CapSolver is opt-in via env. No social-login profiles or cookies are reused.

## Install

```bash
git clone https://github.com/teslashibe/web-scrape.git
cd web-scrape
npm install
npx playwright install chromium   # for `fetch` / browser path
```

Node **≥ 20**. Uses `undici` **v7** (compatible with Node 20).

## CLI (preferred for agents)

```bash
# Readable browser fetch (Playwright)
./bin/web-scrape.mjs fetch https://example.com/ --json

# Structured single-page scrape (HTTP + HTML extract)
./bin/web-scrape.mjs scrape https://example.com/ --json

# MCP HTTP server (default :8091)
./bin/web-scrape.mjs serve --port 8091
```

Exit codes: `0` ok · `2` structured non-ok status · `1` usage/error.

## Library

```js
import { scrape, browserFetchURL } from "web-scrape";

const page = await browserFetchURL({ url: "https://example.com/" });
const brief = await scrape({ url: "https://example.com/" });
```

## Docker / Kubernetes

```bash
docker build -t web-scrape .
docker run --rm -p 8091:8091 \
  -e WEB_FETCH_TURNSTILE_PROVIDER=capsolver \
  -e WEB_FETCH_TURNSTILE_API_KEY=CAP-... \
  -e WEBSHARE_RESIDENTIAL_USER=... \
  -e WEBSHARE_RESIDENTIAL_PASS=... \
  -e WEBSHARE_USERNAME_TEMPLATE='{user}-{country}-1' \
  web-scrape
```

Expose Service port `8091`. Probe `GET /mcp/v1/ready` (200 ready / 503 not_ready).

## MCP

```
POST /mcp/v1          JSON-RPC 2.0 (initialize, tools/list, tools/call)
GET  /mcp/v1/health   liveness + limits + metrics
GET  /mcp/v1/ready    readiness
```

Tools:

| Tool | Purpose |
| --- | --- |
| `browser_fetch_url` | Ephemeral Playwright fetch → readable text or structured status |
| `scrape_website_context` | Single-page HTTP scrape → title/summary/product/audience fields |

## CapSolver + proxy (optional)

| Env | Purpose |
| --- | --- |
| `WEB_FETCH_TURNSTILE_PROVIDER=capsolver` | Enable provider |
| `WEB_FETCH_TURNSTILE_API_KEY` | CapSolver key |
| `WEBSHARE_RESIDENTIAL_USER` / `PASS` | Residential proxy |
| `WEBSHARE_USERNAME_TEMPLATE` | Prefer sticky `{user}-{country}-1` (required for `cf_clearance` IP affinity) |
| `WEB_SCRAPER_PROXY_URL` | Full proxy URL override |
| `BROWSER_FETCH_TIMEOUT_MS` | Playwright deadline (default rises when CapSolver+proxy set; max 180s) |

CapSolver `AntiCloudflareTask` is sent as `socks5:host:port:user:pass` (no page `html` — CapSolver rejects it as `invalid html`). Playwright egress stays HTTP proxy. Managed “Just a moment…” pages use AntiCloudflareTask even when a Turnstile iframe is present. Never log proxy credentials or API keys.

## Tests

```bash
npm test          # unit fixtures (no live network)
npm run validate  # loopback MCP discovery/call
npm run smoke     # status-matrix smoke
```

## License

Apache-2.0. CapSolver / Webshare are optional third-party paid services — you bring your own keys.