web-scrape
by teslashibe
README.md
# web-scrape
CLI-first website scrape + anonymous browser fetch.
- **SSRF-safe** public `http(s)` only
- **robots.txt** honored (fail-open on missing/unreachable robots)
- **Protection taxonomy** — blocked / captcha / Cloudflare / login / rate-limit statuses instead of garbage text
- **Optional CapSolver** — Turnstile + Cloudflare “Just a moment…” clears (sticky residential proxy; CapSolver needs **socks5** for Webshare)
- **Thin MCP adapter** — same library over JSON-RPC HTTP
> Does **not** bypass protections by default. CapSolver is opt-in via env. No social-login profiles or cookies are reused.
## Install
```bash
git clone https://github.com/teslashibe/web-scrape.git
cd web-scrape
npm install
npx playwright install chromium # for `fetch` / browser path
```
Node **≥ 20**. Uses `undici` **v7** (compatible with Node 20).
## CLI (preferred for agents)
```bash
# Readable browser fetch (Playwright)
./bin/web-scrape.mjs fetch https://example.com/ --json
# Structured single-page scrape (HTTP + HTML extract)
./bin/web-scrape.mjs scrape https://example.com/ --json
# MCP HTTP server (default :8091)
./bin/web-scrape.mjs serve --port 8091
```
Exit codes: `0` ok · `2` structured non-ok status · `1` usage/error.
## Library
```js
import { scrape, browserFetchURL } from "web-scrape";
const page = await browserFetchURL({ url: "https://example.com/" });
const brief = await scrape({ url: "https://example.com/" });
```
## Docker / Kubernetes
```bash
docker build -t web-scrape .
docker run --rm -p 8091:8091 \
-e WEB_FETCH_TURNSTILE_PROVIDER=capsolver \
-e WEB_FETCH_TURNSTILE_API_KEY=CAP-... \
-e WEBSHARE_RESIDENTIAL_USER=... \
-e WEBSHARE_RESIDENTIAL_PASS=... \
-e WEBSHARE_USERNAME_TEMPLATE='{user}-{country}-1' \
web-scrape
```
Expose Service port `8091`. Probe `GET /mcp/v1/ready` (200 ready / 503 not_ready).
## MCP
```
POST /mcp/v1 JSON-RPC 2.0 (initialize, tools/list, tools/call)
GET /mcp/v1/health liveness + limits + metrics
GET /mcp/v1/ready readiness
```
Tools:
| Tool | Purpose |
| --- | --- |
| `browser_fetch_url` | Ephemeral Playwright fetch → readable text or structured status |
| `scrape_website_context` | Single-page HTTP scrape → title/summary/product/audience fields |
## CapSolver + proxy (optional)
| Env | Purpose |
| --- | --- |
| `WEB_FETCH_TURNSTILE_PROVIDER=capsolver` | Enable provider |
| `WEB_FETCH_TURNSTILE_API_KEY` | CapSolver key |
| `WEBSHARE_RESIDENTIAL_USER` / `PASS` | Residential proxy |
| `WEBSHARE_USERNAME_TEMPLATE` | Prefer sticky `{user}-{country}-1` (required for `cf_clearance` IP affinity) |
| `WEB_SCRAPER_PROXY_URL` | Full proxy URL override |
| `BROWSER_FETCH_TIMEOUT_MS` | Playwright deadline (default rises when CapSolver+proxy set; max 180s) |
CapSolver `AntiCloudflareTask` is sent as `socks5:host:port:user:pass` (no page `html` — CapSolver rejects it as `invalid html`). Playwright egress stays HTTP proxy. Managed “Just a moment…” pages use AntiCloudflareTask even when a Turnstile iframe is present. Never log proxy credentials or API keys.
## Tests
```bash
npm test # unit fixtures (no live network)
npm run validate # loopback MCP discovery/call
npm run smoke # status-matrix smoke
```
## License
Apache-2.0. CapSolver / Webshare are optional third-party paid services — you bring your own keys.
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues