Skip to main content
Glama

Scrapeline — a web scraper built for Claude, not for clicking

Scrapeline is a browser extension plus a local MCP server. Together they let Claude (Claude Code, Claude Desktop, or any other MCP client) read and scrape the pages you already have open — logged-in sessions included — and crawl, extract, and export data on request, entirely on your own machine. There's no cloud service, no account, and nothing leaves your computer unless you export it yourself.

If you've used a no-code scraper extension before (click one item, get a spreadsheet), the difference here is who's driving: instead of you pointing and clicking, you tell Claude what you want in plain English, and Claude calls the same kind of extraction engine through MCP tools. See § How this compares to a click-to-scrape extension if that's what you're picturing.

See USE_CASES.md for 340 concrete things people use it for, and PLAN.md for the full product plan, architecture rationale, and build history.

Claude ──stdio──▶ MCP server ──ws://127.0.0.1:8765──▶ extension (service worker) ──▶ your tab

Quickstart

Requires Node 20+ (developed on Node 24).

git clone https://github.com/surypratapsingh/simple-web-crawler-scraper-extension.git
cd simple-web-crawler-scraper-extension
npm install
npm run build          # builds mcp-server/dist and extension/.output/chrome-mv3
  1. Load the extension. Open chrome://extensions (works the same in Edge, at edge://extensions; Brave too), turn on Developer mode, click Load unpacked, and pick extension/.output/chrome-mv3.

  2. Register the MCP server.

    • Claude Code: this repo ships a project-scoped .mcp.json — run claude from this folder and approve it when prompted. From anywhere else:

      claude mcp add scrapeline -- node /path/to/simple-web-crawler-scraper-extension/mcp-server/dist/server.mjs
    • Claude Desktop: add to claude_desktop_config.json:

      { "mcpServers": { "scrapeline": {
          "command": "node",
          "args": ["/path/to/simple-web-crawler-scraper-extension/mcp-server/dist/server.mjs"] } } }
  3. Click the extension's toolbar icon — the status dot turns green once it's connected to the MCP server.

  4. Ask Claude something like "list my open tabs", "scrape the product grid on this page to CSV", or "crawl this listing across all its pages and save it".

Several Claude sessions can share one browser: the first MCP process becomes the hub, later ones relay through it automatically.

After pulling changes or running npm run build again, reload the extension in chrome://extensions and restart your Claude session, so both sides pick up the new build.

There's also a one-page project site with the same information laid out visually: site/index.html (open it in a browser).

Related MCP server: gotham-browser

Connecting other AI tools (MCP is an open protocol)

Scrapeline isn't Claude-exclusive — it's a normal local MCP server over stdio, so anything that speaks MCP can use it the same way:

  • Claude Code / Claude Desktop — see Quickstart above.

  • Cursor, Windsurf, or any other MCP-capable editor/agent — point its MCP config at the same built server: { "command": "node", "args": ["/path/to/simple-web-crawler-scraper-extension/mcp-server/dist/server.mjs"] }. Multiple clients can share one running browser at once (hub/relay, as above).

  • claude.ai in the browser — this is the one exception: the web app runs entirely in Anthropic's cloud and has no path to a stdio process on your machine, so it can never see a local server no matter how it's configured there. The only way around that is putting a public URL in front of the local server (e.g. via ngrok or a Cloudflare Tunnel) and adding real authentication first — Scrapeline currently trusts whatever calls it and reads your live, logged-in tabs, so don't expose it publicly without that.

What it can do

Everything below is exposed as an MCP tool Claude calls directly — there's no UI to learn, you just describe what you want. Costs are approximate token counts for a typical page.

Look around (cheap, escalate as needed)

Tool

What it does

Typical cost

browser_status

Is the extension connected?

~20 tok

list_tabs

id / title / url of every open tab

~10 tok/tab

page_info

title, h1, description, element counts

~60 tok

page_outline

headings, forms, tables/lists, buttons, top links (+ refs)

300–600 tok

page_text

readable main text, boilerplate stripped, offset to page through it

you set the budget

page_query

text/attributes for a CSS selector or a ref from page_outline

small

page_scroll

scroll (triggers lazy-loading / infinite scroll)

tiny

open_url

open a URL in a new background tab

tiny

All page-derived text is prefixed with an "untrusted content" marker, so text embedded in a page is treated as data by Claude, never as instructions.

Iframes and shadow DOM are handled transparently: open shadow roots are read like normal DOM, and page_outline lists real cross-origin sub-frames you can target with frameId on any of the tools above.

Extract structured data from one page

page_extract finds the main table or repeating list on a page automatically (product grids, search results, feeds) — or, given a rowSelector and per-field selectors, extracts exactly what you specify. Either way, output is TSV (cheap) or JSON. Add saveAs: "csv" | "json" | "ndjson" and it writes every row to a file under ~/.scrapeline/exports and hands Claude back just the path and a 3-row preview, so a 5,000-row table costs about as many tokens as a 5-row one.

Crawl a whole listing across pages

crawl_start follows pagination — "Next" links/buttons, a {n} URL template, or infinite scroll — collecting de-duplicated rows into its own background tab (your tabs are left alone unless you ask for inPlace). It's a background job: it waits briefly for completion, then hands back a job id to poll with crawl_status. State is saved after every page, so a crawl survives the browser extension being reloaded or restarted and picks up where it left off.

It's polite and safe by construction: at least 500ms between pages (with jitter), hard caps on pages/rows/time, a check of robots.txt (warns by default; respectRobots: true refuses outright if disallowed), and it will only ever click a control that looks like pagination — never a form-submit button, and never anything labelled like "delete", "buy", "sign out" etc., even if you hand it that selector directly.

crawl_start { url: "https://example.com/listing", maxPages: 20, saveAs: "csv" }

Add detail: {} and it will also open each new row's link (same site only) in a second background tab and append fields from that page — either automatically (JSON-LD/meta: price, SKU, availability, description...) or via your own selectors. Every row gets a detail_status so you can see what worked.

Bulk-extract a list of URLs ("Page Extractor")

Already have the URLs — a list of product pages, profiles, or articles? extract_urls visits each one (same background-job machinery as crawling) and returns one row per page, again either automatically from JSON-LD/meta or via your own field selectors.

extract_urls { urls: ["https://shop.example.com/p/1", ".../p/2", ...], saveAs: "json" }

find_contacts scans a page (or, with urls / discoverFrom, many pages) for email addresses, phone numbers, and social profile links (X, Instagram, Facebook, LinkedIn, YouTube, TikTok, GitHub, Reddit, and more) — filtering out the platforms' own generic pages (share buttons, login screens) so you get real profiles. discoverFrom follows same-site links to find a Contact/About page on its own.

find_contacts { discoverFrom: { url: "https://example.com", maxPages: 10 } }

List or download every image on a page

find_images lists image URLs on a page (or across many, like find_contacts) with size filtering so icons and tracking pixels are skipped automatically. Add download: true and it fetches the actual files to ~/.scrapeline/downloads/<folder> instead of just listing URLs.

Map a whole site from its sitemap

sitemap_explore reads a site's sitemap.xml (following a sitemap index, and checking robots.txt for a Sitemap: line if you just give it a bare domain) and returns the URLs it lists — instantly, with no browser tab involved, since sitemaps are public by design. Feed the result straight into extract_urls or crawl_start.

Export a Shopify store's whole catalogue

shopify_catalog reads any live Shopify store's public /products.json endpoint and flattens it to one row per variant (product, variant, SKU, price, compare-at price, availability, vendor, type, tags, image, URL) — again instant and tab-free.

Save a setup and run it again

recipe_save stores the arguments of any call above under a name; recipe_run replays it later (optionally overriding some arguments), and recipe_list / recipe_delete manage what's saved. Recipes are plain JSON files under ~/.scrapeline/recipes, so they're easy to inspect, back up, or hand to a colleague running the same setup.

recipe_save { name: "weekly competitor scan", tool: "crawl_start", args: { url: "...", saveAs: "csv" } }
recipe_run  { name: "weekly competitor scan" }

Safety model

  • The bridge binds 127.0.0.1 only, and rejects any WebSocket connection that isn't from a chrome-extension:// origin — a malicious web page cannot talk to it. An optional shared token (SCRAPELINE_TOKEN) and pinned extension id (SCRAPELINE_EXTENSION_ID) add defense in depth.

  • The popup has a kill switch (Claude access on/off), a per-host block list, and an audit log of every call (method + host only — page content is never logged).

  • Only http(s) pages are readable; the Chrome/Edge Web Store and browser-internal pages are refused. There is no general click/type/fill capability — the only clicks anywhere in the system are the crawler's own pagination controls, which are restricted to things that look like pagination and never anything destructive-sounding.

  • Filenames for exports, downloads, and recipes are sanitized (no directories, no traversal) and files are never overwritten silently.

  • The development build requests <all_urls> so it works everywhere out of the box; a Chrome Web Store submission would instead use activeTab + optional host permissions.

Env vars for the MCP server: SCRAPELINE_PORT (default 8765), SCRAPELINE_TOKEN, SCRAPELINE_EXTENSION_ID, SCRAPELINE_EXPORT_DIR, SCRAPELINE_DOWNLOADS_DIR, SCRAPELINE_RECIPES_DIR. Change the port in the popup to match if you set one.

How this compares to a click-to-scrape extension

Point-and-click scraper extensions (highlight one item, get a spreadsheet) are built for a human sitting at the browser deciding what to click. Scrapeline is built for the case where an AI agent is doing the work on your behalf: you describe the goal, Claude picks the right tool (or chains several — outline, then extract, then crawl, then save), and it can act on many pages unattended within the guardrails above. That trade-off means:

  • No visual point-and-click picker. Claude generates and refines CSS selectors itself from the page structure; you never draw a box around an element. (A human-facing picker is a plausible future addition, but isn't the point of this tool today.)

  • No hosted "run it while your laptop is off" cloud tier. Everything runs locally, by design — see the non-goals in PLAN.md. A crawl only runs while your browser and this MCP server are both up.

  • No spreadsheet UI. Results go to CSV/JSON/NDJSON files (or straight into Claude's answer); open them in whatever you already use.

  • No Google Sheets / cloud export integration yet — plain files only, for now.

Development

npm test               # unit tests (Vitest + jsdom): extraction/crawl/contacts/sitemap/shopify/... logic
npm run test:e2e       # Playwright: a real browser + the built extension + a real MCP client over stdio
npm run typecheck
cd extension && npx wxt  # dev mode with HMR

The E2E suite launches installed Edge by default (channel: 'msedge' in tests/e2e/harness.ts) — Playwright's own bundled Chromium can get blocked by some Windows Application Control / antivirus policies when copied into a temp folder, which installed, signed browsers aren't subject to. Set SCRAPELINE_BROWSER_CHANNEL=chromium (or chrome) or SCRAPELINE_CHROMIUM=/path/to/exe to override.

License

MIT.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Bridges browser content, developer tools data, and web page interactions with Claude through MCP. Enables page inspection, DOM analysis, JavaScript execution, console monitoring, network activity tracking, and screenshot capture across multiple browser tabs.
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables Claude Code to control a real browser using AI for web scraping, competitive intelligence, and UX auditing through the MCP protocol.
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.
    1
    MIT