Skip to main content
Glama
RebaubierWilliam

Web Scraper MCP Server

README.md
# Web Scraping with MCP

A **Model Context Protocol (MCP) server** that gives Claude (and any other MCP-compatible LLM client) the ability to scrape and extract data from live websites.

Inspired by [luminati-io/web-scraping-with-mcp](https://github.com/luminati-io/web-scraping-with-mcp).

---

## What it does

| Tool | Description |
|---|---|
| `fetch_page(url)` | Loads a URL with a headless Chromium browser (JS rendered) and caches the HTML |
| `extract_text(css_selector)` | Extracts text from elements matching a CSS selector |
| `extract_links()` | Returns all hyperlinks found on the cached page |
| `extract_table()` | Extracts the first HTML table as structured JSON |
| `extract_amazon()` | Specialized extractor for Amazon product pages |
| `list_cached_pages()` | Lists all URLs currently in the cache |
| `clear_cache(url?)` | Clears one or all cached pages |

---

## Architecture

```
Claude Desktop / Cursor IDE
        │
        │  MCP (JSON-RPC over stdio)
        ▼
  scraper_mcp.py   ◄── FastMCP server
        │
        ├── Playwright  (headless browser, JS rendering)
        └── lxml        (HTML parsing / XPath / CSS selectors)
```

---

## Installation

```bash
# 1. Clone this repo
git clone https://github.com/RebaubierWilliam/web-scraping-with-mcp.git
cd web-scraping-with-mcp

# 2. Create and activate a virtual environment
python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS / Linux
source .venv/bin/activate

# 3. Install dependencies
pip install -r requirements.txt

# 4. Install Playwright browsers
python -m playwright install chromium
```

---

## Connect to Claude Desktop

1. Open (or create) your Claude Desktop config file:
   - **Windows**: `%APPDATA%\Claude\claude_desktop_config.json`
   - **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json`

2. Add the following block, replacing the path:

```json
{
  "mcpServers": {
    "web-scraper": {
      "command": "python",
      "args": ["C:/Users/YourName/Source/web-scraping-with-mcp/scraper_mcp.py"]
    }
  }
}
```

3. Restart Claude Desktop. You should see a hammer icon confirming the tools are available.

---

## Usage examples

Once connected to Claude Desktop, you can ask:

```
Fetch the page https://www.amazon.fr/dp/B09V3KXJPB and extract the product details.
```

```
Go to https://news.ycombinator.com and give me all the links.
```

```
Scrape the prices table from https://example.com/pricing
```

---

## Run the server manually (for testing)

```bash
python scraper_mcp.py
```

The server communicates over **stdio** — it is meant to be launched by the MCP host, not used interactively.

---

## Credits

- Original concept: [luminati-io/web-scraping-with-mcp](https://github.com/luminati-io/web-scraping-with-mcp)
- [Model Context Protocol](https://modelcontextprotocol.io/)
- [Playwright for Python](https://playwright.dev/python/)
- [lxml](https://lxml.de/)