Web Scraper MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Web Scraper MCP ServerScrape Hacker News for top stories"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Web Scraping with MCP
A Model Context Protocol (MCP) server that gives Claude (and any other MCP-compatible LLM client) the ability to scrape and extract data from live websites.
Inspired by luminati-io/web-scraping-with-mcp.
What it does
Tool | Description |
| Loads a URL with a headless Chromium browser (JS rendered) and caches the HTML |
| Extracts text from elements matching a CSS selector |
| Returns all hyperlinks found on the cached page |
| Extracts the first HTML table as structured JSON |
| Specialized extractor for Amazon product pages |
| Lists all URLs currently in the cache |
| Clears one or all cached pages |
Related MCP server: chromium-mcp
Architecture
Claude Desktop / Cursor IDE
│
│ MCP (JSON-RPC over stdio)
▼
scraper_mcp.py ◄── FastMCP server
│
├── Playwright (headless browser, JS rendering)
└── lxml (HTML parsing / XPath / CSS selectors)Installation
# 1. Clone this repo
git clone https://github.com/RebaubierWilliam/web-scraping-with-mcp.git
cd web-scraping-with-mcp
# 2. Create and activate a virtual environment
python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS / Linux
source .venv/bin/activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Install Playwright browsers
python -m playwright install chromiumConnect to Claude Desktop
Open (or create) your Claude Desktop config file:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.json
Add the following block, replacing the path:
{
"mcpServers": {
"web-scraper": {
"command": "python",
"args": ["C:/Users/YourName/Source/web-scraping-with-mcp/scraper_mcp.py"]
}
}
}Restart Claude Desktop. You should see a hammer icon confirming the tools are available.
Usage examples
Once connected to Claude Desktop, you can ask:
Fetch the page https://www.amazon.fr/dp/B09V3KXJPB and extract the product details.Go to https://news.ycombinator.com and give me all the links.Scrape the prices table from https://example.com/pricingRun the server manually (for testing)
python scraper_mcp.pyThe server communicates over stdio — it is meant to be launched by the MCP host, not used interactively.
Credits
This server cannot be deployed
Maintenance
Related MCP Connectors
One MCP for the Web. Easily search, crawl, navigate, and extract websites without getting blocked.…
Scrape, crawl and search the web for AI agents via MCP.
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Related MCP Servers
AlicenseAqualityCmaintenanceScrape any webpage and return clean markdown, HTML, or structured JSON. Bypasses anti-bot protection, renders JavaScript (React/Vue/Angular), supports premium residential proxies and CSS extraction. Works with any MCP client — no local install required.1635 npm19MIT- AlicenseAqualityDmaintenanceMCP server for web fetching and automation using Chromium headless browser. Enables fetching pages as markdown, taking screenshots, automating interactions, and extracting data via CSS selectors.5MIT
- AlicenseCqualityCmaintenanceProvides browser automation and web scraping as MCP tools, enabling autonomous URL ingestion, crawling, extraction, and anti-bot handling with interactive browser control.625MIT
- FlicenseNot gradedqualityCmaintenanceHeadless browser automation for LLM agents via REST API or MCP tools. Enables navigating pages, reading structured content, clicking elements, filling forms, and executing JavaScript.-