PyScrappy
PyScrappy is a web scraping MCP server that provides AI agents with 22 tools to retrieve structured data from across the web — from general URLs to specialized platforms.
General Web Scraping
scrape_url— Scrape any URL for structured text, links, images, tables, and metadata; supports CSS selectors, pagination, and optional JavaScript rendering
Data & Research
scrape_wikipedia— Fetch Wikipedia articles in full, paragraph, or header modescrape_stock— Get Yahoo Finance stock quotes, historical price data, and company profilesscrape_news— Fetch articles from RSS/Atom feeds, auto-discover feeds from a news site, or extract full text from a single article URLsearch_images— Search for images and return URLs and metadata (default engine: Bing)search_youtube— Search YouTube videos and return titles, channels, links, and metadatasearch_linkedin_jobs— Search public LinkedIn job postings by keyword and locationsearch_github— Search GitHub repositories by query, sortable by stars, forks, or recencysearch_hackernews— Search Hacker News stories by relevance or datesearch_books— Search books via Open Library by title, author, or free textget_weather— Get current weather (temperature, humidity, wind, condition) for any location (no API key required)get_crypto— Get cryptocurrency prices, market cap, and 24h change via CoinGeckoconvert_currency— Get exchange rates and convert amounts between currenciesdefine_word— Look up English word definitions, part of speech, and usage exampleslookup_movie— Look up movie/TV info from IMDB via the OMDb API (requiresOMDB_API_KEY)
E-Commerce
search_amazon— Search Amazon products and return title, price, rating, and imagesearch_newegg— Search Newegg for electronics and computer hardwaresearch_ikea— Search IKEA furniture and home products with per-country pricing
Food Delivery
search_ubereats— List Uber Eats restaurants delivering in a cityget_ubereats_menu— Retrieve a full menu (items and prices) for a specific Uber Eats restaurantscrape_zomato— Search restaurants on Zomato by city (with optional cuisine/name filter)
Entertainment
search_soundcloud— Search SoundCloud tracks (uses browser backend for JS rendering)
Scrapes Amazon marketplace for product listings, including titles, prices, and details.
Scrapes GitHub repositories or profiles (details not fully shown in excerpt but listed as built-in scraper).
Scrapes IKEA product search results per country, including prices and details.
Fetches movie and TV information via OMDb API, such as title, year, rating, and genre.
Scrapes Newegg electronics and hardware product listings.
Scrapes RSS feeds (e.g., news articles) from any provided feed URL.
Searches SoundCloud for tracks and returns metadata like title and plays.
Scrapes Uber Eats restaurant listings and menus by city and locale.
Scrapes Wikipedia articles, summaries, and infoboxes by query.
Searches YouTube for videos and returns metadata such as title, views, and URL.
Scrapes Zomato restaurant listings by city.
PyScrappy is an AI-native web scraping toolkit that turns websites into structured, LLM-ready data. Use it as a Python library or expose it as an MCP server for AI agents.
📖 Documentation: pyscrappy.vercel.app
Key features
Generic scraper — give it any URL, get back structured text, links, images, tables, and metadata
LLM-ready output —
.to_markdown()turns any result into clean Markdown; also.to_json()and.to_dataframe()MCP server — expose the scrapers as tools for AI agents (Claude, Cursor, local LLMs, …)
JS rendering — optional Playwright backend for JavaScript-heavy sites
Custom selectors — pass CSS selectors to extract exactly what you need
Chainable
Selector— navigate HTML directly with CSS/XPath,find_all,find_by_text, andfind_similar(Scrapy/BeautifulSoup-style)Adaptive (self-healing) selectors — remember an element and relocate it by similarity when a site changes its markup, so scrapers don't silently break
Concurrent scraping —
scrape_many/scrape_allrun scrapes in parallelProxy & scraping-API support — route through a proxy or ScraperAPI/ScrapeOps for blocked sites
TLS-fingerprint impersonation —
impersonate="chrome"gets past anti-bot filters that block plain clients (optionalcurl_cffibackend)Command-line extract —
pyscrappy extract <url> out.mdscrapes a URL straight to a file, no codeRetry & rate-limiting — built-in exponential backoff and per-domain rate limiting
Type-safe — full type hints,
py.typedmarker20+ built-in scrapers — Wikipedia, IMDB, stocks, news, GitHub, Amazon/IKEA, YouTube, and more
Related MCP server: webclaw
Installation
pip install pyscrappyOptional extras:
# Browser support (for JS-rendered pages)
pip install 'pyscrappy[browser]'
playwright install chromium
# DataFrame support
pip install 'pyscrappy[dataframe]'
# MCP server (use PyScrappy's scrapers as AI-agent tools)
pip install 'pyscrappy[mcp]'
# Stealth (TLS-fingerprint impersonation to bypass anti-bot filters)
pip install 'pyscrappy[stealth]'
# Everything
pip install 'pyscrappy[all]'For AI agents
PyScrappy ships an MCP server that exposes its scrapers as tools, so an agent (Claude, Cursor, an OpenAI agent, a local LLM) can pull structured web data from any URL and hand it straight to the model:
AI agent ──MCP tool call──▶ PyScrappy ──fetch + extract──▶ Any website
▲ │
└────────────── clean Markdown / JSON ◀───────────────────────┘pip install 'pyscrappy[mcp]'
claude mcp add pyscrappy pyscrappy-mcpThen just ask: "use pyscrappy to summarize the latest headlines from bbc.com." See MCP server for the full setup and tool list.
Local models (Ollama), no MCP host needed
Ollama can't talk MCP on its own, so normally you'd run a host (Goose, Cline, …) in between. PyScrappy skips that with a built-in agent that talks to Ollama directly and lets a local model call the scrapers as tools:
pip install 'pyscrappy[mcp]' # needs Python 3.10+
pyscrappy chat --model qwen2.5 "what's the current AAPL quote?"It exposes the same 22 tools as the MCP server. The only requirement is a model
that supports tool calling (Llama 3.1, Qwen 2.5, Mistral, …); how well it
picks the right tool is up to the model. Point it at a remote Ollama with
--host, and pass -v to see each tool call.
MCP server (use PyScrappy from an AI agent)
PyScrappy ships an optional Model Context Protocol server, so an AI agent (e.g. Claude) can call PyScrappy's scrapers as tools and get structured web data back.
pip install 'pyscrappy[mcp]'The MCP extra installs the standalone fastmcp package and requires Python 3.10
or newer. On Python 3.9 the core scraping library still works, but the MCP server
is unavailable.
This installs the pyscrappy-mcp command. It uses stdio by default for local MCP
clients; Streamable HTTP and legacy SSE are available for remote deployments:
pyscrappy-mcp # stdio (default)
pyscrappy-mcp --http # Streamable HTTP
pyscrappy-mcp --sse # legacy SSEYou can also run the stdio server with python -m pyscrappy.mcp.
Register with Claude Code
claude mcp add pyscrappy pyscrappy-mcpRegister with Claude Desktop
Add to your claude_desktop_config.json and restart the app:
{
"mcpServers": {
"pyscrappy": {
"command": "pyscrappy-mcp"
}
}
}Tip: Claude Desktop does not inherit your shell
PATH. Ifpyscrappy-mcpis not found, use the absolute path to the command (e.g. the one printed bywhich pyscrappy-mcp).
Available tools
The server exposes 20+ tools. The most common ones are scrape_url (any
URL → text, links, images, tables, metadata), scrape_wikipedia,
scrape_stock, scrape_news, and search_github — plus many more
covering image/YouTube/LinkedIn/Hacker News/book search, weather, crypto,
currency, dictionary, Amazon/Newegg/IKEA/SoundCloud, IMDB, and Zomato/Uber Eats.
To see the full, live list, ask the agent to call the list_available_scrapers
tool, or from a shell:
python -c "from pyscrappy import list_scrapers; print(', '.join(sorted(list_scrapers())))"The lookup_movie tool needs a free OMDb API
key. Pass it to the server through your MCP client config, e.g. for Claude Desktop:
{
"mcpServers": {
"pyscrappy": {
"command": "pyscrappy-mcp",
"env": { "OMDB_API_KEY": "your-key" }
}
}
}Once registered, just ask the agent naturally, e.g. "use pyscrappy to get the latest headlines from bbc.co.uk and the AAPL stock quote."
Built-in scrapers
PyScrappy ships 24 built-in scrapers, and every one that works without a proxy is also exposed as an MCP tool.
A few of them:
GenericScraper— scrape any URL with auto-extraction (text, links, images, tables, metadata)Data / research —
WikipediaScraper,StockScraper(Yahoo Finance),NewsScraper(RSS/Atom),GitHubScraper,HackerNewsScraper, plus weather, crypto, currency, dictionary, image, LinkedIn-jobs, and book searchE-commerce —
AmazonScraper,NeweggScraper,IKEAScraperSocial / media / food —
YouTubeScraper, SoundCloud, Zomato, Uber Eats (Instagram / Twitter / Spotify also ship, but are blocked and need a proxy)
…and many more. To see the full, live list:
python -c "from pyscrappy import list_scrapers; print(', '.join(sorted(list_scrapers())))"IMDBScraper (lookup_movie) is the one exception that needs a key — a free
OMDb OMDB_API_KEY (see the
MCP config above for how to pass it).
Plugins
PyScrappy is extensible: you can add your own scrapers, and third parties can
ship them as standalone pyscrappy-<name> packages. A registered scraper works
everywhere a built-in does, including the MCP server and the pyscrappy chat
agent, with no change to PyScrappy core.
In your own code — register with the decorator:
from pyscrappy import BaseScraper, register_scraper, get_scraper
from pyscrappy.core.models import ScrapeResult, ScrapeMetadata
@register_scraper("reddit")
class RedditScraper(BaseScraper):
def scrape(self, subreddit: str, **kwargs) -> ScrapeResult:
data = self.fetch_and_parse(f"https://old.reddit.com/r/{subreddit}/.json")
# ... build a list of dicts ...
return ScrapeResult(data=[...], metadata=ScrapeMetadata(scraper="reddit"))
get_scraper("reddit")().scrape(subreddit="python")As a distributable package — advertise an entry point in your
pyproject.toml, and PyScrappy discovers it once your package is installed:
[project.entry-points."pyscrappy.scrapers"]
reddit = "pyscrappy_reddit:RedditScraper"After pip install pyscrappy-reddit, the scraper shows up in
list_scrapers(), and an AI agent can call it via the scrape_with MCP tool —
no core change required.
First-class MCP tools (optional). Add an mcp_tools mapping and your scraper
becomes a dedicated, typed MCP tool instead of only being reachable through the
generic scrape_with — its schema is derived from the method signature, so
agents get proper named arguments:
@register_scraper("reddit")
class RedditScraper(BaseScraper):
mcp_tools = {"search_reddit": "scrape"} # tool name -> method
def scrape(self, subreddit: str, sort: str = "hot") -> ScrapeResult:
...See the plugin template for a complete, copyable starting point, and the plugin guide for the full walkthrough.
Quick start
Scrape any URL → clean, LLM-ready Markdown
from pyscrappy import scrape
result = scrape("https://en.wikipedia.org/wiki/Web_scraping")
print(result.to_markdown()) # feed straight to an LLM
# ...or result.to_json() / result.to_dataframe()Prefer raw fields? Every result is a ScrapeResult with .data (a list of
dicts):
print(result.data[0]["metadata"]["title"])
print(result.data[0]["text"]["word_count"])Custom CSS selectors
from pyscrappy import GenericScraper
with GenericScraper() as gs:
result = gs.scrape(
url="https://news.ycombinator.com",
selectors={"title": ".titleline a", "score": ".score"},
)
for item in result.data:
print(item["title"], item.get("score", ""))Navigate HTML with Selector
When you want to traverse markup directly (Scrapy/BeautifulSoup-style) rather than
get back structured dicts, use Selector:
from pyscrappy import Selector
page = Selector(html) # or navigate any HTML string
page.css(".title::text").getall() # CSS with ::text / ::attr(name)
page.xpath("//a/@href").getall() # XPath (elements, text(), @attr)
page.find_all("h2", class_="title") # BeautifulSoup-style search
page.find_by_text("Add to cart", tag="button") # search by text content
first = page.css(".product")[0]
first.css(".price::text").get() # chainable
first.find_similar() # sibling elements shaped like this onecss() / xpath() return a SelectorList with .get() / .getall() / .text().
find_similar() locates elements with the same tag and overlapping classes, handy
for pulling every card/row once you've found one.
Adaptive (self-healing) selectors
A hard-coded CSS selector silently breaks the day a site changes its markup. Adaptive selectors survive that: save a fingerprint of the element the first time, and if the selector later matches nothing, relocate it by structural and textual similarity instead of returning empty.
from pyscrappy import Selector
# First run: match normally and remember this element under an id.
page = Selector(html_v1, url="https://shop.example.com")
price = page.css(".price", auto_save=True, adaptive_id="price").get()
# Later, after a redesign renamed ".price" — heal instead of breaking:
page = Selector(html_v2, url="https://shop.example.com")
result = page.css(".price", adaptive=True, adaptive_id="price")
print(result.get(), "→ confidence:", result.adaptive_confidence)How the relocation decides — and where it's stronger than a naive similarity match:
Weighted signals, not a flat average. A stable
id/data-*hook counts far more than a sibling-tag list, so weak signals can't outvote strong ones.Anchor-relative. It remembers the nearest stable ancestor (an id'd /
data-*container) and depth, so it survives layout reshuffles that move absolute positions.Volatility-aware text. Prices, dates, and counts are down-weighted, so healing stays reliable on exactly the fields that change most between scrapes.
Confidence-scored.
SelectorList.adaptive_confidence(0-100) tells you how sure the relocation was;threshold=sets the minimum to accept.
Fingerprints persist in a small JSON store (~/.pyscrappy/adaptive.json by
default, or $PYSCRAPPY_HOME), namespaced by site so the same adaptive_id on
two sites never collides. Adaptive is entirely opt-in: without adaptive=True, a
broken selector still just returns empty, exactly as before.
Site-specific scrapers
Every built-in scraper follows the same pattern — instantiate, scrape(...),
read result.data (or .to_dataframe() / .to_markdown()):
from pyscrappy import WikipediaScraper
with WikipediaScraper() as ws:
result = ws.scrape(query="Python (programming language)", mode="summary")
print(result.data[0]["text"])Each scraper has its own arguments (Wikipedia, stocks, IMDB, news, YouTube, Amazon/Newegg/IKEA, Uber Eats, and more — see the full list). For per-scraper arguments and examples, see the documentation.
From the command line
Scrape a URL straight to a file without writing any code — the output format is inferred from the file extension:
pyscrappy extract https://example.com out.md # clean Markdown
pyscrappy extract https://example.com out.json # structured JSON
pyscrappy extract https://example.com out.txt # extracted page text
pyscrappy extract https://example.com out.html # raw fetched HTML
# Narrow to elements matching a CSS selector, or render JS first:
pyscrappy extract https://example.com items.txt --css-selector ".product"
pyscrappy extract https://example.com page.md --render-jsConfiguration
from pyscrappy import ScraperConfig, GenericScraper
config = ScraperConfig(
timeout=20.0, # request timeout in seconds
max_retries=3, # retry failed requests
rate_limit=2.0, # seconds between requests per domain
proxy="http://...", # proxy URL, or a list to rotate through
scraper_api=None, # route via a scraping-API service (see below)
headless=True, # browser runs headless
render_js="auto", # auto-detect if JS rendering is needed
cache_ttl=0, # response cache TTL in seconds (0 = disabled)
impersonate=None, # e.g. "chrome" to spoof a browser's TLS fingerprint (see below)
)
with GenericScraper(config) as gs:
result = gs.scrape(url="https://example.com")Proxies and blocked sites
Some sites (e.g. eBay, Instagram, Twitter/X, Spotify) block direct automated requests. PyScrappy supports two ways to get through them.
A proxy (or a rotating list) — applies to both the HTTP and browser backends:
from pyscrappy import ScraperConfig, AmazonScraper
# Single proxy
config = ScraperConfig(proxy="http://user:pass@host:port")
# Rotating list (one picked per request)
config = ScraperConfig(proxy=["http://p1:8080", "http://p2:8080"])A scraping-API service (ScraperAPI, ScrapeOps, ScrapingBee) — routes requests through the service, which handles proxies and anti-bot challenges for you:
config = ScraperConfig(scraper_api={
"provider": "scraperapi", # or "scrapeops", "scrapingbee"
"api_key": "YOUR_KEY",
"render_js": True, # optional
})
# Now any scraper works through the service, unchanged:
with AmazonScraper(config) as scraper:
result = scraper.scrape(query="laptop")This is the reliable way to use the scrapers marked "needs proxy" above.
TLS-fingerprint impersonation — many anti-bot systems block a plain HTTP
client by its TLS/JA3 fingerprint before serving any content. Set impersonate
to mimic a real browser's fingerprint and get past that class of block without a
headless browser:
from pyscrappy import ScraperConfig, GenericScraper
# needs the optional extra: pip install 'pyscrappy[stealth]'
config = ScraperConfig(impersonate="chrome") # or "chrome124", "safari", "firefox"
with GenericScraper(config) as gs:
result = gs.scrape("https://example.com")Impersonation currently applies to the synchronous path only; setting it on an async client raises a clear error. All the usual retry, rate-limiting, caching, and robots handling still apply.
Concurrent scraping
Scraping is I/O-bound, so running several scrapes at once parallelizes the
network waits. scrape_many runs one scraper over many inputs; scrape_all
runs a mix of scrapers together. Both preserve input order.
from pyscrappy import scrape_many, scrape_all, AmazonScraper, WikipediaScraper, NewsScraper
# One scraper, many queries, concurrently:
results = scrape_many(AmazonScraper, [{"query": "laptop"}, {"query": "phone"}])
# Different scrapers at once:
results = scrape_all([
lambda: WikipediaScraper().scrape(query="Python"),
lambda: NewsScraper().scrape(feed_url="https://rss.nytimes.com/services/xml/rss/nyt/World.xml"),
])Response caching
Set cache_ttl to a positive number of seconds to cache successful GET
responses. Repeated requests for the same URL (and query params) within the TTL
are served from cache, skipping both the network and the rate limiter. Caching
is disabled by default (cache_ttl=0).
from pyscrappy import WikipediaScraper
from pyscrappy import ScraperConfig
config = ScraperConfig(cache_ttl=300) # cache for 5 minutes
with WikipediaScraper(config) as ws:
ws.scrape(query="Python") # fetched over the network
ws.scrape(query="Python") # served from cacheThe cache is in memory and shared across scraper instances in the same process
(so it also speeds up repeated calls through the MCP server), and is cleared
when the process exits. Call HttpClient.clear_cache() to empty it manually.
Dependencies
Required: httpx, beautifulsoup4, lxml
Optional: playwright (JS rendering), pandas (DataFrames), fastmcp
(MCP server, Python 3.10+)
License
Contributing
All contributions welcome. See Issues.
This package is for educational and research purposes.
Maintenance
Related MCP Servers
- Flicense-qualityDmaintenanceEnables LLMs to extract content from websites using automated static and dynamic scraping engines with built-in anti-bot protections. It provides tools for web data retrieval and stores results in MongoDB with support for JSON and CSV exports.
- AlicenseAqualityAmaintenanceWeb content extraction for AI agents. 10 tools: scrape, crawl, map, batch, extract, summarize, diff, brand, search, research. Uses TLS fingerprinting to bypass anti-bot without a headless browser. Outputs LLM-optimized markdown with 67% fewer tokens than raw HTML.102,130AGPL 3.0
- AlicenseAqualityAmaintenanceWeb scraping, crawling, and structured data extraction for AI agents. 5 tools: scrape (clean markdown from any URL), crawl (entire sites), map (discover URLs), extract (structured JSON), and search. 833ms avg latency, single binary, self-hostable.8571AGPL 3.0
- Alicense-qualityCmaintenanceWeb scraping and search MCP server that wraps Firecrawl API for URL discovery and web search with optional content retrieval.9MIT
Related MCP Connectors
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Turn the web into structured, reliable, actionable enterprise data for AI Agents
Give your agent live data from Twitter, Reddit, the web and GitHub. No API keys, no scraping stack.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/mldsveda/PyScrappy'
If you have feedback or need assistance with the MCP directory API, please join our Discord server