webscout-mcp
# webscout-mcp
[](https://pypi.org/project/webscout-mcp/)
[](https://pypi.org/project/webscout-mcp/)
[](https://github.com/wxs-lang/webscout-mcp/actions/workflows/tests.yml)
[](https://github.com/wxs-lang/webscout-mcp/actions/workflows/quality.yml)
[](https://wxs-lang.github.io/webscout-mcp/)
[](https://hub.docker.com/r/wxslang/webscout-mcp)
[](https://github.com/wxs-lang/webscout-mcp/blob/main/LICENSE)
[](https://github.com/wxs-lang/webscout-mcp/stargazers)
[](https://github.com/wxs-lang/webscout-mcp/network/members)
[](https://github.com/wxs-lang/webscout-mcp/issues)
[](https://github.com/wxs-lang/webscout-mcp/commits/main)
[](https://github.com/wxs-lang/webscout-mcp/commits/main)
Self-healing web access layer for AI Agents. Search, fetch, extract, and read long web content progressively โ with provider routing, deterministic recovery, browser fallback, caching, and observability over the Model Context Protocol (MCP). Everything stays on your machine.
> **๐ฏ Project Positioning**: webscout-mcp is a **self-healing web access layer for AI agents**. The core MCP server exposes **11 tools**: `web_search`, `web_fetch`, `web_crawl`, `web_extract`, `metadata_extract`, `rss_parse`, `content_quality`, `broken_links`, `cache_stats`, `cache_clear`, and `search_health`. AI content understanding, vector search/RAG, headless browser automation, monitoring, SEO, and the optional Jev shadow evaluator are available as Python libraries / optional extras and are **not** required for core use. See [Module Status](MODULE_STATUS.md) for detailed stability and integration status.
[**ไธญๆ็ๆฌ็ฎไป**](README_zh.md) | ๅฟซ้ไบ่งฃ้กน็ฎ๏ผ้ๅไธญๆ็จๆท้
่ฏป
## ๐ฏ What's Included in MCP (Right Now)
The MCP server currently exposes these **11 tools**:
| Tool | Description | Stability |
|------|-------------|-----------|
| `web_search` | Self-healing multi-provider search with deterministic recovery and automatic fallback | โ
Stable |
| `web_fetch` | Fetch a URL and return one window of extracted main content, with progressive continuation (`start_char`) | โ
Stable |
| `web_crawl` | Concurrent website crawling with depth limits | ๐ถ Beta |
| `web_extract` | Structured content extraction with CSS selectors | โ
Stable |
| `metadata_extract` | Extract metadata (JSON-LD, OpenGraph, Twitter cards) from a page | ๐ถ Beta |
| `rss_parse` | Parse an RSS or Atom feed and return its entries | ๐ถ Beta |
| `content_quality` | Analyze content quality of a fetched page | ๐งช Experimental |
| `broken_links` | Check for broken links on a web page | ๐งช Experimental |
| `cache_stats` | View cache statistics and hit rates | โ
Stable |
| `cache_clear` | Clear the search/fetch cache | โ
Stable |
| `search_health` | Health report for all search backends | ๐ถ Beta |
**Available as Python libraries / optional extras (not required for core use)**: AI content understanding, vector search & RAG, headless browser automation, web monitoring & alerting, SEO analysis, OCR, PDF processing, knowledge graphs, and the optional **Jev** shadow evaluator (`pip install "webscout-mcp[jev]"`, disabled by default, BYOK). See [Module Status](MODULE_STATUS.md) for the full list.
## โจ Features
### ๐ Core Web Tools
- **Self-healing multi-provider search** โ Bing, DuckDuckGo, SearXNG, Tavily with health-based ranking, deterministic recovery, and automatic fallback
- **Smart content extraction** โ trafilatura + readability-lxml + html2text fallback, clean article content
- **Concurrent crawler** โ BFS crawl with depth/page limits, robots.txt compliance, retry on failures
- **Structured data extraction** โ CSS selectors, attributes, regex extraction
- **Metadata extraction** โ JSON-LD, OpenGraph, Twitter Cards, article metadata, images, links
- **RSS/Atom support** โ parse feeds and feed indexes
### ๐ค AI Content Understanding
- **Text summarization** โ automatic article and page summarization
- **Question answering** โ ask questions about fetched content
- **Key points extraction** โ extract main ideas and takeaways
- **Content classification** โ categorize content into custom categories
- **Tag generation** โ auto-generate relevant tags
- **Sentiment analysis** โ analyze text sentiment
- **Document comparison** โ compare two documents side by side
- **Entity extraction** โ extract people, places, organizations, dates
- **Multiple LLM backends** โ Ollama (local/free), OpenAI, Doubao, custom OpenAI-compatible
### ๐ง Vector Search & RAG
- **Semantic search** โ search by meaning, not just keywords
- **RAG (Retrieval-Augmented Generation)** โ answer questions based on your crawled content
- **Local vector database** โ ChromaDB persistent storage
- **Multiple embedding backends** โ local sentence-transformers (free), OpenAI, custom
- **Document chunking** โ automatic text splitting with overlap
- **Similarity threshold** โ configurable relevance filtering
### ๐ Headless Browser Automation
- **JavaScript rendering** โ fetch modern SPAs and dynamic content
- **User interaction simulation** โ scroll, click, fill forms
- **Screenshot capture** โ full-page screenshots
- **PDF export** โ convert web pages to PDF
- **Login state management** โ cookie persistence across sessions
- **Anti-detection stealth mode** โ navigator.webdriver, plugins, languages spoofing
- **Resource blocking** โ block images, media, CSS, fonts for faster loading
- **Proxy support** โ HTTP/HTTPS proxy configuration
- **Multiple browsers** โ Chromium, Firefox, WebKit
### ๐ก Web Monitoring & Alerting
- **Scheduled monitoring** โ configurable check intervals
- **Content change detection** โ text, HTML, specific element changes
- **Keyword monitoring** โ appearance, disappearance, count changes
- **Price monitoring** โ track price changes with threshold alerts
- **Change history** โ persistent history with diff generation
- **Multi-channel alerts** โ Webhook, Email (SMTP), DingTalk, WeCom
- **Configurable thresholds** โ minimum change size, similarity thresholds
### โก Performance & Security
- **Smart caching** โ SQLite cache with TTL, size limits, automatic eviction
- **Rate limiting** โ per-domain token-bucket rate limiting
- **SSRF protection** โ blocks localhost, sensitive ports, invalid schemes
- **Browser fingerprint rotation** โ random User-Agents + realistic headers
- **TLS fingerprint simulation** โ realistic TLS ClientHello fingerprints
- **Connection pooling** โ persistent HTTP connections
- **Cookie management** โ automatic cookie handling and persistence
### ๐ Easy Setup & Deployment
- **One-click setup** โ `webscout-mcp setup` auto-installs all dependencies
- **System detection** โ auto-detects OS, CPU, memory, GPU
- **Smart recommendations** โ suggests optimal configuration based on hardware
- **Docker support** โ pre-built images for amd64 and arm64
- **Docker Compose** โ one-command deployment
- **systemd service** โ Linux service file for production
- **Kubernetes** โ deployment manifests for container orchestration
- **Configuration hot-reload** โ reload config without restart
### ๐ Website Analysis & Optimization
- **SEO analyzer** โ comprehensive SEO audit: meta tags, headings, images, links, URL structure, content length, Open Graph, Twitter Cards, Schema markup, with multi-dimensional scoring and actionable recommendations
- **Broken link checker** โ detect broken links, redirect chains, invalid URLs, mixed content; classify internal/external/mailto/tel/javascript links; detailed reporting with statistics
- **Performance analyzer** โ page performance audit: HTML size, DOM size, resource counts, render-blocking resources, inline CSS/JS, optimization techniques (lazy loading, preconnect, preload), compression/cache detection, performance scoring
- **Content quality assessor** โ readability scores (Flesch-Kincaid, Gunning Fog, SMOG), keyword density, content structure analysis, duplicate content detection, quality scoring
### ๐ Export & Integration
- **Multiple export formats** โ JSON, CSV, Excel, Parquet, SQLite, Markdown, HTML
- **Field selection & ordering** โ export only specified fields with custom column order
- **Append mode** โ incremental exports for CSV and SQLite
- **MCP server** โ native Model Context Protocol support
- **CLI interface** โ command-line tools for search, fetch, crawl
- **Python API** โ full programmatic access to all features
- **Sitemap support** โ parse sitemap.xml and sitemap indexes
- **Incremental crawling** โ only re-fetch changed pages via ETag/Last-Modified
## โ ๏ธ Search Backend Stability Notice
webscout-mcp uses **direct HTML scraping** for search backends (Bing, DuckDuckGo, Google, Brave) by default โ no API keys required. This makes it **free to use**, but please be aware of the stability trade-offs:
### What can go wrong
- **DOM changes**: Search engines frequently update their HTML structure, which can break scrapers
- **CAPTCHAs**: Automated requests may trigger CAPTCHAs (especially Google and Brave)
- **Bot detection**: Advanced bot detection may block or rate-limit requests
- **IP blocking**: Sustained automated requests can lead to IP bans
- **Parameter changes**: Search engines may change request parameters or headers
### Mitigations built in
- โ
**Multi-backend failover**: If one backend fails, automatically try the next one
- โ
**Realistic browser headers**: Random User-Agents and realistic request headers
- โ
**Rate limiting**: Per-domain rate limiting to avoid overwhelming search engines
- โ
**Caching**: SQLite cache reduces repeated requests to the same queries
- โ
**Retry with backoff**: Exponential backoff on transient failures
### For production use
For production workloads requiring **higher reliability**, consider:
1. Using official search APIs (Bing Search API, SerpAPI, etc.) โ planned for future releases
2. Deploying with rotating proxies
3. Increasing cache TTL to reduce request frequency
4. Monitoring search backend health and adjusting backends accordingly
**Bottom line**: webscout-mcp's default search is **great for development, personal use, and low-volume workloads**. For high-volume production use, plan for additional reliability measures.
## ๐ฆ Installation
### Quick Install
```bash
pip install webscout-mcp
```
Requires Python 3.10+.
### One-Click Full Setup (Recommended)
```bash
# Install core package
pip install webscout-mcp
# Run setup to install all optional dependencies
webscout-mcp setup --playwright --ollama --vector-store
```
The setup command will:
- Detect your system configuration (OS, CPU, memory, GPU)
- Install Playwright and Chromium browser
- Install Ollama and download a local LLM (optional)
- Install ChromaDB and sentence-transformers for vector search (optional)
- Generate a configuration file
- Run a health check to verify everything works
### Optional Dependencies
```bash
# Browser automation (Playwright)
pip install webscout-mcp[browser]
playwright install chromium
# Vector search & RAG (ChromaDB + sentence-transformers)
pip install webscout-mcp[vector]
# AI content understanding (OpenAI client)
pip install webscout-mcp[ai]
# All features
pip install webscout-mcp[all]
```
### Docker
```bash
docker pull wxslang/webscout-mcp:latest
docker run -p 8000:8000 wxslang/webscout-mcp:latest
```
## ๐ Quick Start
### MCP Client Configuration
Add to your MCP client config:
**Claude Desktop** (`~/Library/Application Support/Claude/claude_desktop_config.json` on macOS, `%APPDATA%\Claude\claude_desktop_config.json` on Windows):
```json
{
"mcpServers": {
"webscout": {
"command": "webscout-mcp",
"args": []
}
}
}
```
**Cursor** (Settings โ MCP โ Add new MCP server):
```json
{
"mcpServers": {
"webscout": {
"command": "webscout-mcp",
"args": []
}
}
}
```
### CLI Usage
```bash
# Search the web
webscout-mcp search "python web scraping" --max-results 10
# Fetch a page
webscout-mcp fetch https://example.com --extract --format markdown
# Crawl a site
webscout-mcp crawl https://example.com --depth 2 --pages 50
# Run setup
webscout-mcp setup --playwright
# Start MCP server
webscout-mcp serve
```
### Python API
```python
from webscout_mcp import WebScout
# Initialize
scout = WebScout()
# Search
results = scout.search("AI agents", max_results=5)
for result in results:
print(result.title, result.url)
# Fetch and extract
page = scout.fetch("https://example.com/article")
print(page.title)
print(page.content) # Clean article text
# Crawl
pages = scout.crawl("https://example.com", depth=2, max_pages=20)
```
## ๐ค AI Content Understanding
```python
from webscout_mcp.ai_processor import AIProcessor, AIConfig
# Use local Ollama (free, no API key needed)
config = AIConfig(backend="ollama", model="qwen2.5:7b")
ai = AIProcessor(config=config)
# Summarize
summary = ai.summarize(page.content, max_length=500)
print(summary.content)
# Ask questions
answer = ai.answer_question(page.content, "What are the main points?")
print(answer.content)
# Extract key points
points = ai.extract_key_points(page.content, num_points=5)
print(points.content)
# Analyze sentiment
sentiment = ai.analyze_sentiment(page.content)
print(sentiment.content)
```
**Using OpenAI API:**
```python
config = AIConfig(
backend="openai",
model="gpt-4o",
api_key="your-api-key",
)
```
**Using Doubao (่ฑๅ
):**
```python
config = AIConfig(
backend="doubao",
model="ep-20240101",
api_key="your-api-key",
)
```
## ๐ง Vector Search & RAG
```python
from webscout_mcp.vector_store import VectorStore, RAGEngine, Document
# Initialize vector store (local, free)
store = VectorStore()
# Add documents
doc = Document.from_text(page.content, source=page.url)
store.add_document(doc)
# Semantic search
results = store.search("how to build AI agents", n_results=5)
for result in results:
print(f"[{result.score:.2f}] {result.document.content[:100]}")
# RAG: Ask questions based on your documents
rag = RAGEngine(vector_store=store)
answer = rag.query("What is the best approach for web scraping?")
print(answer["answer"])
print("Sources:", answer["sources"])
```
## ๐ Headless Browser Automation
```python
from webscout_mcp.browser_fetcher import BrowserFetcher, BrowserConfig
# Initialize
config = BrowserConfig(headless=True, block_media=True)
browser = BrowserFetcher(config=config)
# Fetch JS-rendered page
result = browser.fetch(
"https://example.com/spa",
wait_for_selector=".content",
scroll_to_bottom=True,
)
print(result.title)
print(result.content)
# Take screenshot
browser.take_screenshot("https://example.com", "screenshot.png", full_page=True)
# Export to PDF
browser.export_pdf("https://example.com", "page.pdf")
# Click element
result = browser.click_element("https://example.com", "button.load-more")
# Fill form
result = browser.fill_form(
"https://example.com/login",
{"#username": "user", "#password": "pass"},
submit_selector="button[type=submit]",
)
browser.close()
```
## ๐ก Web Monitoring & Alerting
```python
from webscout_mcp.monitor import WebMonitor, MonitorConfig, WebhookAlert, EmailAlert
# Initialize
config = MonitorConfig(check_interval=300, min_change_size=10)
monitor = WebMonitor(config=config)
# Add alert channels
monitor.add_alert_channel(WebhookAlert("https://hooks.example.com/alert"))
monitor.add_alert_channel(EmailAlert(
smtp_server="smtp.gmail.com",
smtp_port=587,
username="you@gmail.com",
password="app-password",
from_addr="you@gmail.com",
to_addrs=["recipient@example.com"],
))
# Check for changes
changes = monitor.check_url("https://example.com/pricing")
for change in changes:
print(f"{change.change_type}: {change.old_value} -> {change.new_value}")
# Get history
history = monitor.get_history("https://example.com/pricing")
```
## ๐ Website Analysis & Optimization
### SEO Analysis
```python
from webscout_mcp.seo_analyzer import SEOAnalyzer
# Initialize
analyzer = SEOAnalyzer()
# Analyze a page
html = "<html>...</html>"
metrics = analyzer.analyze(html, url="https://example.com")
# Check scores
print(f"Overall SEO Score: {metrics.overall_score}/100")
print(f"Meta Score: {metrics.meta_score}")
print(f"Heading Score: {metrics.heading_score}")
print(f"Image Score: {metrics.image_score}")
# Check issues and recommendations
print("Issues:", metrics.issues)
print("Recommendations:", metrics.recommendations)
```
### Broken Link Checking
```python
from webscout_mcp.broken_link_checker import BrokenLinkChecker
# Initialize
checker = BrokenLinkChecker(timeout=10.0, max_redirects=5)
# Check all links on a page
html = "<html>...</html>"
report = checker.check_page(html, base_url="https://example.com")
# Check statistics
print(f"Total links: {report.total_links}")
print(f"OK: {report.ok_links}")
print(f"Broken: {report.broken_links}")
print(f"Redirects: {report.redirect_links}")
print(f"Broken percentage: {report.broken_link_percentage}%")
# Get only broken links
broken = checker.get_broken_links(report)
for link in broken:
print(f"[{link.status}] {link.url} - {link.error_message}")
# Generate human-readable summary
print(checker.generate_summary(report))
```
### Performance Analysis
```python
from webscout_mcp.performance_analyzer import PerformanceAnalyzer
# Initialize
analyzer = PerformanceAnalyzer()
# Analyze page performance
html = "<html>...</html>"
headers = {"Content-Encoding": "gzip", "Cache-Control": "max-age=3600"}
metrics = analyzer.analyze(html, url="https://example.com", response_headers=headers)
# Check scores
print(f"Overall Performance Score: {metrics.overall_score}/100")
print(f"HTML Size: {metrics.html_size_kb}KB (score: {metrics.html_size_score})")
print(f"DOM Nodes: {metrics.dom_node_count} (score: {metrics.dom_size_score})")
print(f"Requests: {metrics.request_count} (score: {metrics.request_count_score})")
# Check optimization techniques
print(f"Has gzip: {metrics.has_gzip}")
print(f"Has brotli: {metrics.has_brotli}")
print(f"Has lazy loading: {metrics.has_lazy_loading}")
print(f"Has preconnect: {metrics.has_preconnect}")
# Check issues and recommendations
print("Issues:", metrics.issues)
print("Warnings:", metrics.warnings)
print("Recommendations:", metrics.recommendations)
```
### Enhanced Data Export
```python
from webscout_mcp.data_exporter import DataExporter, ExportConfig
# Sample data
data = [
{"title": "Result 1", "url": "https://example.com/1", "score": 0.95},
{"title": "Result 2", "url": "https://example.com/2", "score": 0.85},
]
# Export to JSON
config = ExportConfig(format="json", output_path="results.json", pretty_json=True)
exporter = DataExporter(config=config)
result = exporter.export(data)
print(f"Exported {result.record_count} records to {result.output_path}")
# Export to CSV with field selection
config = ExportConfig(
format="csv",
output_path="results.csv",
fields=["title", "url"], # Only export these fields
csv_delimiter=",",
)
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to Excel
config = ExportConfig(format="excel", output_path="results.xlsx", excel_sheet_name="Results")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to SQLite
config = ExportConfig(format="sqlite", output_path="results.db", sqlite_table_name="search_results")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to Parquet (columnar storage)
config = ExportConfig(format="parquet", output_path="results.parquet")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to Markdown
config = ExportConfig(format="markdown", output_path="results.md")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to HTML
config = ExportConfig(format="html", output_path="results.html")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Using convenience function
from webscout_mcp.data_exporter import export_data
result = export_data(data, "results.json", export_format="json", fields=["title", "url"])
```
## โ๏ธ Configuration
### Environment Variables
```bash
# Core
WEBSCOUT_CACHE_ENABLED=true
WEBSCOUT_CACHE_TTL=3600
WEBSCOUT_RATE_LIMIT_ENABLED=true
# Search
WEBSCOUT_SEARCH_DEFAULT_BACKEND=bing
WEBSCOUT_SEARCH_MAX_RESULTS=10
# AI
WEBSCOUT_AI_BACKEND=ollama
WEBSCOUT_AI_MODEL=qwen2.5:7b
WEBSCOUT_AI_API_KEY=your-key
# Vector Store
WEBSCOUT_VECTOR_DB=chroma
WEBSCOUT_EMBEDDING_BACKEND=local
WEBSCOUT_EMBEDDING_MODEL=BAAI/bge-small-zh-v1.5
# Browser
WEBSCOUT_BROWSER_TYPE=chromium
WEBSCOUT_BROWSER_HEADLESS=true
WEBSCOUT_BROWSER_BLOCK_MEDIA=true
# Monitor
WEBSCOUT_MONITOR_INTERVAL=300
WEBSCOUT_MONITOR_MIN_CHANGE=10
# Setup
WEBSCOUT_SETUP_PLAYWRIGHT=true
WEBSCOUT_SETUP_OLLAMA=false
WEBSCOUT_SETUP_CHROMADB=false
```
### Config File
Create `~/.config/webscout/config.toml`:
```toml
[server]
host = "127.0.0.1"
port = 8000
[cache]
enabled = true
ttl = 3600
[search]
default_backend = "bing"
max_results = 10
[ai]
backend = "ollama"
model = "qwen2.5:7b"
[vector_store]
enabled = true
vector_db = "chroma"
embedding_backend = "local"
[browser]
enabled = true
headless = true
[monitor]
check_interval = 300
```
## ๐ Documentation
- [README](README.md) โ This file
- [Module Status](MODULE_STATUS.md) โ Module stability levels and MCP integration status
- [Project Introduction](PROJECT_INTRODUCTION.md) โ Detailed project overview and architecture
- [Deployment Guide](DEPLOYMENT.md) โ Docker, systemd, Kubernetes deployment
- [Examples](examples/) โ Usage examples and sample code
- [CHANGELOG](CHANGELOG.md) โ Version history
- [CONTRIBUTING](CONTRIBUTING.md) โ Contributing guidelines
- [CODE OF CONDUCT](CODE_OF_CONDUCT.md) โ Community code of conduct
- [SECURITY](SECURITY.md) โ Security policy and vulnerability reporting
## ๐งช Testing
```bash
# Install dev dependencies
pip install webscout-mcp[dev]
# Run all tests
pytest tests/
# Run with coverage
pytest tests/ --cov=webscout_mcp --cov-report=html
```
Test coverage: **395+ tests** covering all modules.
## ๐ค Contributing
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Add tests
5. Submit a pull request
## ๐ License
MIT License โ see [LICENSE](LICENSE) for details.
## ๐ Acknowledgments
- [trafilatura](https://github.com/adbar/trafilatura) โ content extraction
- [readability-lxml](https://github.com/buriy/python-readability) โ readability fallback
- [Playwright](https://playwright.dev/) โ browser automation
- [ChromaDB](https://www.trychroma.com/) โ vector database
- [sentence-transformers](https://www.sbert.net/) โ text embeddings
- [Ollama](https://ollama.com/) โ local LLM runtime
- [httpx](https://www.python-httpx.org/) โ HTTP client
- [BeautifulSoup](https://www.crummy.com/software/BeautifulSoup/) โ HTML parsing
---
**Built with โค๏ธ for the AI agent community.**
TDQS
Scored across 11 tools
Most tools have clearly distinct purposes, e.g., web_search, web_fetch, web_crawl, rss_parse. A small amount of overlap exists between web_extract and metadata_extract, but their descriptions clarify the difference between CSS-selector extraction and metadata extraction.
Names are readable and mostly use snake_case, but the pattern is mixed: some are object_verb (web_fetch, rss_parse, cache_clear), while others are noun phrases (broken_links, content_quality, search_health). The web_ prefix helps, but there is no consistent verb_noun convention.
11 tools is well within the ideal range and each tool serves a distinct web-research purpose, from searching and fetching to crawling, parsing, extracting, and cache management. No tool feels redundant or unnecessary.
The tool surface covers the core web-scouting workflow well: search, fetch, crawl, parse RSS, extract structured data and metadata, check links, and assess quality. Minor gaps exist, such as no search pagination or browser-level automation, but agents can generally complete tasks without dead ends.