webscout-mcp
webscout-mcp
AI-powered web intelligence platform for AI agents. Search, fetch, crawl, extract, understand, and monitor the web โ with built-in AI, vector search, browser automation, and alerting. Everything stays on your machine.
๐ฏ Project Positioning: webscout-mcp is primarily a Web Search / Fetch MCP server with extensive extension modules. The core MCP server exposes 6 stable tools (search, fetch, crawl, extract, cache stats, cache clear). Additional modules (AI, RAG, browser, monitoring, SEO, etc.) are available as Python libraries and are planned for MCP integration. See Module Status for detailed stability and integration status.
ไธญๆ็ๆฌ็ฎไป | ๅฟซ้ไบ่งฃ้กน็ฎ๏ผ้ๅไธญๆ็จๆท้ ่ฏป
๐ฏ What's Included in MCP (Right Now)
The MCP server currently exposes these 6 core tools:
Tool | Description | Stability |
| Multi-backend web search with result merging | โ Stable |
| Fetch and parse web pages with content extraction | โ Stable |
| Concurrent website crawling with depth limits | ๐ถ Beta |
| Structured content extraction with CSS selectors | โ Stable |
| View cache statistics and hit rates | โ Stable |
| Clear the search/fetch cache | โ Stable |
Available as Python libraries (not yet MCP tools): AI content understanding, vector search & RAG, headless browser automation, web monitoring & alerting, SEO analysis, OCR, PDF processing, knowledge graphs, and more. See Module Status for the full list.
โจ Features
๐ Core Web Tools
Multi-backend search โ Bing, DuckDuckGo, Google, Brave HTML with automatic failover and result merging
Smart content extraction โ trafilatura + readability-lxml + html2text fallback, clean article content
Concurrent crawler โ BFS crawl with depth/page limits, robots.txt compliance, retry on failures
Structured data extraction โ CSS selectors, attributes, regex extraction
Metadata extraction โ JSON-LD, OpenGraph, Twitter Cards, article metadata, images, links
RSS/Atom support โ parse feeds and feed indexes
๐ค AI Content Understanding
Text summarization โ automatic article and page summarization
Question answering โ ask questions about fetched content
Key points extraction โ extract main ideas and takeaways
Content classification โ categorize content into custom categories
Tag generation โ auto-generate relevant tags
Sentiment analysis โ analyze text sentiment
Document comparison โ compare two documents side by side
Entity extraction โ extract people, places, organizations, dates
Multiple LLM backends โ Ollama (local/free), OpenAI, Doubao, custom OpenAI-compatible
๐ง Vector Search & RAG
Semantic search โ search by meaning, not just keywords
RAG (Retrieval-Augmented Generation) โ answer questions based on your crawled content
Local vector database โ ChromaDB persistent storage
Multiple embedding backends โ local sentence-transformers (free), OpenAI, custom
Document chunking โ automatic text splitting with overlap
Similarity threshold โ configurable relevance filtering
๐ Headless Browser Automation
JavaScript rendering โ fetch modern SPAs and dynamic content
User interaction simulation โ scroll, click, fill forms
Screenshot capture โ full-page screenshots
PDF export โ convert web pages to PDF
Login state management โ cookie persistence across sessions
Anti-detection stealth mode โ navigator.webdriver, plugins, languages spoofing
Resource blocking โ block images, media, CSS, fonts for faster loading
Proxy support โ HTTP/HTTPS proxy configuration
Multiple browsers โ Chromium, Firefox, WebKit
๐ก Web Monitoring & Alerting
Scheduled monitoring โ configurable check intervals
Content change detection โ text, HTML, specific element changes
Keyword monitoring โ appearance, disappearance, count changes
Price monitoring โ track price changes with threshold alerts
Change history โ persistent history with diff generation
Multi-channel alerts โ Webhook, Email (SMTP), DingTalk, WeCom
Configurable thresholds โ minimum change size, similarity thresholds
โก Performance & Security
Smart caching โ SQLite cache with TTL, size limits, automatic eviction
Rate limiting โ per-domain token-bucket rate limiting
SSRF protection โ blocks localhost, sensitive ports, invalid schemes
Browser fingerprint rotation โ random User-Agents + realistic headers
TLS fingerprint simulation โ realistic TLS ClientHello fingerprints
Connection pooling โ persistent HTTP connections
Cookie management โ automatic cookie handling and persistence
๐ Easy Setup & Deployment
One-click setup โ
webscout-mcp setupauto-installs all dependenciesSystem detection โ auto-detects OS, CPU, memory, GPU
Smart recommendations โ suggests optimal configuration based on hardware
Docker support โ pre-built images for amd64 and arm64
Docker Compose โ one-command deployment
systemd service โ Linux service file for production
Kubernetes โ deployment manifests for container orchestration
Configuration hot-reload โ reload config without restart
๐ Website Analysis & Optimization
SEO analyzer โ comprehensive SEO audit: meta tags, headings, images, links, URL structure, content length, Open Graph, Twitter Cards, Schema markup, with multi-dimensional scoring and actionable recommendations
Broken link checker โ detect broken links, redirect chains, invalid URLs, mixed content; classify internal/external/mailto/tel/javascript links; detailed reporting with statistics
Performance analyzer โ page performance audit: HTML size, DOM size, resource counts, render-blocking resources, inline CSS/JS, optimization techniques (lazy loading, preconnect, preload), compression/cache detection, performance scoring
Content quality assessor โ readability scores (Flesch-Kincaid, Gunning Fog, SMOG), keyword density, content structure analysis, duplicate content detection, quality scoring
๐ Export & Integration
Multiple export formats โ JSON, CSV, Excel, Parquet, SQLite, Markdown, HTML
Field selection & ordering โ export only specified fields with custom column order
Append mode โ incremental exports for CSV and SQLite
MCP server โ native Model Context Protocol support
CLI interface โ command-line tools for search, fetch, crawl
Python API โ full programmatic access to all features
Sitemap support โ parse sitemap.xml and sitemap indexes
Incremental crawling โ only re-fetch changed pages via ETag/Last-Modified
โ ๏ธ Search Backend Stability Notice
webscout-mcp uses direct HTML scraping for search backends (Bing, DuckDuckGo, Google, Brave) by default โ no API keys required. This makes it free to use, but please be aware of the stability trade-offs:
What can go wrong
DOM changes: Search engines frequently update their HTML structure, which can break scrapers
CAPTCHAs: Automated requests may trigger CAPTCHAs (especially Google and Brave)
Bot detection: Advanced bot detection may block or rate-limit requests
IP blocking: Sustained automated requests can lead to IP bans
Parameter changes: Search engines may change request parameters or headers
Mitigations built in
โ Multi-backend failover: If one backend fails, automatically try the next one
โ Realistic browser headers: Random User-Agents and realistic request headers
โ Rate limiting: Per-domain rate limiting to avoid overwhelming search engines
โ Caching: SQLite cache reduces repeated requests to the same queries
โ Retry with backoff: Exponential backoff on transient failures
For production use
For production workloads requiring higher reliability, consider:
Using official search APIs (Bing Search API, SerpAPI, etc.) โ planned for future releases
Deploying with rotating proxies
Increasing cache TTL to reduce request frequency
Monitoring search backend health and adjusting backends accordingly
Bottom line: webscout-mcp's default search is great for development, personal use, and low-volume workloads. For high-volume production use, plan for additional reliability measures.
๐ฆ Installation
Quick Install
pip install webscout-mcpRequires Python 3.10+.
One-Click Full Setup (Recommended)
# Install core package
pip install webscout-mcp
# Run setup to install all optional dependencies
webscout-mcp setup --playwright --ollama --vector-storeThe setup command will:
Detect your system configuration (OS, CPU, memory, GPU)
Install Playwright and Chromium browser
Install Ollama and download a local LLM (optional)
Install ChromaDB and sentence-transformers for vector search (optional)
Generate a configuration file
Run a health check to verify everything works
Optional Dependencies
# Browser automation (Playwright)
pip install webscout-mcp[browser]
playwright install chromium
# Vector search & RAG (ChromaDB + sentence-transformers)
pip install webscout-mcp[vector]
# AI content understanding (OpenAI client)
pip install webscout-mcp[ai]
# All features
pip install webscout-mcp[all]Docker
docker pull wxslang/webscout-mcp:latest
docker run -p 8000:8000 wxslang/webscout-mcp:latest๐ Quick Start
MCP Client Configuration
Add to your MCP client config:
Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):
{
"mcpServers": {
"webscout": {
"command": "webscout-mcp",
"args": []
}
}
}Cursor (Settings โ MCP โ Add new MCP server):
{
"mcpServers": {
"webscout": {
"command": "webscout-mcp",
"args": []
}
}
}CLI Usage
# Search the web
webscout-mcp search "python web scraping" --max-results 10
# Fetch a page
webscout-mcp fetch https://example.com --extract --format markdown
# Crawl a site
webscout-mcp crawl https://example.com --depth 2 --pages 50
# Run setup
webscout-mcp setup --playwright
# Start MCP server
webscout-mcp servePython API
from webscout_mcp import WebScout
# Initialize
scout = WebScout()
# Search
results = scout.search("AI agents", max_results=5)
for result in results:
print(result.title, result.url)
# Fetch and extract
page = scout.fetch("https://example.com/article")
print(page.title)
print(page.content) # Clean article text
# Crawl
pages = scout.crawl("https://example.com", depth=2, max_pages=20)๐ค AI Content Understanding
from webscout_mcp.ai_processor import AIProcessor, AIConfig
# Use local Ollama (free, no API key needed)
config = AIConfig(backend="ollama", model="qwen2.5:7b")
ai = AIProcessor(config=config)
# Summarize
summary = ai.summarize(page.content, max_length=500)
print(summary.content)
# Ask questions
answer = ai.answer_question(page.content, "What are the main points?")
print(answer.content)
# Extract key points
points = ai.extract_key_points(page.content, num_points=5)
print(points.content)
# Analyze sentiment
sentiment = ai.analyze_sentiment(page.content)
print(sentiment.content)Using OpenAI API:
config = AIConfig(
backend="openai",
model="gpt-4o",
api_key="your-api-key",
)Using Doubao (่ฑๅ ):
config = AIConfig(
backend="doubao",
model="ep-20240101",
api_key="your-api-key",
)๐ง Vector Search & RAG
from webscout_mcp.vector_store import VectorStore, RAGEngine, Document
# Initialize vector store (local, free)
store = VectorStore()
# Add documents
doc = Document.from_text(page.content, source=page.url)
store.add_document(doc)
# Semantic search
results = store.search("how to build AI agents", n_results=5)
for result in results:
print(f"[{result.score:.2f}] {result.document.content[:100]}")
# RAG: Ask questions based on your documents
rag = RAGEngine(vector_store=store)
answer = rag.query("What is the best approach for web scraping?")
print(answer["answer"])
print("Sources:", answer["sources"])๐ Headless Browser Automation
from webscout_mcp.browser_fetcher import BrowserFetcher, BrowserConfig
# Initialize
config = BrowserConfig(headless=True, block_media=True)
browser = BrowserFetcher(config=config)
# Fetch JS-rendered page
result = browser.fetch(
"https://example.com/spa",
wait_for_selector=".content",
scroll_to_bottom=True,
)
print(result.title)
print(result.content)
# Take screenshot
browser.take_screenshot("https://example.com", "screenshot.png", full_page=True)
# Export to PDF
browser.export_pdf("https://example.com", "page.pdf")
# Click element
result = browser.click_element("https://example.com", "button.load-more")
# Fill form
result = browser.fill_form(
"https://example.com/login",
{"#username": "user", "#password": "pass"},
submit_selector="button[type=submit]",
)
browser.close()๐ก Web Monitoring & Alerting
from webscout_mcp.monitor import WebMonitor, MonitorConfig, WebhookAlert, EmailAlert
# Initialize
config = MonitorConfig(check_interval=300, min_change_size=10)
monitor = WebMonitor(config=config)
# Add alert channels
monitor.add_alert_channel(WebhookAlert("https://hooks.example.com/alert"))
monitor.add_alert_channel(EmailAlert(
smtp_server="smtp.gmail.com",
smtp_port=587,
username="you@gmail.com",
password="app-password",
from_addr="you@gmail.com",
to_addrs=["recipient@example.com"],
))
# Check for changes
changes = monitor.check_url("https://example.com/pricing")
for change in changes:
print(f"{change.change_type}: {change.old_value} -> {change.new_value}")
# Get history
history = monitor.get_history("https://example.com/pricing")๐ Website Analysis & Optimization
SEO Analysis
from webscout_mcp.seo_analyzer import SEOAnalyzer
# Initialize
analyzer = SEOAnalyzer()
# Analyze a page
html = "<html>...</html>"
metrics = analyzer.analyze(html, url="https://example.com")
# Check scores
print(f"Overall SEO Score: {metrics.overall_score}/100")
print(f"Meta Score: {metrics.meta_score}")
print(f"Heading Score: {metrics.heading_score}")
print(f"Image Score: {metrics.image_score}")
# Check issues and recommendations
print("Issues:", metrics.issues)
print("Recommendations:", metrics.recommendations)Broken Link Checking
from webscout_mcp.broken_link_checker import BrokenLinkChecker
# Initialize
checker = BrokenLinkChecker(timeout=10.0, max_redirects=5)
# Check all links on a page
html = "<html>...</html>"
report = checker.check_page(html, base_url="https://example.com")
# Check statistics
print(f"Total links: {report.total_links}")
print(f"OK: {report.ok_links}")
print(f"Broken: {report.broken_links}")
print(f"Redirects: {report.redirect_links}")
print(f"Broken percentage: {report.broken_link_percentage}%")
# Get only broken links
broken = checker.get_broken_links(report)
for link in broken:
print(f"[{link.status}] {link.url} - {link.error_message}")
# Generate human-readable summary
print(checker.generate_summary(report))Performance Analysis
from webscout_mcp.performance_analyzer import PerformanceAnalyzer
# Initialize
analyzer = PerformanceAnalyzer()
# Analyze page performance
html = "<html>...</html>"
headers = {"Content-Encoding": "gzip", "Cache-Control": "max-age=3600"}
metrics = analyzer.analyze(html, url="https://example.com", response_headers=headers)
# Check scores
print(f"Overall Performance Score: {metrics.overall_score}/100")
print(f"HTML Size: {metrics.html_size_kb}KB (score: {metrics.html_size_score})")
print(f"DOM Nodes: {metrics.dom_node_count} (score: {metrics.dom_size_score})")
print(f"Requests: {metrics.request_count} (score: {metrics.request_count_score})")
# Check optimization techniques
print(f"Has gzip: {metrics.has_gzip}")
print(f"Has brotli: {metrics.has_brotli}")
print(f"Has lazy loading: {metrics.has_lazy_loading}")
print(f"Has preconnect: {metrics.has_preconnect}")
# Check issues and recommendations
print("Issues:", metrics.issues)
print("Warnings:", metrics.warnings)
print("Recommendations:", metrics.recommendations)Enhanced Data Export
from webscout_mcp.data_exporter import DataExporter, ExportConfig
# Sample data
data = [
{"title": "Result 1", "url": "https://example.com/1", "score": 0.95},
{"title": "Result 2", "url": "https://example.com/2", "score": 0.85},
]
# Export to JSON
config = ExportConfig(format="json", output_path="results.json", pretty_json=True)
exporter = DataExporter(config=config)
result = exporter.export(data)
print(f"Exported {result.record_count} records to {result.output_path}")
# Export to CSV with field selection
config = ExportConfig(
format="csv",
output_path="results.csv",
fields=["title", "url"], # Only export these fields
csv_delimiter=",",
)
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to Excel
config = ExportConfig(format="excel", output_path="results.xlsx", excel_sheet_name="Results")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to SQLite
config = ExportConfig(format="sqlite", output_path="results.db", sqlite_table_name="search_results")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to Parquet (columnar storage)
config = ExportConfig(format="parquet", output_path="results.parquet")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to Markdown
config = ExportConfig(format="markdown", output_path="results.md")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Export to HTML
config = ExportConfig(format="html", output_path="results.html")
exporter = DataExporter(config=config)
result = exporter.export(data)
# Using convenience function
from webscout_mcp.data_exporter import export_data
result = export_data(data, "results.json", export_format="json", fields=["title", "url"])โ๏ธ Configuration
Environment Variables
# Core
WEBSCOUT_CACHE_ENABLED=true
WEBSCOUT_CACHE_TTL=3600
WEBSCOUT_RATE_LIMIT_ENABLED=true
# Search
WEBSCOUT_SEARCH_DEFAULT_BACKEND=bing
WEBSCOUT_SEARCH_MAX_RESULTS=10
# AI
WEBSCOUT_AI_BACKEND=ollama
WEBSCOUT_AI_MODEL=qwen2.5:7b
WEBSCOUT_AI_API_KEY=your-key
# Vector Store
WEBSCOUT_VECTOR_DB=chroma
WEBSCOUT_EMBEDDING_BACKEND=local
WEBSCOUT_EMBEDDING_MODEL=BAAI/bge-small-zh-v1.5
# Browser
WEBSCOUT_BROWSER_TYPE=chromium
WEBSCOUT_BROWSER_HEADLESS=true
WEBSCOUT_BROWSER_BLOCK_MEDIA=true
# Monitor
WEBSCOUT_MONITOR_INTERVAL=300
WEBSCOUT_MONITOR_MIN_CHANGE=10
# Setup
WEBSCOUT_SETUP_PLAYWRIGHT=true
WEBSCOUT_SETUP_OLLAMA=false
WEBSCOUT_SETUP_CHROMADB=falseConfig File
Create ~/.config/webscout/config.toml:
[server]
host = "127.0.0.1"
port = 8000
[cache]
enabled = true
ttl = 3600
[search]
default_backend = "bing"
max_results = 10
[ai]
backend = "ollama"
model = "qwen2.5:7b"
[vector_store]
enabled = true
vector_db = "chroma"
embedding_backend = "local"
[browser]
enabled = true
headless = true
[monitor]
check_interval = 300๐ Documentation
README โ This file
Module Status โ Module stability levels and MCP integration status
Project Introduction โ Detailed project overview and architecture
Deployment Guide โ Docker, systemd, Kubernetes deployment
Examples โ Usage examples and sample code
CHANGELOG โ Version history
CONTRIBUTING โ Contributing guidelines
CODE OF CONDUCT โ Community code of conduct
SECURITY โ Security policy and vulnerability reporting
๐งช Testing
# Install dev dependencies
pip install webscout-mcp[dev]
# Run all tests
pytest tests/
# Run with coverage
pytest tests/ --cov=webscout_mcp --cov-report=htmlTest coverage: 395+ tests covering all modules.
๐ค Contributing
Contributions are welcome! Please see CONTRIBUTING.md for guidelines.
Fork the repository
Create a feature branch
Make your changes
Add tests
Submit a pull request
๐ License
MIT License โ see LICENSE for details.
๐ Acknowledgments
trafilatura โ content extraction
readability-lxml โ readability fallback
Playwright โ browser automation
ChromaDB โ vector database
sentence-transformers โ text embeddings
Ollama โ local LLM runtime
httpx โ HTTP client
BeautifulSoup โ HTML parsing
Built with โค๏ธ for the AI agent community.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/wxs-lang/webscout-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server