clean-search-mcp
Integrates DuckDuckGo as a search engine source to retrieve clean, spam-free search results filtered by domain blocklists and content quality scoring.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@clean-search-mcpsearch for best practices in React hooks"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
clean-search-mcp ๐งน
A lightweight MCP (Model Context Protocol) service that provides clean, spam-free search results for AI agents. Filters out content farms, malware sites, SEO garbage, and low-quality content before they reach your LLM.
Features
Three search engines โ Yandex + Bing + DuckDuckGo auto-fallback
176K+ domain blocklist โ auto-updated from 25+ community sources, covers malware/scam/ads/trackers/content farms
Three-layer filtering โ domain blacklist โ content rules โ quality scoring
Content extraction โ full page text via trafilatura + selectolax
Result scoring โ 0-1 quality score (official docs 0.8 > tutorials 0.6 > garbage 0)
LRU cache โ search cache 6h, fetch cache 24h, auto-cleanup
User blacklist โ add domains on the fly, report bad results
Deep mode โ optional Playwright fallback for JS-heavy pages
No heavy dependencies โ pure HTTP, no browser required
Related MCP server: evo-scry
Quick Start
pip install -r requirements.txt
python main.pyMCP Client Config
{
"mcpServers": {
"clean-search": {
"command": "python",
"args": ["/path/to/clean_search_mcp/main.py"]
}
}
}Test Locally
python test_local.py "your search query" -n 5
python test_local.py "your query" -n 5 --no-content # skip page content
python test_local.py "your query" --deep # use Playwright fallbackAPI
clean_search(query, max_results=5, with_content=True, deep_mode=False)
Param | Default | Description |
| required | Search query |
| 5 | Results to return (max 10) |
| True | Include extracted page text |
| False | Use Playwright fallback for JS pages |
Returns [{title, url, snippet, content, score}] sorted by quality.
add_user_blacklist(domain)
Add a domain to personal blocklist.
report_bad_result(url)
Report a low-quality URL (domain auto-blocked).
Configuration
Edit config.py to tune:
Search providers: enable/disable Yandex, Bing, DuckDuckGo
Blacklist sources: add or remove community blocklist URLs
Scoring weights: adjust domain authority, content quality bonuses
Caching: TTL, max files, cleanup interval
Proxy: set
PROXYfor HTTP/Playwright
Dependencies
mcp, httpx, selectolax, trafilatura, duckduckgo-searchAll lightweight pip packages. Playwright is optional (deep mode only).
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
Scrape, crawl and search the web for AI agents via MCP.
Live AI-native web search with citations. One tool for every MCP client. Flat per-request pricing.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceA read-only MCP server that gives AI agents the web as compact, ranked, verified evidence โ no API keys, no cloud retrieval, all models local.21 npm1MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.MIT
- AlicenseAqualityBmaintenanceEnables AI agents to perform multi-engine web search, fetch web pages, and extract clean Markdown content via MCP, with no API keys required.38MIT
- AlicenseNot gradedqualityCmaintenanceA self-hosted, MCP-native web-search backend for AI agents that provides meta-search, clean extraction, RAG with citations, and GitHub project selection.3MIT