Web Explorer MCP
Provides web search capabilities using a local SearxNG instance for private, API-key-free search.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Web Explorer MCPsearch for latest privacy news"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Web Explorer MCP
A Model Context Protocol (MCP) server that provides web search and webpage content extraction using a local SearxNG instance.
Why Web Explorer MCP?
Unlike commercial solutions (GitHub Copilot, Cursor IDE), Web Explorer MCP prioritizes privacy and autonomy:
Feature | Web Explorer MCP | GitHub Copilot | Cursor IDE |
Privacy | ✅ Local SearxNG, zero tracking | ❌ Bing API, Microsoft servers | ❌ Cloud search, third-party APIs |
Cost | ✅ Free, no limits | 💰 $10-20/month subscription | 💰 $20/month Pro plan |
API Keys | ✅ None required | ⚠️ GitHub account required | ⚠️ Account & subscription |
Data Control | ✅ All data stays local | ❌ Queries sent to Microsoft | ❌ Queries sent to external services |
Setup | ✅ 2 commands | ⚠️ Account setup, policy config | ⚠️ Account, payment setup |
Open Source | ✅ Fully auditable | ⚠️ Partial (client only) | ❌ Proprietary |
Perfect for: Developers who value privacy, work with sensitive data, or prefer not to depend on external services and subscriptions.
Related MCP server: mcp-searxng
⚠️ Responsible Use
This tool is designed for human-assisted AI interactions, not for automated high-volume scraping:
🚫 Not for DDoS - Do not use for overwhelming websites or search engines
🚫 Not for High-Speed Automation - Avoid usage speeds significantly higher than a real user
🚫 Not for Fully Automated AI Agents - Not recommended for high-performance autonomous agents
✅ Respect Infrastructure - Honor website owners' business scenarios and infrastructure capabilities
✅ Follow robots.txt - Respect crawling policies and rate limits
Use responsibly: This tool is meant for legitimate research and development, not for abuse.
Features
🔍 Web Search - Search using local SearxNG (private, no API keys)
📄 Content Extraction - Extract clean text from webpages with Playwright rendering
🐳 Zero Pollution - Runs in Docker, leaves no traces
🚀 Simple Setup - Install in 2 commands
Quick Start
1. Install Services (SearxNG + Playwright)
git clone https://github.com/l0kifs/web-explorer-mcp.git
cd web-explorer-mcp
./install.sh # or ./install.fish for Fish shell2. Configure Claude Desktop
Add to your Claude config (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):
{
"mcpServers": {
"web-explorer": {
"command": "uvx",
"args": ["web-explorer-mcp"]
}
}
}3. Restart Claude
That's it! Ask Claude to search the web.
Tools
web_search_tool(query, page, page_size)- Search the webwebpage_content_tool(url, max_chars, page)- Extract webpage content with pagination support
Configuration & Usage
See docs/CONFIGURATION.md for:
Other AI clients (Continue.dev, Cline)
Environment variables
Troubleshooting
Management commands
Update
uvx --force web-explorer-mcp # MCP server
docker compose pull && docker compose up -d # SearxNG + PlaywrightUninstall
docker compose down -v
cd .. && rm -rf web-explorer-mcpDevelopment
uv sync # Install dependencies
docker compose up -d # Start SearxNG + Playwright
uv run web-explorer-mcp # Run locallySee CONTRIBUTING.md for details.
License
MIT - see LICENSE
Available Tools
2 toolswebpage_content_toolA
Extract and clean webpage content for a provided URL.
This tool extracts full content from webpages using Playwright with JavaScript rendering. Content is automatically paginated for display if it exceeds max_chars.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to fetch and extract. | |
| page | No | Page number to return (default 1). Pagination is applied to main_content for readability, but full content is always extracted. | |
| max_chars | No | Maximum characters per page to include in the main text. If not provided, 5000 characters are used. Pagination is applied to main_content only. | |
| raw_content | No | If True, return raw HTML content without processing. Defaults to False. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that it uses Playwright with JavaScript rendering and that content is paginated when exceeding max_chars. However, it does not explicitly state that it is read-only or disclose limitations like site-specific failures, leaving some behavioral traits implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, with the first sentence immediately stating the purpose. There is no redundant content; every clause contributes meaning. This is an exemplary concise structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with the 100% schema coverage and output schema, the description provides a complete functional overview: extraction, cleaning, JS rendering, and pagination. It lacks explicit guidance on when to use this tool versus the sibling and does not clarify 'clean', but these are minor gaps given the existing structured context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers all four parameters with descriptions, so the high-coverage baseline applies. The description adds a note about pagination behavior, but this mostly restates the schema's 'pagination is applied to main_content only' without deepening parameter understanding. Therefore, a score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb+resource: 'Extract and clean webpage content for a provided URL.' This clearly distinguishes it from the sibling web_search_tool, which searches rather than fetches a known URL. The wording is unambiguous and informative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context that this tool is for extracting content from a specific URL, implying use when a URL is already known. It does not explicitly state when not to use it or mention alternatives, but the distinction from web_search_tool is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_search_toolA
Perform web search using SearxNG instance.
This tool searches the web using a local SearxNG instance and returns structured search results. It provides title, description, and URL for each result with support for pagination. The tool handles errors gracefully and returns them in the response rather than raising exceptions.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination, starting from 1. Each page contains `page_size` results. Defaults to 1 (first page). | |
| query | Yes | The search query string. Must be non-empty and will be trimmed of leading/trailing whitespace. Examples: "python programming", "machine learning tutorials", "fastapi documentation". | |
| page_size | No | Maximum number of results to return per page. If not provided, uses the default from application settings. Must be positive. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It discloses the result structure, pagination support, and graceful error handling ('returns them in the response rather than raising exceptions'). It does not mention rate limits or authentication, but for a local search tool these are minor omissions. It provides valuable behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence gives the core purpose, and the second adds essential behavioral details (result fields, pagination, error handling). It is front-loaded and efficiently worded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so the description needn't detail return values. It covers the core function, result fields, pagination, and error behavior. It lacks explicit guidance on when to choose this over webpage_content_tool, but that is covered under usage guidelines. Overall, it is a well-rounded description for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage for query, page, and page_size. The description mentions pagination support generally, which reaffirms the page/page_size purpose, but adds no specific parameter semantics beyond what the schema already states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Perform web search using SearxNG instance' – a specific verb and resource. It clearly states the tool searches the web and returns structured results (title, description, URL), distinguishing it from the sibling webpage_content_tool which fetches page content. No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool is for web searches and provides context about results and pagination, but it does not explicitly mention when to use it over webpage_content_tool or state any exclusion criteria. The use case is clear, but alternative guidance is not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.3.1- First observed
web_search_tool - First observed
webpage_content_tool
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: one performs web searches and the other extracts content from a given URL. There is no overlap or ambiguity between them.
Both tool names follow a consistent pattern: a descriptive noun phrase followed by '_tool' (web_search_tool, webpage_content_tool). This makes the naming predictable and clear.
With only 2 tools, the server feels minimal for a 'Web Explorer' purpose. While the two tools cover search and content retrieval, the count is borderline and could benefit from additional tools like link extraction or history management.
The tool set covers the core web exploration workflow: searching and fetching content. It lacks some advanced operations like extracting specific elements or managing browsing sessions, but the primary needs are met without major gaps.
Maintenance
Related MCP Connectors
Web search and page extraction across several independent search providers.
Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
- fastCRWOAuthio.github.us
Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents.
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables web search, image search, and news search through a self-hosted SearXNG instance. Provides privacy-focused meta-search capabilities aggregating results from multiple search engines.31MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to perform web searches and read URL content via a SearXNG instance.214 npmMIT
- AlicenseAqualityBmaintenanceEnables local LLMs to search the web and fetch clean content from URLs without API keys, using SearxNG and Mozilla Readability.236MIT
- FlicenseNot gradedqualityDmaintenanceEnables web search and content scraping from multiple engines via a local SearXNG instance, allowing AI assistants to retrieve and extract web content.1-