Skip to main content
Glama
l0kifs

Web Explorer MCP

by l0kifs

Web Explorer MCP

A Model Context Protocol (MCP) server that provides web search and webpage content extraction using a local SearxNG instance.

Why Web Explorer MCP?

Unlike commercial solutions (GitHub Copilot, Cursor IDE), Web Explorer MCP prioritizes privacy and autonomy:

Feature

Web Explorer MCP

GitHub Copilot

Cursor IDE

Privacy

✅ Local SearxNG, zero tracking

❌ Bing API, Microsoft servers

❌ Cloud search, third-party APIs

Cost

✅ Free, no limits

💰 $10-20/month subscription

💰 $20/month Pro plan

API Keys

✅ None required

⚠️ GitHub account required

⚠️ Account & subscription

Data Control

✅ All data stays local

❌ Queries sent to Microsoft

❌ Queries sent to external services

Setup

✅ 2 commands

⚠️ Account setup, policy config

⚠️ Account, payment setup

Open Source

✅ Fully auditable

⚠️ Partial (client only)

❌ Proprietary

Perfect for: Developers who value privacy, work with sensitive data, or prefer not to depend on external services and subscriptions.

Related MCP server: mcp-searxng

⚠️ Responsible Use

This tool is designed for human-assisted AI interactions, not for automated high-volume scraping:

  • 🚫 Not for DDoS - Do not use for overwhelming websites or search engines

  • 🚫 Not for High-Speed Automation - Avoid usage speeds significantly higher than a real user

  • 🚫 Not for Fully Automated AI Agents - Not recommended for high-performance autonomous agents

  • Respect Infrastructure - Honor website owners' business scenarios and infrastructure capabilities

  • Follow robots.txt - Respect crawling policies and rate limits

Use responsibly: This tool is meant for legitimate research and development, not for abuse.

Features

  • 🔍 Web Search - Search using local SearxNG (private, no API keys)

  • 📄 Content Extraction - Extract clean text from webpages with Playwright rendering

  • 🐳 Zero Pollution - Runs in Docker, leaves no traces

  • 🚀 Simple Setup - Install in 2 commands

Quick Start

1. Install Services (SearxNG + Playwright)

git clone https://github.com/l0kifs/web-explorer-mcp.git
cd web-explorer-mcp
./install.sh  # or ./install.fish for Fish shell

2. Configure Claude Desktop

Add to your Claude config (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):

{
  "mcpServers": {
    "web-explorer": {
      "command": "uvx",
      "args": ["web-explorer-mcp"]
    }
  }
}

3. Restart Claude

That's it! Ask Claude to search the web.

Tools

  • web_search_tool(query, page, page_size) - Search the web

  • webpage_content_tool(url, max_chars, page) - Extract webpage content with pagination support

Configuration & Usage

See docs/CONFIGURATION.md for:

  • Other AI clients (Continue.dev, Cline)

  • Environment variables

  • Troubleshooting

  • Management commands

Update

uvx --force web-explorer-mcp  # MCP server
docker compose pull && docker compose up -d  # SearxNG + Playwright

Uninstall

docker compose down -v
cd .. && rm -rf web-explorer-mcp

Development

uv sync              # Install dependencies
docker compose up -d # Start SearxNG + Playwright
uv run web-explorer-mcp  # Run locally

See CONTRIBUTING.md for details.

License

MIT - see LICENSE

Available Tools

2 tools
webpage_content_toolA

Extract and clean webpage content for a provided URL.

This tool extracts full content from webpages using Playwright with JavaScript rendering. Content is automatically paginated for display if it exceeds max_chars.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to fetch and extract.
pageNoPage number to return (default 1). Pagination is applied to main_content for readability, but full content is always extracted.
max_charsNoMaximum characters per page to include in the main text. If not provided, 5000 characters are used. Pagination is applied to main_content only.
raw_contentNoIf True, return raw HTML content without processing. Defaults to False.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that it uses Playwright with JavaScript rendering and that content is paginated when exceeding max_chars. However, it does not explicitly state that it is read-only or disclose limitations like site-specific failures, leaving some behavioral traits implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, with the first sentence immediately stating the purpose. There is no redundant content; every clause contributes meaning. This is an exemplary concise structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Combined with the 100% schema coverage and output schema, the description provides a complete functional overview: extraction, cleaning, JS rendering, and pagination. It lacks explicit guidance on when to use this tool versus the sibling and does not clarify 'clean', but these are minor gaps given the existing structured context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all four parameters with descriptions, so the high-coverage baseline applies. The description adds a note about pagination behavior, but this mostly restates the schema's 'pagination is applied to main_content only' without deepening parameter understanding. Therefore, a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific verb+resource: 'Extract and clean webpage content for a provided URL.' This clearly distinguishes it from the sibling web_search_tool, which searches rather than fetches a known URL. The wording is unambiguous and informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context that this tool is for extracting content from a specific URL, implying use when a URL is already known. It does not explicitly state when not to use it or mention alternatives, but the distinction from web_search_tool is clear enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

web_search_toolA

Perform web search using SearxNG instance.

This tool searches the web using a local SearxNG instance and returns structured search results. It provides title, description, and URL for each result with support for pagination. The tool handles errors gracefully and returns them in the response rather than raising exceptions.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number for pagination, starting from 1. Each page contains `page_size` results. Defaults to 1 (first page).
queryYesThe search query string. Must be non-empty and will be trimmed of leading/trailing whitespace. Examples: "python programming", "machine learning tutorials", "fastapi documentation".
page_sizeNoMaximum number of results to return per page. If not provided, uses the default from application settings. Must be positive.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It discloses the result structure, pagination support, and graceful error handling ('returns them in the response rather than raising exceptions'). It does not mention rate limits or authentication, but for a local search tool these are minor omissions. It provides valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence gives the core purpose, and the second adds essential behavioral details (result fields, pagination, error handling). It is front-loaded and efficiently worded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so the description needn't detail return values. It covers the core function, result fields, pagination, and error behavior. It lacks explicit guidance on when to choose this over webpage_content_tool, but that is covered under usage guidelines. Overall, it is a well-rounded description for a search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for query, page, and page_size. The description mentions pagination support generally, which reaffirms the page/page_size purpose, but adds no specific parameter semantics beyond what the schema already states. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Perform web search using SearxNG instance' – a specific verb and resource. It clearly states the tool searches the web and returns structured results (title, description, URL), distinguishing it from the sibling webpage_content_tool which fetches page content. No ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the tool is for web searches and provides context about results and pagination, but it does not explicitly mention when to use it over webpage_content_tool or state any exclusion criteria. The use case is clear, but alternative guidance is not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.3.1
    • First observedweb_search_tool
    • First observedwebpage_content_tool

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: one performs web searches and the other extracts content from a given URL. There is no overlap or ambiguity between them.

Naming Consistency5/5

Both tool names follow a consistent pattern: a descriptive noun phrase followed by '_tool' (web_search_tool, webpage_content_tool). This makes the naming predictable and clear.

Tool Count3/5

With only 2 tools, the server feels minimal for a 'Web Explorer' purpose. While the two tools cover search and content retrieval, the count is borderline and could benefit from additional tools like link extraction or history management.

Completeness4/5

The tool set covers the core web exploration workflow: searching and fetching content. It lacks some advanced operations like extracting specific elements or managing browsing sessions, but the primary needs are met without major gaps.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables web search, image search, and news search through a self-hosted SearXNG instance. Provides privacy-focused meta-search capabilities aggregating results from multiple search engines.
    3
    1
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to perform web searches and read URL content via a SearXNG instance.
    2
    14 npm
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables local LLMs to search the web and fetch clean content from URLs without API keys, using SearxNG and Mozilla Readability.
    2
    36
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables web search and content scraping from multiple engines via a local SearXNG instance, allowing AI assistants to retrieve and extract web content.
    1
    -