Skip to main content
Glama
PXSR

MCP Smart Searcher

by PXSR

MCP Smart Searcher

A smart MCP (Model Context Protocol) server for multi-engine web search with AI-powered results.

Features

  • Multi-engine search — Search across 8 engines simultaneously: DuckDuckGo, Baidu, Juejin, GitHub, GitHub Code, Tavily, Brave, Startpage

  • Web content extraction — Fetch and extract clean content from any public URL in multiple formats:

    • markdown (default) — Structured Markdown with headings, lists, code blocks, tables, and links preserved

    • article — Reader-mode extraction via Mozilla Readability, ideal for blogs/docs/news

    • text — Plain text, legacy behavior

    • outline — Page structure overview (headings, regions, interactive elements) without full content

    • Plus noise removal, hidden-element stripping, and prompt-guided filtering

  • Rate limiting — Built-in concurrency control via semaphore

  • Proxy support — Per-engine proxy configuration

  • Engine allowlist — Restrict which engines can be used

Related MCP server: evo-scry

Installation

# For users
pip install mcp-smart-searcher

# For development
pip install -e ".[dev]"

Usage

Run the server

# Direct command (after pip install)
mcp-smart-searcher

# Or via Python module
python -m mcp_smart_searcher

# Or via uvx (no install required)
uvx mcp-smart-searcher

MCP client configuration

Add to your MCP client config (e.g., Claude Desktop):

{
  "mcpServers": {
    "smart-searcher": {
      "command": "mcp-smart-searcher"
    }
  }
}

Or with uvx (no install required):

{
  "mcpServers": {
    "smart-searcher": {
      "command": "uvx",
      "args": ["mcp-smart-searcher"]
    }
  }
}

Development

# Run with MCP inspector
mcp dev src/mcp_smart_searcher/server.py

# Run tests
PYTHONPATH=src pytest

# Build
python -m build

Configuration

All settings are configured via environment variables:

Variable

Description

Default

DEFAULT_SEARCH_ENGINES

Comma-separated default engines when none specified

duckduckgo,baidu,startpage,tavily,brave

ALLOWED_SEARCH_ENGINES

Comma-separated allowlist; unset = all allowed

(all)

TAVILY_API_KEY

Tavily AI Search API key

(none)

GITHUB_TOKEN

GitHub API token (for github/github_code engines)

(none)

USE_PROXY

Enable proxy for engines that need it

true

PROXY_URL

Proxy address

http://127.0.0.1:10809

PROXY_ENGINES

Override: comma-separated engines that use proxy

(auto)

MAX_CONCURRENT_SEARCH

Max parallel search requests

5

LOG_LEVEL

Logging level (DEBUG/INFO/WARNING/ERROR)

INFO

Proxy behavior

By default, domestic engines (baidu, juejin) skip proxy, while all others use proxy. You can override this with PROXY_ENGINES:

# Only use proxy for DuckDuckGo and GitHub
PROXY_ENGINES=duckduckgo,github,github_code

# Disable proxy entirely
USE_PROXY=false

MCP client configuration with env vars

{
  "mcpServers": {
    "smart-searcher": {
      "command": "mcp-smart-searcher",
      "env": {
        "TAVILY_API_KEY": "tvly-xxx",
        "GITHUB_TOKEN": "ghp_xxx",
        "PROXY_URL": "http://127.0.0.1:10809",
        "LOG_LEVEL": "INFO"
      }
    }
  }
}

Or with uvx (no install required):

{
  "mcpServers": {
    "smart-searcher": {
      "command": "uvx",
      "args": ["mcp-smart-searcher"],
      "env": {
        "TAVILY_API_KEY": "tvly-xxx",
        "GITHUB_TOKEN": "ghp_xxx"
      }
    }
  }
}

Quick Start

1. Install

pip install mcp-smart-searcher

2. Configure (optional)

Create a .env file or set environment variables:

# .env
TAVILY_API_KEY=tvly-your-key-here
GITHUB_TOKEN=ghp_your-token-here
PROXY_URL=http://127.0.0.1:10809
LOG_LEVEL=INFO

3. Add to your MCP client

{
  "mcpServers": {
    "smart-searcher": {
      "command": "mcp-smart-searcher"
    }
  }
}

4. Done!

Your AI agent can now search the web and fetch web pages.

License

Apache-2.0

Available Tools

2 tools
fetch_web_contentA

Fetch and extract text content from any public URL.

Args: url: Public HTTP/HTTPS URL to fetch prompt: Optional hint for what to extract. When provided, the content is filtered to prioritize paragraphs most relevant to the prompt keywords. Useful for AI agents to focus on specific content (e.g., "extract code examples only", "summarize the main argument"). max_chars: Maximum characters to return (default 30000)

Returns: Extracted text content from the webpage, optionally filtered by prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
promptNo
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must cover behavioral traits. It mentions 'any public URL' and optional filtering, but omits details on error handling, redirects, authentication, or rate limits. The transparency is basic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with a clear header, Args list, and Returns section. It is not overly long, though the prompt examples could be slightly trimmed. Overall, it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the basic inputs and output, but lacks details on error scenarios, maximum allowable characters, or non-HTTP content handling. With no annotations, more context would be beneficial for complete understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, requiring the description to add meaning. It does so by describing the URL as 'public HTTP/HTTPS URL,' the prompt as an optional filter with examples, and max_chars with its default. This adds significant value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Fetch and extract text content from any public URL,' providing a specific verb and resource. It distinguishes from the sibling 'web_search' by focusing on fetching a known URL rather than searching for URLs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the prompt parameter and default max_chars, implying usage for targeted extraction. However, it does not explicitly contrast with the sibling 'web_search' or state when not to use this tool, leaving room for ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.2.1
    • First observedfetch_web_content
    • First observedweb_search

TDQS

A3.8/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: one fetches content from a specific URL, the other performs web searches. There is no overlap or ambiguity.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern (fetch_web_content, web_search) using snake_case, making them predictable and easy to understand.

Tool Count3/5

With only 2 tools, the server is minimal but covers the basic search-and-fetch workflow. It feels slightly thin for a server named 'Smart Searcher' as it lacks additional features like content summarization or multi-URL fetch.

Completeness3/5

The surface covers the core operations of searching and fetching web content, but there are notable gaps such as no tool to fetch search results directly or to handle multiple URLs. Agents may need to combine calls manually.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    MCP server for web crawling, searching, and AI-powered content extraction, supporting single-page, batch, and full-site crawling along with text, news, image, book, and video search.
    8
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Multi-engine aggregated search MCP server that combines results from 7 search engines with deduplication, relevance ranking, and web page content extraction.
    1
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/PXSR/mcp-smart-searcher'

If you have feedback or need assistance with the MCP directory API, please join our Discord server