Heventure Search MCP
中文 | English
MCP Web Search Server
Бесплатный MCP-сервер (Model Context Protocol) для поиска в интернете без API-ключей, поддерживающий DuckDuckGo, Bing, Google и опционально SerpAPI/Tavily для повышения качества поиска.
Возможности
🔍 Поиск по нескольким движкам: DuckDuckGo + Bing + Google (бесплатно, API-ключи не требуются)
🔑 Опциональные API-ключи: SerpAPI и Tavily для лучшего качества поиска
📄 Извлечение контента: Получение текстового содержимого с любой веб-страницы
🚀 Асинхронная обработка: Высокопроизводительная обработка на основе asyncio
Related MCP server: Web Search MCP Server
Установка
PyPI (Рекомендуется)
pip install heventure-search-mcp
heventure-search-mcpuvx
uvx heventure-search-mcpИз исходного кода
pip install git+https://github.com/HughesCuit/heventure-search-mcp.git
python -m serverИспользование
Конфигурация MCP-клиента
{
"mcpServers": {
"web-search": {
"command": "python",
"args": ["/path/to/server.py"]
}
}
}Trae AI
{
"mcpServers": {
"heventure-search-mcp": {
"command": "uvx",
"args": ["heventure-search-mcp"]
}
}
}Доступные инструменты
web_search
Поиск контента в интернете с использованием нескольких движков.
Параметры:
query(строка, обязательно): Поисковый запросmax_results(целое число, опционально): Максимальное количество результатов (по умолчанию: 10, диапазон: 1-20)search_engine(строка, опционально): Выбор движка (по умолчанию: "both")"duckduckgo": Только DuckDuckGo"bing": Только Bing"google": Только Google"both": DuckDuckGo + Google + Bing
Опциональные API-ключи (для улучшенного поиска)
Вы можете опционально установить переменные окружения для включения платных поисковых систем:
# SerpAPI (Google search results via API, 100 searches/month free)
export SERPAPI_KEY="your_serpapi_key"
# Tavily (AI-optimized search, 1000 searches/month free)
export TAVILY_API_KEY="your_tavily_api_key"Когда API-ключи настроены, они будут автоматически использоваться вместе с бесплатными движками для улучшения качества поиска.
Пример:
{
"query": "Python tutorial",
"max_results": 5,
"search_engine": "both"
}get_webpage_content
Получение текстового содержимого с веб-страницы.
Параметры:
url(строка, обязательно): URL целевой веб-страницы
Пример:
{
"url": "https://example.com"
}Обработка ошибок
Автоматический повтор при сбоях сети
Корректная обработка ошибок парсинга
Понятные пользователю сообщения об ошибках
Лицензия
MIT License
Участие в разработке
Приветствуются сообщения об ошибках (Issues) и Pull Requests!
Available Tools
2 toolsget_webpage_contentA
Extract readable text content from a specific webpage URL. Use this tool after web_search when you need the full text of a page found in search results. The tool strips scripts, styles, and navigation elements, returning only the meaningful text content. Output is truncated to 2000 characters to stay within context limits. Returns an empty string if the page cannot be fetched or parsed.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The full URL of the webpage to extract content from (e.g., 'https://example.com/article'). Must be a valid HTTP or HTTPS URL. Supports most standard web pages; JavaScript-rendered content may not be fully captured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, but description discloses stripping of non-content elements, truncation to 2000 characters, and empty string on failure. Provides adequate behavioral insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five concise sentences front-load the purpose. No redundant information; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given low complexity and no output schema, description fully covers return behavior, truncation, and failure modes. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single URL parameter; description adds context about JavaScript-rendered content limitations, which enhances understanding beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool extracts readable text from a URL, specifying it strips scripts/styles/navigation. Distinguishes from sibling 'web_search' by indicating usage after search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use: after web_search when full text of a search result page is needed. Does not explicitly state when not to use, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchA
Search the web using multiple search engines. Use this tool when you need to find current information, articles, or references from the internet. Supports DuckDuckGo, Google, and Bing as free engines, with optional SerpAPI (requires SERPAPI_KEY env var) and Tavily (requires TAVILY_API_KEY env var) for higher-quality, more stable results. By default ('both'), queries DuckDuckGo and Bing concurrently and merges results, providing good coverage without any API keys. Each result includes title, URL, and description snippet. Rate limiting and caching are applied automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query string. Use specific keywords for best results. Supports natural language queries (e.g., 'best practices for Python async') as well as keyword-style queries (e.g., 'Python async await tutorial'). | |
| max_results | No | Maximum number of results to return per search engine. Default is 10. Range: 1-20. Higher values increase latency. Actual count may be lower if the engine returns fewer matches. | |
| search_engine | No | Which search engine(s) to query. Options: 'duckduckgo' (free, no API key needed), 'bing' (free), 'google' (free), 'serpapi' (requires SERPAPI_KEY), 'tavily' (requires TAVILY_API_KEY), 'both' (default — uses DuckDuckGo + Bing concurrently for broader coverage without API keys). Use 'serpapi' or 'tavily' when you need higher-quality results and have the API key configured. | both |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Discloses automatic rate limiting and caching, concurrent querying for 'both', and that actual results may be fewer. Clearly indicates read-only behavior (no destructive actions).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise paragraph (5 sentences) with front-loaded purpose. Every sentence contributes essential information: purpose, use cases, engine options, default behavior, result format, and automatic features.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lack of output schema, description details result format (title, URL, snippet). Covers all three parameters, engine behaviors, caching/rate limiting, and result count caveats. No missing context for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already has 100% coverage, but description adds value by explaining default behavior ('both' queries DuckDuckGo and Bing concurrently), environment variable requirements for serpapi/tavily, and that higher max_results increases latency.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'Search the web using multiple search engines' and specifies the resource (web) and action (search). Distinguishes from sibling tool 'get_webpage_content' by focusing on search rather than content retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this tool when you need to find current information, articles, or references from the internet.' Provides context for engine selection (free vs. API-based), though it doesn't explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
get_webpage_content1 field changed- changed
Input schema / properties / url / descriptionPrevious value: -"要获取内容的网页URL"New value: +"The full URL of the webpage to extract content from (e.g., 'https://example.com/article'). Must be a valid HTTP or HTTPS URL. Supports most standard web pages; JavaScript-rendered content may not be fully captured."
- Changed
web_search3 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"最大结果数量"New value: +"Maximum number of results to return per search engine. Default is 10. Range: 1-20. Higher values increase latency. Actual count may be lower if the engine returns fewer matches." - changed
Input schema / properties / query / descriptionPrevious value: -"搜索查询字符串"New value: +"The search query string. Use specific keywords for best results. Supports natural language queries (e.g., 'best practices for Python async') as well as keyword-style queries (e.g., 'Python async await tutorial')." - changed
Input schema / properties / search_engine / descriptionPrevious value: -"搜索引擎选择:duckduckgo / bing / google / serpapi / tavily / both"New value: +"Which search engine(s) to query. Options: 'duckduckgo' (free, no API key needed), 'bing' (free), 'google' (free), 'serpapi' (requires SERPAPI_KEY), 'tavily' (requires TAVILY_API_KEY), 'both' (default — uses DuckDuckGo + Bing concurrently for broader coverage without API keys). Use 'serpapi' or 'tavily' when you need higher-quality results and have the API key configured."
2 tool updates
v0.1.0- First observed
get_webpage_content - First observed
web_search
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one searches the web, the other retrieves content from a specific URL. There is no overlap in functionality, and their descriptions make the differentiation explicit.
Both tool names follow a consistent verb_noun pattern: 'web_search' and 'get_webpage_content'. The naming is clear, predictable, and uses underscores consistently.
With only two tools, the server is minimal but well-scoped for a search-and-retrieve workflow. While more tools could be added (e.g., for advanced filtering or caching), the current count fits the server's stated purpose without being too thin.
The tool set covers the basic search workflow: find pages via web_search, then extract content via get_webpage_content. There are no obvious gaps for the intended use case of retrieving text from web pages.
Maintenance
Related MCP Connectors
Web search and page extraction across several independent search providers.
1Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables browser automation using Puppeteer through the MCP interface. Allows launching browsers, creating pages, and executing arbitrary JavaScript for web scraping, testing, and debugging tasks.5149 npm2MIT
- FlicenseNot gradedqualityCmaintenanceEnables web search across multiple search engines (DuckDuckGo, Bing, Startpage) with parallel execution and result deduplication. Also provides web page content extraction capabilities.2-
- AlicenseNot gradedqualityDmaintenanceEnables web searching through Google, DuckDuckGo, and Bing using a headless Chrome browser, returning structured results with titles, URLs, and snippets. Also supports fetching and extracting text content from any webpage.13MIT
- AlicenseAqualityDmaintenanceEnables AI agents to securely read and extract information from PDF files including text content, metadata, and page counts from both local files and URLs within the project context.11,612 npmMIT