Heventure Search MCP
中文 | English
MCP Web Search Server
一个免费、无需 API 密钥的 Web 搜索 MCP (Model Context Protocol) 服务器,支持 DuckDuckGo、Bing、Google,并可选配 SerpAPI/Tavily 以提升搜索质量。
功能特性
🔍 多引擎搜索:DuckDuckGo + Bing + Google(免费,无需 API 密钥)
🔑 可选 API 密钥:支持 SerpAPI 和 Tavily 以获得更好的搜索质量
📄 网页内容抓取:获取任意网页的文本内容
🚀 异步处理:基于 asyncio 的高性能异步处理
Related MCP server: Web Search MCP Server
安装
PyPI (推荐)
pip install heventure-search-mcp
heventure-search-mcpuvx
uvx heventure-search-mcp从源码安装
pip install git+https://github.com/HughesCuit/heventure-search-mcp.git
python -m server使用方法
MCP 客户端配置
{
"mcpServers": {
"web-search": {
"command": "python",
"args": ["/path/to/server.py"]
}
}
}Trae AI
{
"mcpServers": {
"heventure-search-mcp": {
"command": "uvx",
"args": ["heventure-search-mcp"]
}
}
}可用工具
web_search
使用多个引擎搜索网页内容。
参数:
query(字符串,必填):搜索查询词max_results(整数,可选):最大结果数(默认:10,范围:1-20)search_engine(字符串,可选):引擎选择(默认:"both")"duckduckgo":仅 DuckDuckGo"bing":仅 Bing"google":仅 Google"both":DuckDuckGo + Google + Bing
可选 API 密钥(用于增强搜索)
您可以选择设置环境变量来启用付费搜索引擎:
# SerpAPI (Google search results via API, 100 searches/month free)
export SERPAPI_KEY="your_serpapi_key"
# Tavily (AI-optimized search, 1000 searches/month free)
export TAVILY_API_KEY="your_tavily_api_key"配置 API 密钥后,系统将自动结合免费引擎使用它们,以提高搜索质量。
示例:
{
"query": "Python tutorial",
"max_results": 5,
"search_engine": "both"
}get_webpage_content
获取网页的文本内容。
参数:
url(字符串,必填):目标网页 URL
示例:
{
"url": "https://example.com"
}错误处理
网络故障时自动重试
解析错误时优雅降级
用户友好的错误提示
许可证
MIT License
贡献
欢迎提交 Issue 和 Pull Request!
Available Tools
2 toolsget_webpage_contentA
Extract readable text content from a specific webpage URL. Use this tool after web_search when you need the full text of a page found in search results. The tool strips scripts, styles, and navigation elements, returning only the meaningful text content. Output is truncated to 2000 characters to stay within context limits. Returns an empty string if the page cannot be fetched or parsed.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The full URL of the webpage to extract content from (e.g., 'https://example.com/article'). Must be a valid HTTP or HTTPS URL. Supports most standard web pages; JavaScript-rendered content may not be fully captured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, but description discloses stripping of non-content elements, truncation to 2000 characters, and empty string on failure. Provides adequate behavioral insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five concise sentences front-load the purpose. No redundant information; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given low complexity and no output schema, description fully covers return behavior, truncation, and failure modes. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single URL parameter; description adds context about JavaScript-rendered content limitations, which enhances understanding beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool extracts readable text from a URL, specifying it strips scripts/styles/navigation. Distinguishes from sibling 'web_search' by indicating usage after search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use: after web_search when full text of a search result page is needed. Does not explicitly state when not to use, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchA
Search the web using multiple search engines. Use this tool when you need to find current information, articles, or references from the internet. Supports DuckDuckGo, Google, and Bing as free engines, with optional SerpAPI (requires SERPAPI_KEY env var) and Tavily (requires TAVILY_API_KEY env var) for higher-quality, more stable results. By default ('both'), queries DuckDuckGo and Bing concurrently and merges results, providing good coverage without any API keys. Each result includes title, URL, and description snippet. Rate limiting and caching are applied automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query string. Use specific keywords for best results. Supports natural language queries (e.g., 'best practices for Python async') as well as keyword-style queries (e.g., 'Python async await tutorial'). | |
| max_results | No | Maximum number of results to return per search engine. Default is 10. Range: 1-20. Higher values increase latency. Actual count may be lower if the engine returns fewer matches. | |
| search_engine | No | Which search engine(s) to query. Options: 'duckduckgo' (free, no API key needed), 'bing' (free), 'google' (free), 'serpapi' (requires SERPAPI_KEY), 'tavily' (requires TAVILY_API_KEY), 'both' (default — uses DuckDuckGo + Bing concurrently for broader coverage without API keys). Use 'serpapi' or 'tavily' when you need higher-quality results and have the API key configured. | both |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Discloses automatic rate limiting and caching, concurrent querying for 'both', and that actual results may be fewer. Clearly indicates read-only behavior (no destructive actions).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise paragraph (5 sentences) with front-loaded purpose. Every sentence contributes essential information: purpose, use cases, engine options, default behavior, result format, and automatic features.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lack of output schema, description details result format (title, URL, snippet). Covers all three parameters, engine behaviors, caching/rate limiting, and result count caveats. No missing context for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already has 100% coverage, but description adds value by explaining default behavior ('both' queries DuckDuckGo and Bing concurrently), environment variable requirements for serpapi/tavily, and that higher max_results increases latency.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'Search the web using multiple search engines' and specifies the resource (web) and action (search). Distinguishes from sibling tool 'get_webpage_content' by focusing on search rather than content retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this tool when you need to find current information, articles, or references from the internet.' Provides context for engine selection (free vs. API-based), though it doesn't explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
get_webpage_content1 field changed- changed
Input schema / properties / url / descriptionPrevious value: -"要获取内容的网页URL"New value: +"The full URL of the webpage to extract content from (e.g., 'https://example.com/article'). Must be a valid HTTP or HTTPS URL. Supports most standard web pages; JavaScript-rendered content may not be fully captured."
- Changed
web_search3 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"最大结果数量"New value: +"Maximum number of results to return per search engine. Default is 10. Range: 1-20. Higher values increase latency. Actual count may be lower if the engine returns fewer matches." - changed
Input schema / properties / query / descriptionPrevious value: -"搜索查询字符串"New value: +"The search query string. Use specific keywords for best results. Supports natural language queries (e.g., 'best practices for Python async') as well as keyword-style queries (e.g., 'Python async await tutorial')." - changed
Input schema / properties / search_engine / descriptionPrevious value: -"搜索引擎选择:duckduckgo / bing / google / serpapi / tavily / both"New value: +"Which search engine(s) to query. Options: 'duckduckgo' (free, no API key needed), 'bing' (free), 'google' (free), 'serpapi' (requires SERPAPI_KEY), 'tavily' (requires TAVILY_API_KEY), 'both' (default — uses DuckDuckGo + Bing concurrently for broader coverage without API keys). Use 'serpapi' or 'tavily' when you need higher-quality results and have the API key configured."
2 tool updates
v0.1.0- First observed
get_webpage_content - First observed
web_search
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one searches the web, the other retrieves content from a specific URL. There is no overlap in functionality, and their descriptions make the differentiation explicit.
Both tool names follow a consistent verb_noun pattern: 'web_search' and 'get_webpage_content'. The naming is clear, predictable, and uses underscores consistently.
With only two tools, the server is minimal but well-scoped for a search-and-retrieve workflow. While more tools could be added (e.g., for advanced filtering or caching), the current count fits the server's stated purpose without being too thin.
The tool set covers the basic search workflow: find pages via web_search, then extract content via get_webpage_content. There are no obvious gaps for the intended use case of retrieving text from web pages.
Maintenance
Related MCP Connectors
Web search and page extraction across several independent search providers.
1Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables browser automation using Puppeteer through the MCP interface. Allows launching browsers, creating pages, and executing arbitrary JavaScript for web scraping, testing, and debugging tasks.5149 npm2MIT
- FlicenseNot gradedqualityCmaintenanceEnables web search across multiple search engines (DuckDuckGo, Bing, Startpage) with parallel execution and result deduplication. Also provides web page content extraction capabilities.2-
- AlicenseNot gradedqualityDmaintenanceEnables web searching through Google, DuckDuckGo, and Bing using a headless Chrome browser, returning structured results with titles, URLs, and snippets. Also supports fetching and extracting text content from any webpage.13MIT
- AlicenseAqualityDmaintenanceEnables AI agents to securely read and extract information from PDF files including text content, metadata, and page counts from both local files and URLs within the project context.11,612 npmMIT