Heventure Search MCP
The Heventure Search MCP server provides web search and webpage content extraction for AI tools like Claude Desktop and Cursor.
Web search using the
web_searchtool with multiple engines: DuckDuckGo, Bing, Google, SerpAPI, or Tavily. Usebothto query multiple engines simultaneously.Control results: specify
max_results(1–20, default 10) and choose a specificsearch_engine.Extract webpage content via the
get_webpage_contenttool — provide any URL to retrieve its readable text.No API keys required for free engines (DuckDuckGo, Bing, Google); optionally configure SerpAPI or Tavily keys for enhanced quality.
Caching with LRU cache (300s TTL, 100 entries) for performance.
Provides web search capabilities using DuckDuckGo's search engine, supporting both API and HTML parsing methods to retrieve search results without requiring an API key.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Heventure Search MCPfind recent developments in quantum computing"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
中文 | English
🔍 MCP Web Search Server
Free forever. No API key required. A web search MCP server that works out of the box with Claude Desktop, Cursor, and any MCP-compatible AI tool.
pip install heventure-search-mcp✨ Why?
Most MCP search servers require you to sign up for API keys (Bing, Google, SerpAPI...). This one works immediately — zero configuration, zero cost, zero sign-ups.
Feature | This Server | Others |
No API key needed | ✅ | ❌ |
DuckDuckGo (free) | ✅ | varies |
Bing (free) | ✅ | ❌ |
Google (free) | ✅ | ❌ |
Optional SerpAPI/Tavily | ✅ | ✅ |
Async + caching | ✅ | varies |
Install in 10 seconds | ✅ | varies |
Related MCP server: Web Search MCP Server
🚀 Quick Start
Option 1: Claude Desktop / Cursor
Add to your MCP config:
{
"mcpServers": {
"web-search": {
"command": "uvx",
"args": ["heventure-search-mcp"]
}
}
}Option 2: Command Line
pip install heventure-search-mcp
heventure-search-mcpOption 3: Docker
docker run -p 8080:8080 heventure-search-mcp🔧 Available Tools
web_search
Search the web with multiple engines simultaneously.
Parameter | Type | Default | Description |
| string | required | Search query |
| int | 10 | Number of results (1-20) |
| string |
|
|
get_page_content
Extract readable text from any webpage.
Parameter | Type | Default | Description |
| string | required | Page URL to fetch |
🔑 Optional: Enhanced Search
The free engines work great for most use cases. For higher quality results, you can optionally add paid API keys:
# SerpAPI — 100 free searches/month
export SERPAPI_KEY="your_key"
# Tavily — 1,000 free searches/month
export TAVILY_API_KEY="your_key"🏗️ Architecture
Engines: DuckDuckGo, Bing, Google, SerpAPI, Tavily
Caching: LRU cache with 300s TTL (100 entries max)
Protocol: MCP (Model Context Protocol)
Runtime: Python 3.10+ with asyncio
🤝 Contributing
Issues and Pull Requests are welcome! See CONTRIBUTING.md for guidelines.
📄 License
MIT License — use it however you want.
Available Tools
2 toolsget_webpage_contentA
Extract readable text content from a specific webpage URL. Use this tool after web_search when you need the full text of a page found in search results. The tool strips scripts, styles, and navigation elements, returning only the meaningful text content. Output is truncated to 2000 characters to stay within context limits. Returns an empty string if the page cannot be fetched or parsed.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The full URL of the webpage to extract content from (e.g., 'https://example.com/article'). Must be a valid HTTP or HTTPS URL. Supports most standard web pages; JavaScript-rendered content may not be fully captured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, but description discloses stripping of non-content elements, truncation to 2000 characters, and empty string on failure. Provides adequate behavioral insight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five concise sentences front-load the purpose. No redundant information; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given low complexity and no output schema, description fully covers return behavior, truncation, and failure modes. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single URL parameter; description adds context about JavaScript-rendered content limitations, which enhances understanding beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool extracts readable text from a URL, specifying it strips scripts/styles/navigation. Distinguishes from sibling 'web_search' by indicating usage after search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use: after web_search when full text of a search result page is needed. Does not explicitly state when not to use, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchA
Search the web using multiple search engines. Use this tool when you need to find current information, articles, or references from the internet. Supports DuckDuckGo, Google, and Bing as free engines, with optional SerpAPI (requires SERPAPI_KEY env var) and Tavily (requires TAVILY_API_KEY env var) for higher-quality, more stable results. By default ('both'), queries DuckDuckGo and Bing concurrently and merges results, providing good coverage without any API keys. Each result includes title, URL, and description snippet. Rate limiting and caching are applied automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query string. Use specific keywords for best results. Supports natural language queries (e.g., 'best practices for Python async') as well as keyword-style queries (e.g., 'Python async await tutorial'). | |
| max_results | No | Maximum number of results to return per search engine. Default is 10. Range: 1-20. Higher values increase latency. Actual count may be lower if the engine returns fewer matches. | |
| search_engine | No | Which search engine(s) to query. Options: 'duckduckgo' (free, no API key needed), 'bing' (free), 'google' (free), 'serpapi' (requires SERPAPI_KEY), 'tavily' (requires TAVILY_API_KEY), 'both' (default — uses DuckDuckGo + Bing concurrently for broader coverage without API keys). Use 'serpapi' or 'tavily' when you need higher-quality results and have the API key configured. | both |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Discloses automatic rate limiting and caching, concurrent querying for 'both', and that actual results may be fewer. Clearly indicates read-only behavior (no destructive actions).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise paragraph (5 sentences) with front-loaded purpose. Every sentence contributes essential information: purpose, use cases, engine options, default behavior, result format, and automatic features.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lack of output schema, description details result format (title, URL, snippet). Covers all three parameters, engine behaviors, caching/rate limiting, and result count caveats. No missing context for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already has 100% coverage, but description adds value by explaining default behavior ('both' queries DuckDuckGo and Bing concurrently), environment variable requirements for serpapi/tavily, and that higher max_results increases latency.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'Search the web using multiple search engines' and specifies the resource (web) and action (search). Distinguishes from sibling tool 'get_webpage_content' by focusing on search rather than content retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this tool when you need to find current information, articles, or references from the internet.' Provides context for engine selection (free vs. API-based), though it doesn't explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
get_webpage_content1 field changed- changed
Input schema / properties / url / descriptionPrevious value: -"要获取内容的网页URL"New value: +"The full URL of the webpage to extract content from (e.g., 'https://example.com/article'). Must be a valid HTTP or HTTPS URL. Supports most standard web pages; JavaScript-rendered content may not be fully captured."
- Changed
web_search3 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"最大结果数量"New value: +"Maximum number of results to return per search engine. Default is 10. Range: 1-20. Higher values increase latency. Actual count may be lower if the engine returns fewer matches." - changed
Input schema / properties / query / descriptionPrevious value: -"搜索查询字符串"New value: +"The search query string. Use specific keywords for best results. Supports natural language queries (e.g., 'best practices for Python async') as well as keyword-style queries (e.g., 'Python async await tutorial')." - changed
Input schema / properties / search_engine / descriptionPrevious value: -"搜索引擎选择:duckduckgo / bing / google / serpapi / tavily / both"New value: +"Which search engine(s) to query. Options: 'duckduckgo' (free, no API key needed), 'bing' (free), 'google' (free), 'serpapi' (requires SERPAPI_KEY), 'tavily' (requires TAVILY_API_KEY), 'both' (default — uses DuckDuckGo + Bing concurrently for broader coverage without API keys). Use 'serpapi' or 'tavily' when you need higher-quality results and have the API key configured."
2 tool updates
v0.1.0- First observed
get_webpage_content - First observed
web_search
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one searches the web, the other retrieves content from a specific URL. There is no overlap in functionality, and their descriptions make the differentiation explicit.
Both tool names follow a consistent verb_noun pattern: 'web_search' and 'get_webpage_content'. The naming is clear, predictable, and uses underscores consistently.
With only two tools, the server is minimal but well-scoped for a search-and-retrieve workflow. While more tools could be added (e.g., for advanced filtering or caching), the current count fits the server's stated purpose without being too thin.
The tool set covers the basic search workflow: find pages via web_search, then extract content via get_webpage_content. There are no obvious gaps for the intended use case of retrieving text from web pages.
Maintenance
Related MCP Connectors
Web search and page extraction across several independent search providers.
1Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables browser automation using Puppeteer through the MCP interface. Allows launching browsers, creating pages, and executing arbitrary JavaScript for web scraping, testing, and debugging tasks.5149 npm2MIT
- FlicenseNot gradedqualityCmaintenanceEnables web search across multiple search engines (DuckDuckGo, Bing, Startpage) with parallel execution and result deduplication. Also provides web page content extraction capabilities.2-
- AlicenseNot gradedqualityDmaintenanceEnables web searching through Google, DuckDuckGo, and Bing using a headless Chrome browser, returning structured results with titles, URLs, and snippets. Also supports fetching and extracting text content from any webpage.13MIT
- AlicenseAqualityDmaintenanceEnables AI agents to securely read and extract information from PDF files including text content, metadata, and page counts from both local files and URLs within the project context.11,612 npmMIT