MyWebSearch
This server provides keyless multi-engine web search and clean content/documentation retrieval tools for AI agents.
Search the web across 9 engines (Bing, Baidu, CSDN, Juejin, Sogou, DuckDuckGo, Brave, Startpage, Exa) with auto-routing, parallel queries, deduplication, and optional min-result cascading.
Fetch and extract web page content as clean text or Markdown, with readability mode, link inclusion, character limits, and pagination.
Fetch full CSDN articles from blog.csdn.net URLs.
Fetch full Juejin (掘金) articles in text or Markdown format.
Fetch GitHub or Gitee repository READMEs directly from repo URLs.
Resolve library/package names to Context7 IDs for official documentation lookups.
Query version-specific official library docs and code snippets via Context7, without needing a proxy.
Search the web using Baidu as the search engine, returning structured results with titles, URLs, and descriptions.
Search the web using Brave Search, returning structured results with titles, URLs, and descriptions.
Search the web using CSDN and fetch individual article content from CSDN pages.
Search the web using DuckDuckGo, returning structured results with titles, URLs, and descriptions.
Fetch README files from GitHub repositories as content.
Search the web using Juejin, returning structured results with titles, URLs, and descriptions.
Search the web using Sogou, returning structured results with titles, URLs, and descriptions.
Search the web using Startpage, returning structured results with titles, URLs, and descriptions.
🔍 MyWebSearch
Keyless Multi-Engine AI Web Search & High-Purity Content Extraction Engine
MCP Server · CLI · Local HTTP Daemon · Skill-Guided Agent Workflows
✨ Why MyWebSearch?
my-websearch is a full-stack web retrieval and documentation engine built for AI coding agents and MCP clients (Claude Desktop, Cursor, Windsurf, Cherry Studio, Cline, ZCode, etc.). It delivers 9-engine federated search, version-pinned official library documentation via Context7, and clean Markdown web page extraction—with zero paid API keys required.
🌐 9-Engine Smart Orchestration: Combines direct domestic engines (Bing, Baidu, CSDN, Juejin, Sogou) and global engines (DuckDuckGo, Brave, Startpage, Exa) with language-aware
autorouting, parallel multi-query execution (queries: string[]), cross-engine URL deduplication/ranking, circuit breakers, andminResultsautomatic cascade fallback.🛡️ Native Anti-Bot & Challenge Solvers (Zero-Browser Fast Path):
Chrome TLS/HTTP2 Fingerprint Impersonation: Powered by
wreq-jswith persistent session cookie jars (e.g., automatic Alibaba Cloudhttps_waf_cookiepersistence on CSDN and Chrome 131/133 TLS handshakes on Bing/Brave/Startpage).Millisecond Cryptographic & JS Challenge Solvers: Built-in pure-JS/Rust solvers for Startpage's Anubis SHA-256 Proof-of-Work (PoW) and DuckDuckGo's
d.js(isJsaChallenge/window.execDeep) HTML5-parser + arithmetic challenge—bypassing HTTP 202 / 429 anti-bot walls in milliseconds without spawning a 400MB headless browser.Deep Ad-Stripping & Real URL Resolution: Automatically strips sponsored ads on Brave (
data-type="ad",/a/redirect) and Sogou, and resolves encrypted redirect links (/link?url=,uigs_paratoken replay, BaiduLocationheaders, Bingu=a1...Base64 links) to clean target URLs.
📄 AI-Ready Content Extraction & GFM Markdown:
Combines
@mozilla/readabilitywith container-level noise stripping (stripChromeNoiseWithGuard) to remove<nav>,<aside>,<footer>, sidebars, and breadcrumbs while preserving<article><header><h1>article titles.Full
format: "markdown"support (turndown+ GFM tables/fenced code blocks) across both Readability and container-fallback paths, plus automatic GBK/GB2312 decoding andstartIndexpagination.
📚 Built-in Context7 Official Library Docs: Native
resolveLibraryIdandqueryDocstools fetch up-to-date, version-specific documentation and code snippets without running a separate Context7 MCP server.🌏 Split-Horizon Proxy (
PROXY_ENGINES) & Clash Fake-IP Ready:Route only overseas engines (
duckduckgo,exa,brave,startpage) through your proxy while keeping domestic engines on fast direct connections.Built-in
FAKE_IP_CIDRS(198.18.0.0/15enabled by default) works seamlessly with Clash TUN / Fake-IP setups while enforcing strict SSRF protection against private-network access.
Related MCP server: grok-mcp
🏗️ Architecture
flowchart TB
subgraph Clients["🤖 Entrypoints"]
MCP["MCP Server<br/>(STDIO / Streamable HTTP / SSE)"]
CLI["CLI One-Shot Commands<br/>(my-websearch search / fetch-*)"]
Daemon["Local HTTP Daemon<br/>(127.0.0.1:3210 · /health · /metrics)"]
end
subgraph Core["🧠 Search & Fetch Orchestrator"]
Router["Language-Aware Auto Router<br/>ZH → Baidu | EN/Tech → Bing + DuckDuckGo"]
Cascade["minResults Cascade & Circuit Breaker<br/>(Auto Fallback + 5min TTL Cache)"]
Ranker["Cross-Engine URL Deduplication & Relevance Ranking"]
end
subgraph Transport["🛡️ Anti-Bot & Secure Transport Layer"]
Wreq["wreq-js Native Chrome TLS/H2 Fingerprint<br/>+ Automatic Session Cookie Jars"]
Solvers["Millisecond Challenge Solvers<br/>Startpage Anubis PoW | DDG JSA Solver"]
PW["Playwright Stealth Browser Fallback<br/>(Auto-discovers System Edge/Chrome or CDP)"]
Guard["SSRF Guard & Clash Fake-IP Support<br/>(PROXY_ENGINES Split Routing + 198.18.0.0/15)"]
end
subgraph Engines["🌍 9 Search Engines + 6 Content/Docs Tools"]
CN["🇨🇳 Direct Engines<br/>Bing · Baidu · CSDN · Juejin · Sogou"]
INTL["🌐 Global Engines<br/>DuckDuckGo · Brave · Startpage · Exa"]
Docs["📚 Content & Official Docs<br/>fetchWebContent · Context7 · GitHub/Gitee · CSDN/Juejin"]
end
Clients --> Core
Core --> Transport
Transport --> Engines🌍 9 Search Engines Overview
Engine | Connectivity (Mainland China) | API Key | Anti-Bot & Parsing Architecture | Best For |
| 🇨🇳 Direct | None |
| General technical search, mixed EN/ZH queries |
| 🇨🇳 Direct | None | Parallel redirect resolution ( | Chinese news, documentation, domestic communities |
| 🇨🇳 Direct | None |
| Chinese error messages, debugging notes |
| 🇨🇳 Direct | None | Official Juejin search API integration | Modern frontend/backend/mobile Chinese articles |
| 🇨🇳 Direct | None | Desktop | WeChat ecosystem articles & Chinese long-tail queries |
| 🌐 Proxy in CN | None | Preload | English technical search, open-source discussions |
| 🌐 Proxy in CN | None | Built-in Anubis SHA-256 PoW solver + | Google-backed search results with high privacy |
| 🌐 Proxy in CN | None |
| Independent English index & technical blogs |
| 🌐 Direct API | Optional Free Key | Official semantic search API (enabled via | AI papers, GitHub repositories, semantic lookup |
🛠️ 7 MCP Tools Reference
Tool Name | Purpose | Key Parameters & Highlights |
| Multi-engine federated web search | Supports |
| Generic web page & Markdown extraction | Supports |
| Resolve package name to Context7 ID | Turns |
| Fetch official versioned library docs | Retrieves code examples and API docs by Context7 ID (supports version pinning like |
| Fetch GitHub or Gitee repo README | Supports HTTPS, SSH, |
| Fetch full CSDN blog article | Clean |
| Fetch full Juejin article | Direct API extraction returning clean article body; pass |
🚀 Quick Start
1. Run Immediately with NPX
# Basic startup (STDIO + HTTP)
npx -y my-websearch@latest
# 🇨🇳 Recommended for Mainland China (Overseas engines via proxy, domestic engines direct)
USE_PROXY=true PROXY_URL=http://127.0.0.1:7890 PROXY_ENGINES=duckduckgo,exa,brave,startpage npx -y my-websearch@latest2. Configure in MCP Clients
🔹 Claude Desktop / Cursor / Windsurf / Cline (mcpServers Config)
{
"mcpServers": {
"my-websearch": {
"command": "npx",
"args": ["-y", "my-websearch@latest"],
"env": {
"MODE": "stdio",
"DEFAULT_SEARCH_ENGINE": "auto",
"DEFAULT_MIN_RESULTS": "5",
"USE_PROXY": "true",
"PROXY_URL": "http://127.0.0.1:7890",
"PROXY_ENGINES": "duckduckgo,exa,brave,startpage",
"FAKE_IP_CIDRS": "198.18.0.0/15"
}
}
}
}💡 Windows CMD Wrapper (if your client requires
cmd /cto locatenpx):{ "mcpServers": { "my-websearch": { "command": "cmd", "args": ["/c", "npx", "-y", "my-websearch@latest"], "env": { "MODE": "stdio", "DEFAULT_SEARCH_ENGINE": "auto", "SYSTEMROOT": "C:/Windows" } } } }
🔹 Cherry Studio (STDIO or Streamable HTTP)
STDIO Mode: Use the standard JSON config above.
Streamable HTTP Mode (start server with
npx my-websearch@latest, default port3211):{ "mcpServers": { "web-search": { "name": "MyWebSearch", "type": "streamableHttp", "baseUrl": "http://localhost:3211/mcp" } } }
🔹 Client Config File Locations
Client / Harness | Config File Location |
Claude Desktop | macOS: |
Cursor |
|
ZCode / zcode |
|
Gemini CLI / Antigravity |
|
DSH (DeepSeek Harness) |
|
Reasonix |
|
💻 CLI, Local HTTP Daemon & Agent Skill
Beyond MCP, my-websearch provides a fast CLI and a long-lived Local HTTP Daemon (127.0.0.1:3210) that keeps connection pools, 5-minute search caches, and solved anti-bot sessions warm across calls.
1. Install the Agent Skill
npx skills add https://gitee.com/wtznicy/my-websearch --skill my-websearch2. Common CLI & Daemon Commands
# Install globally
npm install -g my-websearch
# Start the background-ready local HTTP daemon (port 3210)
my-websearch serve
# Check daemon health and active configuration
my-websearch status --json
# Run a one-shot search (automatically reuses the local daemon if running)
my-websearch search "Model Context Protocol specification" --limit 5 --min-results 5 --json
# Extract clean Markdown from any web page
my-websearch fetch-web "https://blog.vuejs.org/posts/vue-3-5" --max-chars 15000 --json
# Clear the 5-minute in-memory search cache
my-websearch cache-clear📊 Prometheus Metrics & Health Endpoints: The daemon exposes
GET /health,GET /status, andGET /metrics(engine latency, success/failure counters, cache hit ratio, memory/uptime). See docs/http-api.md for the complete HTTP API.
📖 Tool Usage Examples
1. search — Multi-Engine & Multi-Query Search
{
"queries": [
"Rust tokio async runtime tutorial",
"tokio spawn blocking best practices"
],
"engines": ["duckduckgo", "bing", "startpage"],
"limit": 8,
"minResults": 6
}2. fetchWebContent — Clean Markdown Extraction
{
"url": "https://blog.vuejs.org/posts/vue-3-5",
"format": "markdown",
"readability": true,
"includeLinks": true,
"maxChars": 20000,
"startIndex": 0
}When
truncated: true, pass the returnednextStartIndexasstartIndexin your next call to page through long documents.
3. resolveLibraryId + queryDocs — Official Library Documentation
// Step 1: Resolve library ID
{
"libraryName": "Next.js",
"query": "App Router middleware authentication"
}
// Step 2: Query version-specific documentation snippets
{
"libraryId": "/vercel/next.js",
"query": "how to protect routes in middleware.ts",
"limit": 5
}⚙️ Environment Variables Reference
Variable | Default | Options / Format | Description |
|
|
| Default engine. |
|
| Comma-separated engines | Parallel engine group for English/technical queries under |
|
| Comma-separated engines | Primary engine group for Chinese queries under |
|
| Non-negative integer | Automatically cascades to other engines when initial engines return fewer than |
| empty (all) | Comma-separated engines | Restrict which search engines can be used |
|
|
| Enable explicit HTTP/HTTPS proxy (if unset, OS system proxy is auto-detected when needed) |
|
| Valid proxy URL | Proxy URL (automatically passed to both |
| empty (all) | Comma-separated engines | Recommended for Mainland China: |
|
| Comma-separated CIDRs | Required for Clash TUN / Fake-IP: treats DNS answers in these ranges as synthetic proxy IPs instead of blocking them as private IPs |
|
|
| Browser TLS/HTTP2 fingerprint target used for Bing HTTP requests |
|
|
| Set |
|
|
| Startpage uses the built-in Anubis SHA-256 PoW solver first; set |
| empty | Exa API Key | Optional: only required if you explicitly use the |
| empty | Context7 API Key | Optional: anonymous usage includes 200 requests/month per egress IP; set a free key (context7.com/dashboard) for higher quotas |
| empty | GitHub Personal Access Token | Optional: raises the rate limit for |
|
|
| Disable TLS verification for |
|
|
| MCP server transport mode |
|
|
| MCP HTTP/SSE listen port (CLI local daemon uses |
|
| Milliseconds | Idle TTL for MCP HTTP/SSE sessions before the reaper closes them (default 30 min) |
|
| Positive integer | Max retained MCP HTTP sessions; the least-recently-active ones are closed first |
|
| Milliseconds | How often the session reaper runs (default 5 min) |
|
| Non-negative integer | Max concurrent searches in daemon mode ( |
|
|
| Enable Prometheus metrics collection ( |
|
|
| Enable security audit logging (SSRF blocks, TLS overrides) |
|
|
| Logging verbosity ( |
🤝 Contributing & Acknowledgements
Issues and Pull Requests are welcome! To build and run the test suite locally:
npm install
npm run build
npm run test:vitest # Run 134+ Vitest unit tests
npm test # Run bounded-concurrency integration test suiteAcknowledgements
Author: wtznicy
This project evolved from Open-WebSearch (originally created by Aas-ee) — special thanks to the original author. Thanks also to:
wreq-js for native Chrome TLS/HTTP2 fingerprint impersonation and session cookie management
Context7 (Upstash) for powering
resolveLibraryIdandqueryDocsMozilla Readability & Turndown for clean article extraction and GFM Markdown conversion
Available Tools
7 toolsfetchCsdnArticleARead-onlyIdempotent
Fetch full article content from a CSDN article URL (blog.csdn.net /article/details/ only; for other sites use fetchWebContent)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, covering safety. The description adds the behavioral constraint that only CSDN article details URLs are accepted and that it fetches full content, which goes beyond the annotations and helps set expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and resource, then provides the URL restriction and a clear alternative. Every clause earns its place with zero redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (one required parameter, no output schema, and annotations covering safety), the description is largely complete. It specifies the action, scope, and usage boundary. The only minor gap is not describing the exact return format or error behavior, but this is not critical for a fetch tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines url as a uri string, but the description adds critical meaning: the URL must be a CSDN article details URL (blog.csdn.net /article/details/). This is essential for correct invocation and is not conveyed by the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action (fetch full article content), the target resource (CSDN article URL), and specifies the exact URL pattern (blog.csdn.net /article/details/). It also distinguishes itself from the sibling tool fetchWebContent by name, making it easy for an agent to select correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides a usage boundary: 'only; for other sites use fetchWebContent'. This tells the agent when to use this tool and which alternative to fall back on, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchGithubReadmeARead-onlyIdempotent
Fetch README content from a GitHub/Gitee repository URL (github.com or gitee.com repo URLs only; for other pages use fetchWebContent)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the URL scope restriction but doesn't disclose further behavioral details such as response format or potential rate limiting. With annotations carrying the safety burden, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero waste. The main action is front-loaded, followed by the URL restriction and the alternative tool, all in one breath.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter fetch tool with annotations covering safety, the description is fully adequate. It states the input scope, the output concept (README content), and routes other URLs elsewhere. No output schema is needed for such a simple read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does by clarifying the url parameter must be a GitHub or Gitee repository URL, which is essential meaning beyond the bare 'url' property name. It could be more precise (e.g., full format https://github.com/owner/repo), but it provides sufficient guidance for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Fetch README content from a GitHub/Gitee repository URL'. It also distinguishes itself from the sibling fetchWebContent by explicitly restricting input to github.com/gitee.com repo URLs, so an agent can tell them apart without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool ('github.com or gitee.com repo URLs only') and provides the alternative for other cases ('for other pages use fetchWebContent'). This gives clear routing with no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchJuejinArticleARead-onlyIdempotent
Fetch full article content from a Juejin(掘金) post URL (juejin.cn or article.juejin.cn /post/ only; for other sites use fetchWebContent)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| format | No | Output format (default: text). 'markdown' keeps fenced code blocks (with language) and GFM tables — recommended for technical posts |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds extra behavioral context by constraining the accepted URL to Juejin post URLs only, which is useful beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the purpose, scope, and primary alternative with zero filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only fetch tool with 2 parameters and no output schema, the description plus schema provides all needed context: what tool does, what URLs to pass, and what format options exist. Annotations cover side effects, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (url has no schema description), but the description compensates by specifying the allowed domains and path for the url parameter. The format parameter is already documented in the schema with enum and description, so the description's added value for semantics is solid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the specific verb/resource: 'Fetch full article content from a Juejin post URL'. It also specifies the exact URL pattern (juejin.cn or article.juejin.cn /post/ only) and names a sibling alternative (fetchWebContent), making it distinguishable from other fetch tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance ('from a Juejin post URL') and when-not-to-use guidance ('for other sites use fetchWebContent'), directly routing agents to the correct alternative. This is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchWebContentCRead-onlyIdempotent
Fetch content from a public HTTP(S) URL (supports Markdown files and normal web pages)
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | Return the raw response body (HTML/plain text) without extraction | |
| url | Yes | ||
| format | No | Content format: 'markdown' preserves fenced code blocks (with language) and GFM tables — better for technical docs | |
| maxChars | No | ||
| startIndex | No | Character offset to start reading from (for paging through long content) | |
| readability | No | ||
| includeLinks | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, open-world behavior, and the description does not contradict them. It adds only that the target should be public and that Markdown pages/pages are supported, but does not disclose extraction, paging, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler; it earns its place. It is concise but so minimal that it leans toward under-specification, which is penalized elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, no output schema, and specialized siblings, this one-liner leaves critical invocation details (defaults, extraction behavior, link inclusion, paging) undocumented. The agent lacks enough context to reliably call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, and the description doesn't compensate: it never explains raw, readability, includeLinks, maxChars, or startIndex behavior. The Markdown mention is the only added parameter-related context, aligning loosely with the format enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fetch content from a public HTTP(S) URL') and signals supported content types ('Markdown files and normal web pages'). It is clear, but it does not explicitly differentiate itself from the specialized sibling fetch tools (fetchGithubReadme, fetchCsdnArticle, fetchJuejinArticle), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No sentence addresses when to choose this tool over siblings; the specialized fetch tools aren't mentioned, and there are no exclusions or preconditions. The only usage signal is the general 'public HTTP(S) URL' phrase, which is implicit rather than explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryDocsARead-onlyIdempotent
Get up-to-date, version-specific official docs and code snippets for a Context7 library ID (e.g. /vercel/next.js, version-pinnable like /vercel/next.js@v15.1.8). Use with resolveLibraryId for official documentation lookups — direct, reliable, no proxy needed.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of code snippets to return (default 5) | |
| query | No | The question or task to get relevant documentation for (optional; defaults to an overview when omitted, e.g. 'how to set up middleware with auth') | |
| libraryId | Yes | Exact Context7-compatible library ID (e.g. /vercel/next.js, /packages/express; optional version like /vercel/next.js@v15.1.8) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent behavior. The description adds meaningful behavioral context: results are up-to-date, version-pinnable, and obtained without a proxy. It does not describe return format or pagination, but those are minor for a read-only docs lookup tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, and each sentence carries useful information: what it does and how to use it. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read-only tool with rich annotations and full schema coverage, the description covers purpose, usage workflow, and result expectations. It lacks an explicit return-value shape and edge-case behavior, but the overall picture is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already provides examples for libraryId, including version suffixes. The tool description repeats those examples rather than adding new semantic detail for query or limit, so it does not go beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('official docs and code snippets for a Context7 library ID'), with concrete examples. It also distinguishes this tool from web-fetch and search siblings by framing it as 'official documentation lookups' and naming resolveLibraryId as its companion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises using this tool with resolveLibraryId for official documentation lookups and gives a clear benefit ('direct, reliable, no proxy needed'). It does not explicitly state when to prefer alternatives like search or fetchWebContent, so exclusions are missing, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolveLibraryIdARead-onlyIdempotent
PREFERRED for official docs: resolve a library/package name to a Context7 library ID (e.g. /vercel/next.js). Use this (then queryDocs) FIRST when the task needs official library/framework documentation — more reliable than web search and works without a proxy.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of library matches to return (default 5) | |
| query | No | The user's question or task, used to rank results by relevance (optional; defaults to the library name when omitted, e.g. 'how to implement authentication') | |
| libraryName | Yes | The library or package name to search for (e.g. 'Next.js', 'express', 'prisma') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and non-destructive. The description adds operational context beyond annotations: proxy-independent behavior and reliability relative to web search. These behavioral traits help the agent decide under network constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first states purpose with an example, the second gives usage priority and comparison. There is no filler, and the key directives ('PREFERRED', 'FIRST') are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple lookup tool, the description covers when to use it, what it returns (a Context7 library ID with an example format), and how to chain it with queryDocs. With no output schema, the gap around empty results or error behavior is minor but not fully addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 100% of parameters with clear descriptions, so the baseline applies. The description adds a useful output example and implies the query-reranking behavior, but does not materially enhance parameter understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'resolve a library/package name to a Context7 library ID' — and provides a concrete example ('/vercel/next.js'). It clearly differentiates the tool from queryDocs and web search by positioning it as the resolution step before queryDocs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use the tool ('when the task needs official library/framework documentation'), in what order ('then queryDocs'), and why it beats alternatives ('more reliable than web search and works without a proxy'). This is direct, actionable selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchARead-onlyIdempotent
Search the web across multiple engines with no API key required. searchMode: omit/auto = server SEARCH_MODE; request/playwright force that mode. Engine guidance: for Chinese queries prefer engines=["baidu","sogou","csdn","juejin"] (domestic engines have far better Chinese content coverage); for English/official docs use bing; overseas engines (duckduckgo/brave/startpage/exa) need a proxy. Omit engines to use server-side auto routing. Cite the relevant result URLs as markdown links in your answer. For OFFICIAL library/framework documentation, prefer resolveLibraryId + queryDocs (more reliable, works without proxy). Use the site: operator for site-specific queries (e.g. "update site:docs.elastic.co").
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default: server DEFAULT_SEARCH_LIMIT, usually 10) | |
| query | No | Single search query (use queries for multi-intent fan-out) | |
| engines | No | Search engines to use (default: server default, which may auto-route by query language) | |
| queries | No | 1-4 queries run concurrently and merged (dedup by URL) — for covering multiple intents in one call (e.g. official docs / Chinese community / English) | |
| minResults | No | Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5; max 50 — larger values only burn the search deadline) | |
| searchMode | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint. The description adds significant behavioral context beyond those: no API key required, server-side auto-routing behavior, proxy dependencies for certain engines, concurrent query merging, and the minResults auto-run mechanism that can burn the search deadline. It does not contradict annotations. A minor gap is that it doesn't describe the exact return format, but the citing instruction covers how to present results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is relatively long but every sentence carries meaningful information: purpose, engine guidance, alternative tool routing, and a site operator tip. It is well-organized, front-loading the core purpose and key routing rules before secondary details. Minor redundancy exists ('Omit engines to use server-side auto routing' echoes the schema's default), but overall it's efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema and complex routing logic, the description is exceptionally complete. It covers when to use it, engine selection per locale, proxy constraints, concurrency, result thresholds, and even alternative tools. An agent can make correct decisions without further investigation, given the annotations and schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the baseline is 3. The description adds substantial meaning beyond the schema: it explains the semantics of searchMode (omit/auto vs request/playwright), gives engine-specific guidance (baidu/sogou for Chinese, bing for English, proxy for overseas), clarifies queries for multi-intent, and warns about minResults trade-offs. This compensates well for the 17% of parameters lacking schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search the web across multiple engines with no API key required.' It distinguishes itself from siblings by explicitly naming resolveLibraryId + queryDocs as the preferred alternative for official documentation, and it covers the full scope of search behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit, actionable usage guidance: engine selection per query language (Chinese vs English vs overseas), proxy requirements, auto-routing when engines omitted, the site: operator, and multi-intent fan-out via queries. It also states when NOT to use this tool (for official docs, prefer queryDocs). This is comprehensive and clearly routes the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.17- Changed
fetchJuejinArticle1 field changed- added
Input schema / properties / formatAdded value: +{ + "description": "Output format (default: text). 'markdown' keeps fenced code blocks (with language) and GFM tables — recommended for technical posts", + "enum": [ + "text", + "markdown" + ], + "type": "string" +}
- Changed
search2 fields changed- changed
Input schema / properties / minResults / descriptionPrevious value: -"Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5)"New value: +"Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5; max 50 — larger values only burn the search deadline)" - added
Input schema / properties / minResults / maximumAdded value: +50
2 tool updates
v1.0.15- Changed
fetchWebContent1 field changed- added
Input schema / properties / formatAdded value: +{ + "description": "Content format: 'markdown' preserves fenced code blocks (with language) and GFM tables — better for technical docs", + "enum": [ + "text", + "markdown" + ], + "type": "string" +}
- Changed
search8 fields changed- removed
Input schema / properties / engines / defaultRemoved value: -[ - "bing" -] - added
Input schema / properties / engines / descriptionAdded value: +"Search engines to use (default: server default, which may auto-route by query language)" - removed
Input schema / properties / limit / defaultRemoved value: -10 - added
Input schema / properties / limit / descriptionAdded value: +"Max results (default: server DEFAULT_SEARCH_LIMIT, usually 10)" - changed
Input schema / properties / minResults / descriptionPrevious value: -"Auto-run additional engines when fewer than this many results come back (default: disabled)"New value: +"Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5)" - added
Input schema / properties / queriesAdded value: +{ + "description": "1-4 queries run concurrently and merged (dedup by URL) — for covering multiple intents in one call (e.g. official docs / Chinese community / English)", + "items": { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + "maxItems": 4, + "minItems": 1, + "type": "array" +} - added
Input schema / properties / query / descriptionAdded value: +"Single search query (use queries for multi-intent fan-out)" - removed
Input schema / requiredRemoved value: -[ - "query" -]
7 tool updates
v1.0.11- First observed
fetchCsdnArticle - First observed
fetchGithubReadme - First observed
fetchJuejinArticle - First observed
fetchWebContent - First observed
queryDocs - First observed
resolveLibraryId - First observed
search
TDQS
Scored across 7 tools
Each tool has a well-defined target: search vs. generic fetching vs. site-specific fetching vs. official docs. The specialized fetchers explicitly delegate non-matching URLs to fetchWebContent, reducing misselection.
Most tools follow a verb-first camelCase pattern (fetchX, queryDocs, resolveLibraryId). 'search' is a bare verb without an object, and object specificity varies, but the pattern remains predictable.
Seven tools cover the server's purpose without redundancy or bloat. The count is within the ideal 3–15 range and each tool contributes a distinct capability.
The set covers generic web fetching, specialized extraction for CSDN, Juejin, and GitHub READMEs, multi-engine web search, and official documentation lookups. No critical functionality appears missing for a web search and content-fetch server.
Maintenance
Related MCP Connectors
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Free web search for AI agents. No API key required. Hosted MCP in active development.
Serper MCP — wraps the Serper Google Search API (serper.dev)
Related MCP Servers
- FlicenseAqualityAmaintenanceAn MCP server for SearXNG that provides web search capabilities with concise model-visible output while preserving full result payloads in metadata. It supports search, parallel fetching, URL extraction, and research workflows through both local stdio and streamable HTTP transports.74-
- FlicenseNot gradedqualityCmaintenanceMCP server providing web search, news search, and X/Twitter search capabilities via HTTP or stdio.-
- AlicenseNot gradedqualityCmaintenanceA multi-function Streamable HTTP MCP tool aggregation server that provides web search via Brave, Exa, and SearXNG with multi-key rotation and cross-provider fallback, and supports extensible tool families (URL fetch, code search, RAG) through a pluggable architecture.MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for multi-engine web search and web page fetching, supporting parallel search, content extraction, and optional LLM-powered search summarization and deep search.5MIT