Skip to main content
Glama

🔍 MyWebSearch

Keyless Multi-Engine AI Web Search & High-Purity Content Extraction Engine
MCP Server · CLI · Local HTTP Daemon · Skill-Guided Agent Workflows

🇨🇳 简体中文 | 🇺🇸 English

npm version npm downloads license GitHub stars M8ven Live Monitored


my-websearch is a full-stack web retrieval and documentation engine built for AI coding agents and MCP clients (Claude Desktop, Cursor, Windsurf, Cherry Studio, Cline, ZCode, etc.). It delivers 9-engine federated search, version-pinned official library documentation via Context7, and clean Markdown web page extraction—with zero paid API keys required.

  • 🌐 9-Engine Smart Orchestration: Combines direct domestic engines (Bing, Baidu, CSDN, Juejin, Sogou) and global engines (DuckDuckGo, Brave, Startpage, Exa) with language-aware auto routing, parallel multi-query execution (queries: string[]), cross-engine URL deduplication/ranking, circuit breakers, and minResults automatic cascade fallback.

  • 🛡️ Native Anti-Bot & Challenge Solvers (Zero-Browser Fast Path):

    • Chrome TLS/HTTP2 Fingerprint Impersonation: Powered by wreq-js with persistent session cookie jars (e.g., automatic Alibaba Cloud https_waf_cookie persistence on CSDN and Chrome 131/133 TLS handshakes on Bing/Brave/Startpage).

    • Millisecond Cryptographic & JS Challenge Solvers: Built-in pure-JS/Rust solvers for Startpage's Anubis SHA-256 Proof-of-Work (PoW) and DuckDuckGo's d.js (isJsaChallenge / window.execDeep) HTML5-parser + arithmetic challenge—bypassing HTTP 202 / 429 anti-bot walls in milliseconds without spawning a 400MB headless browser.

    • Deep Ad-Stripping & Real URL Resolution: Automatically strips sponsored ads on Brave (data-type="ad", /a/redirect) and Sogou, and resolves encrypted redirect links (/link?url=, uigs_para token replay, Baidu Location headers, Bing u=a1... Base64 links) to clean target URLs.

  • 📄 AI-Ready Content Extraction & GFM Markdown:

    • Combines @mozilla/readability with container-level noise stripping (stripChromeNoiseWithGuard) to remove <nav>, <aside>, <footer>, sidebars, and breadcrumbs while preserving <article><header><h1> article titles.

    • Full format: "markdown" support (turndown + GFM tables/fenced code blocks) across both Readability and container-fallback paths, plus automatic GBK/GB2312 decoding and startIndex pagination.

  • 📚 Built-in Context7 Official Library Docs: Native resolveLibraryId and queryDocs tools fetch up-to-date, version-specific documentation and code snippets without running a separate Context7 MCP server.

  • 🌏 Split-Horizon Proxy (PROXY_ENGINES) & Clash Fake-IP Ready:

    • Route only overseas engines (duckduckgo,exa,brave,startpage) through your proxy while keeping domestic engines on fast direct connections.

    • Built-in FAKE_IP_CIDRS (198.18.0.0/15 enabled by default) works seamlessly with Clash TUN / Fake-IP setups while enforcing strict SSRF protection against private-network access.


Related MCP server: grok-mcp

🏗️ Architecture

flowchart TB
    subgraph Clients["🤖 Entrypoints"]
        MCP["MCP Server<br/>(STDIO / Streamable HTTP / SSE)"]
        CLI["CLI One-Shot Commands<br/>(my-websearch search / fetch-*)"]
        Daemon["Local HTTP Daemon<br/>(127.0.0.1:3210 · /health · /metrics)"]
    end

    subgraph Core["🧠 Search & Fetch Orchestrator"]
        Router["Language-Aware Auto Router<br/>ZH → Baidu | EN/Tech → Bing + DuckDuckGo"]
        Cascade["minResults Cascade & Circuit Breaker<br/>(Auto Fallback + 5min TTL Cache)"]
        Ranker["Cross-Engine URL Deduplication & Relevance Ranking"]
    end

    subgraph Transport["🛡️ Anti-Bot & Secure Transport Layer"]
        Wreq["wreq-js Native Chrome TLS/H2 Fingerprint<br/>+ Automatic Session Cookie Jars"]
        Solvers["Millisecond Challenge Solvers<br/>Startpage Anubis PoW | DDG JSA Solver"]
        PW["Playwright Stealth Browser Fallback<br/>(Auto-discovers System Edge/Chrome or CDP)"]
        Guard["SSRF Guard & Clash Fake-IP Support<br/>(PROXY_ENGINES Split Routing + 198.18.0.0/15)"]
    end

    subgraph Engines["🌍 9 Search Engines + 6 Content/Docs Tools"]
        CN["🇨🇳 Direct Engines<br/>Bing · Baidu · CSDN · Juejin · Sogou"]
        INTL["🌐 Global Engines<br/>DuckDuckGo · Brave · Startpage · Exa"]
        Docs["📚 Content & Official Docs<br/>fetchWebContent · Context7 · GitHub/Gitee · CSDN/Juejin"]
    end

    Clients --> Core
    Core --> Transport
    Transport --> Engines

🌍 9 Search Engines Overview

Engine

Connectivity (Mainland China)

API Key

Anti-Bot & Parsing Architecture

Best For

bing

🇨🇳 Direct

None

wreq-js Chrome TLS impersonation + Base64 u=a1 URL decoding + optional Playwright fallback

General technical search, mixed EN/ZH queries

baidu

🇨🇳 Direct

None

Parallel redirect resolution (Location header extraction) + anti-bot page detection

Chinese news, documentation, domestic communities

csdn

🇨🇳 Direct

None

wreq-js session auto-persisting Alibaba Cloud https_waf_cookie + empty-shell auto-retry

Chinese error messages, debugging notes

juejin

🇨🇳 Direct

None

Official Juejin search API integration

Modern frontend/backend/mobile Chinese articles

sogou

🇨🇳 Direct

None

Desktop /web + Mobile m.sogou.com dual parser + uigs_para real URL resolution + ad filtering

WeChat ecosystem articles & Chinese long-tail queries

duckduckgo

🌐 Proxy in CN

None

Preload d.js JSONP + built-in isJsaChallenge (window.execDeep) HTML5/math solver

English technical search, open-source discussions

startpage

🌐 Proxy in CN

None

Built-in Anubis SHA-256 PoW solver + wreq-js session cookies (no browser needed)

Google-backed search results with high privacy

brave

🌐 Proxy in CN

None

wreq-js TLS fingerprint + strict sponsored-ad filtering (data-type="ad", /a/redirect)

Independent English index & technical blogs

exa

🌐 Direct API

Optional Free Key

Official semantic search API (enabled via EXA_API_KEY; fails fast if unset)

AI papers, GitHub repositories, semantic lookup


🛠️ 7 MCP Tools Reference

Tool Name

Purpose

Key Parameters & Highlights

search

Multi-engine federated web search

Supports query or parallel queries: string[], engines, limit, minResults (auto-cascades to additional engines when results are insufficient)

fetchWebContent

Generic web page & Markdown extraction

Supports format: "markdown" (preserves code blocks & tables), readability: true, includeLinks, startIndex pagination, GBK/UTF-8 auto-decoding, chrome noise stripping while rescuing <article><header><h1/h2> titles

resolveLibraryId

Resolve package name to Context7 ID

Turns "Next.js", "prisma", etc. into Context7 library IDs with trust & snippet count metadata

queryDocs

Fetch official versioned library docs

Retrieves code examples and API docs by Context7 ID (supports version pinning like "/vercel/next.js@v15.1.8")

fetchGithubReadme

Fetch GitHub or Gitee repo README

Supports HTTPS, SSH, .git URLs; Gitee uses official API (reachable in mainland China without proxy)

fetchCsdnArticle

Fetch full CSDN blog article

Clean #content_views extraction with automatic browser-cookie fallback if blocked

fetchJuejinArticle

Fetch full Juejin article

Direct API extraction returning clean article body; pass format: "markdown" to keep fenced code blocks (with language) and GFM tables


🚀 Quick Start

1. Run Immediately with NPX

# Basic startup (STDIO + HTTP)
npx -y my-websearch@latest

# 🇨🇳 Recommended for Mainland China (Overseas engines via proxy, domestic engines direct)
USE_PROXY=true PROXY_URL=http://127.0.0.1:7890 PROXY_ENGINES=duckduckgo,exa,brave,startpage npx -y my-websearch@latest

2. Configure in MCP Clients

🔹 Claude Desktop / Cursor / Windsurf / Cline (mcpServers Config)

{
  "mcpServers": {
    "my-websearch": {
      "command": "npx",
      "args": ["-y", "my-websearch@latest"],
      "env": {
        "MODE": "stdio",
        "DEFAULT_SEARCH_ENGINE": "auto",
        "DEFAULT_MIN_RESULTS": "5",
        "USE_PROXY": "true",
        "PROXY_URL": "http://127.0.0.1:7890",
        "PROXY_ENGINES": "duckduckgo,exa,brave,startpage",
        "FAKE_IP_CIDRS": "198.18.0.0/15"
      }
    }
  }
}

💡 Windows CMD Wrapper (if your client requires cmd /c to locate npx):

{
  "mcpServers": {
    "my-websearch": {
      "command": "cmd",
      "args": ["/c", "npx", "-y", "my-websearch@latest"],
      "env": {
        "MODE": "stdio",
        "DEFAULT_SEARCH_ENGINE": "auto",
        "SYSTEMROOT": "C:/Windows"
      }
    }
  }
}

🔹 Cherry Studio (STDIO or Streamable HTTP)

  • STDIO Mode: Use the standard JSON config above.

  • Streamable HTTP Mode (start server with npx my-websearch@latest, default port 3211):

    {
      "mcpServers": {
        "web-search": {
          "name": "MyWebSearch",
          "type": "streamableHttp",
          "baseUrl": "http://localhost:3211/mcp"
        }
      }
    }

🔹 Client Config File Locations

Client / Harness

Config File Location

Claude Desktop

macOS: ~/Library/Application Support/Claude/claude_desktop_config.jsonWindows: %APPDATA%\Claude\claude_desktop_config.json

Cursor

~/.cursor/mcp.json or workspace .cursor/mcp.json

ZCode / zcode

~/.zcode/cli/config.json → mcp.servers.mywebsearch.env

Gemini CLI / Antigravity

mcp_config.json → mcpServers.mywebsearch.env

DSH (DeepSeek Harness)

cordis.patch.yml → mcp-mywebsearch.env

Reasonix

~/.reasonix/config.toml


💻 CLI, Local HTTP Daemon & Agent Skill

Beyond MCP, my-websearch provides a fast CLI and a long-lived Local HTTP Daemon (127.0.0.1:3210) that keeps connection pools, 5-minute search caches, and solved anti-bot sessions warm across calls.

1. Install the Agent Skill

npx skills add https://gitee.com/wtznicy/my-websearch --skill my-websearch

2. Common CLI & Daemon Commands

# Install globally
npm install -g my-websearch

# Start the background-ready local HTTP daemon (port 3210)
my-websearch serve

# Check daemon health and active configuration
my-websearch status --json

# Run a one-shot search (automatically reuses the local daemon if running)
my-websearch search "Model Context Protocol specification" --limit 5 --min-results 5 --json

# Extract clean Markdown from any web page
my-websearch fetch-web "https://blog.vuejs.org/posts/vue-3-5" --max-chars 15000 --json

# Clear the 5-minute in-memory search cache
my-websearch cache-clear

📊 Prometheus Metrics & Health Endpoints: The daemon exposes GET /health, GET /status, and GET /metrics (engine latency, success/failure counters, cache hit ratio, memory/uptime). See docs/http-api.md for the complete HTTP API.


📖 Tool Usage Examples

1. search — Multi-Engine & Multi-Query Search

{
  "queries": [
    "Rust tokio async runtime tutorial",
    "tokio spawn blocking best practices"
  ],
  "engines": ["duckduckgo", "bing", "startpage"],
  "limit": 8,
  "minResults": 6
}

2. fetchWebContent — Clean Markdown Extraction

{
  "url": "https://blog.vuejs.org/posts/vue-3-5",
  "format": "markdown",
  "readability": true,
  "includeLinks": true,
  "maxChars": 20000,
  "startIndex": 0
}

When truncated: true, pass the returned nextStartIndex as startIndex in your next call to page through long documents.

3. resolveLibraryId + queryDocs — Official Library Documentation

// Step 1: Resolve library ID
{
  "libraryName": "Next.js",
  "query": "App Router middleware authentication"
}

// Step 2: Query version-specific documentation snippets
{
  "libraryId": "/vercel/next.js",
  "query": "how to protect routes in middleware.ts",
  "limit": 5
}

⚙️ Environment Variables Reference

Variable

Default

Options / Format

Description

DEFAULT_SEARCH_ENGINE

auto

auto, bing, baidu, csdn, juejin, sogou, duckduckgo, brave, startpage, exa

Default engine. auto routes Chinese queries to AUTO_ROUTE_ZH_ENGINES and English/technical queries to AUTO_ROUTE_EN_ENGINES

AUTO_ROUTE_EN_ENGINES

bing,duckduckgo

Comma-separated engines

Parallel engine group for English/technical queries under auto routing

AUTO_ROUTE_ZH_ENGINES

baidu

Comma-separated engines

Primary engine group for Chinese queries under auto routing (auto-cascades via minResults when needed)

DEFAULT_MIN_RESULTS

5

Non-negative integer

Automatically cascades to other engines when initial engines return fewer than N results

ALLOWED_SEARCH_ENGINES

empty (all)

Comma-separated engines

Restrict which search engines can be used

USE_PROXY

false

true, false

Enable explicit HTTP/HTTPS proxy (if unset, OS system proxy is auto-detected when needed)

PROXY_URL

http://127.0.0.1:7890

Valid proxy URL

Proxy URL (automatically passed to both axios and wreq-js native TLS sessions)

PROXY_ENGINES

empty (all)

Comma-separated engines

Recommended for Mainland China: duckduckgo,exa,brave,startpage. Routes only listed engines via proxy while domestic engines stay direct

FAKE_IP_CIDRS

198.18.0.0/15

Comma-separated CIDRs

Required for Clash TUN / Fake-IP: treats DNS answers in these ranges as synthetic proxy IPs instead of blocking them as private IPs

BING_IMPERSONATE_TARGET

chrome131

wreq-js browser target

Browser TLS/HTTP2 fingerprint target used for Bing HTTP requests

BING_PLAYWRIGHT_FALLBACK

true

true, false

Set false to skip launching Playwright when Bing is challenged (saves ~400MB RAM and lets minResults cascade to lighter engines)

STARTPAGE_PLAYWRIGHT_FALLBACK

true

true, false

Startpage uses the built-in Anubis SHA-256 PoW solver first; set false to disable Playwright fallback if PoW fails

EXA_API_KEY

empty

Exa API Key

Optional: only required if you explicitly use the exa engine (get a free key at dashboard.exa.ai)

CONTEXT7_API_KEY

empty

Context7 API Key

Optional: anonymous usage includes 200 requests/month per egress IP; set a free key (context7.com/dashboard) for higher quotas

GITHUB_TOKEN

empty

GitHub Personal Access Token

Optional: raises the rate limit for fetchGithubReadme (anonymous raw quota returns 403 under frequent re-fetching, failing the fetch entirely)

FETCH_WEB_INSECURE_TLS

false

true, false

Disable TLS verification for fetchWebContent only (use only for legacy sites with broken certificate chains)

MODE

both

both, http, stdio

MCP server transport mode

PORT

3211

1-65535

MCP HTTP/SSE listen port (CLI local daemon uses 3210 by default)

MCP_SESSION_TTL_MS

1800000

Milliseconds

Idle TTL for MCP HTTP/SSE sessions before the reaper closes them (default 30 min)

MCP_MAX_SESSIONS

100

Positive integer

Max retained MCP HTTP sessions; the least-recently-active ones are closed first

MCP_SESSION_REAPER_MS

300000

Milliseconds

How often the session reaper runs (default 5 min)

MAX_CONCURRENT_SEARCHES

0

Non-negative integer

Max concurrent searches in daemon mode (0 = unlimited)

METRICS_ENABLED

false

true, false

Enable Prometheus metrics collection (GET /metrics)

SECURITY_AUDIT

false

true, false

Enable security audit logging (SSRF blocks, TLS overrides)

LOG_LEVEL

info

quiet, debug, info, warn, error

Logging verbosity (quiet silences startup and runtime logs)


🤝 Contributing & Acknowledgements

Issues and Pull Requests are welcome! To build and run the test suite locally:

npm install
npm run build
npm run test:vitest   # Run 134+ Vitest unit tests
npm test              # Run bounded-concurrency integration test suite

Acknowledgements

Author: wtznicy

This project evolved from Open-WebSearch (originally created by Aas-ee) — special thanks to the original author. Thanks also to:

  • wreq-js for native Chrome TLS/HTTP2 fingerprint impersonation and session cookie management

  • Context7 (Upstash) for powering resolveLibraryId and queryDocs

  • Mozilla Readability & Turndown for clean article extraction and GFM Markdown conversion

Available Tools

7 tools
fetchCsdnArticleA
Read-onlyIdempotent

Fetch full article content from a CSDN article URL (blog.csdn.net /article/details/ only; for other sites use fetchWebContent)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, covering safety. The description adds the behavioral constraint that only CSDN article details URLs are accepted and that it fetches full content, which goes beyond the annotations and helps set expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource, then provides the URL restriction and a clear alternative. Every clause earns its place with zero redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (one required parameter, no output schema, and annotations covering safety), the description is largely complete. It specifies the action, scope, and usage boundary. The only minor gap is not describing the exact return format or error behavior, but this is not critical for a fetch tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only defines url as a uri string, but the description adds critical meaning: the URL must be a CSDN article details URL (blog.csdn.net /article/details/). This is essential for correct invocation and is not conveyed by the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action (fetch full article content), the target resource (CSDN article URL), and specifies the exact URL pattern (blog.csdn.net /article/details/). It also distinguishes itself from the sibling tool fetchWebContent by name, making it easy for an agent to select correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides a usage boundary: 'only; for other sites use fetchWebContent'. This tells the agent when to use this tool and which alternative to fall back on, leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetchGithubReadmeA
Read-onlyIdempotent

Fetch README content from a GitHub/Gitee repository URL (github.com or gitee.com repo URLs only; for other pages use fetchWebContent)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

TDQS

A4.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds the URL scope restriction but doesn't disclose further behavioral details such as response format or potential rate limiting. With annotations carrying the safety burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero waste. The main action is front-loaded, followed by the URL restriction and the alternative tool, all in one breath.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter fetch tool with annotations covering safety, the description is fully adequate. It states the input scope, the output concept (README content), and routes other URLs elsewhere. No output schema is needed for such a simple read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by clarifying the url parameter must be a GitHub or Gitee repository URL, which is essential meaning beyond the bare 'url' property name. It could be more precise (e.g., full format https://github.com/owner/repo), but it provides sufficient guidance for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Fetch README content from a GitHub/Gitee repository URL'. It also distinguishes itself from the sibling fetchWebContent by explicitly restricting input to github.com/gitee.com repo URLs, so an agent can tell them apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool ('github.com or gitee.com repo URLs only') and provides the alternative for other cases ('for other pages use fetchWebContent'). This gives clear routing with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetchJuejinArticleA
Read-onlyIdempotent

Fetch full article content from a Juejin(掘金) post URL (juejin.cn or article.juejin.cn /post/ only; for other sites use fetchWebContent)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
formatNoOutput format (default: text). 'markdown' keeps fenced code blocks (with language) and GFM tables — recommended for technical posts

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds extra behavioral context by constraining the accepted URL to Juejin post URLs only, which is useful beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the purpose, scope, and primary alternative with zero filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only fetch tool with 2 parameters and no output schema, the description plus schema provides all needed context: what tool does, what URLs to pass, and what format options exist. Annotations cover side effects, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50% (url has no schema description), but the description compensates by specifying the allowed domains and path for the url parameter. The format parameter is already documented in the schema with enum and description, so the description's added value for semantics is solid.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the specific verb/resource: 'Fetch full article content from a Juejin post URL'. It also specifies the exact URL pattern (juejin.cn or article.juejin.cn /post/ only) and names a sibling alternative (fetchWebContent), making it distinguishable from other fetch tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance ('from a Juejin post URL') and when-not-to-use guidance ('for other sites use fetchWebContent'), directly routing agents to the correct alternative. This is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetchWebContentC
Read-onlyIdempotent

Fetch content from a public HTTP(S) URL (supports Markdown files and normal web pages)

ParametersJSON Schema
NameRequiredDescriptionDefault
rawNoReturn the raw response body (HTML/plain text) without extraction
urlYes
formatNoContent format: 'markdown' preserves fenced code blocks (with language) and GFM tables — better for technical docs
maxCharsNo
startIndexNoCharacter offset to start reading from (for paging through long content)
readabilityNo
includeLinksNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, open-world behavior, and the description does not contradict them. It adds only that the target should be public and that Markdown pages/pages are supported, but does not disclose extraction, paging, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler; it earns its place. It is concise but so minimal that it leans toward under-specification, which is penalized elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, no output schema, and specialized siblings, this one-liner leaves critical invocation details (defaults, extraction behavior, link inclusion, paging) undocumented. The agent lacks enough context to reliably call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, and the description doesn't compensate: it never explains raw, readability, includeLinks, maxChars, or startIndex behavior. The Markdown mention is the only added parameter-related context, aligning loosely with the format enum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Fetch content from a public HTTP(S) URL') and signals supported content types ('Markdown files and normal web pages'). It is clear, but it does not explicitly differentiate itself from the specialized sibling fetch tools (fetchGithubReadme, fetchCsdnArticle, fetchJuejinArticle), so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No sentence addresses when to choose this tool over siblings; the specialized fetch tools aren't mentioned, and there are no exclusions or preconditions. The only usage signal is the general 'public HTTP(S) URL' phrase, which is implicit rather than explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queryDocsA
Read-onlyIdempotent

Get up-to-date, version-specific official docs and code snippets for a Context7 library ID (e.g. /vercel/next.js, version-pinnable like /vercel/next.js@v15.1.8). Use with resolveLibraryId for official documentation lookups — direct, reliable, no proxy needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of code snippets to return (default 5)
queryNoThe question or task to get relevant documentation for (optional; defaults to an overview when omitted, e.g. 'how to set up middleware with auth')
libraryIdYesExact Context7-compatible library ID (e.g. /vercel/next.js, /packages/express; optional version like /vercel/next.js@v15.1.8)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent behavior. The description adds meaningful behavioral context: results are up-to-date, version-pinnable, and obtained without a proxy. It does not describe return format or pagination, but those are minor for a read-only docs lookup tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, and each sentence carries useful information: what it does and how to use it. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter read-only tool with rich annotations and full schema coverage, the description covers purpose, usage workflow, and result expectations. It lacks an explicit return-value shape and edge-case behavior, but the overall picture is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already provides examples for libraryId, including version suffixes. The tool description repeats those examples rather than adding new semantic detail for query or limit, so it does not go beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('official docs and code snippets for a Context7 library ID'), with concrete examples. It also distinguishes this tool from web-fetch and search siblings by framing it as 'official documentation lookups' and naming resolveLibraryId as its companion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises using this tool with resolveLibraryId for official documentation lookups and gives a clear benefit ('direct, reliable, no proxy needed'). It does not explicitly state when to prefer alternatives like search or fetchWebContent, so exclusions are missing, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolveLibraryIdA
Read-onlyIdempotent

PREFERRED for official docs: resolve a library/package name to a Context7 library ID (e.g. /vercel/next.js). Use this (then queryDocs) FIRST when the task needs official library/framework documentation — more reliable than web search and works without a proxy.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of library matches to return (default 5)
queryNoThe user's question or task, used to rank results by relevance (optional; defaults to the library name when omitted, e.g. 'how to implement authentication')
libraryNameYesThe library or package name to search for (e.g. 'Next.js', 'express', 'prisma')

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and non-destructive. The description adds operational context beyond annotations: proxy-independent behavior and reliability relative to web search. These behavioral traits help the agent decide under network constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first states purpose with an example, the second gives usage priority and comparison. There is no filler, and the key directives ('PREFERRED', 'FIRST') are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple lookup tool, the description covers when to use it, what it returns (a Context7 library ID with an example format), and how to chain it with queryDocs. With no output schema, the gap around empty results or error behavior is minor but not fully addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of parameters with clear descriptions, so the baseline applies. The description adds a useful output example and implies the query-reranking behavior, but does not materially enhance parameter understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource — 'resolve a library/package name to a Context7 library ID' — and provides a concrete example ('/vercel/next.js'). It clearly differentiates the tool from queryDocs and web search by positioning it as the resolution step before queryDocs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when to use the tool ('when the task needs official library/framework documentation'), in what order ('then queryDocs'), and why it beats alternatives ('more reliable than web search and works without a proxy'). This is direct, actionable selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.17
    • ChangedfetchJuejinArticle1 field changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "Output format (default: text). 'markdown' keeps fenced code blocks (with language) and GFM tables — recommended for technical posts",
        +  "enum": [
        +    "text",
        +    "markdown"
        +  ],
        +  "type": "string"
        +}
    • Changedsearch2 fields changed
      • changedInput schema / properties / minResults / description
        Previous value: -"Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5)"New value: +"Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5; max 50 — larger values only burn the search deadline)"
      • addedInput schema / properties / minResults / maximum
        Added value: +50
  2. 2 tool updatesv1.0.15
    • ChangedfetchWebContent1 field changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "Content format: 'markdown' preserves fenced code blocks (with language) and GFM tables — better for technical docs",
        +  "enum": [
        +    "text",
        +    "markdown"
        +  ],
        +  "type": "string"
        +}
    • Changedsearch8 fields changed
      • removedInput schema / properties / engines / default
        Removed value: -[
        -  "bing"
        -]
      • addedInput schema / properties / engines / description
        Added value: +"Search engines to use (default: server default, which may auto-route by query language)"
      • removedInput schema / properties / limit / default
        Removed value: -10
      • addedInput schema / properties / limit / description
        Added value: +"Max results (default: server DEFAULT_SEARCH_LIMIT, usually 10)"
      • changedInput schema / properties / minResults / description
        Previous value: -"Auto-run additional engines when fewer than this many results come back (default: disabled)"New value: +"Auto-run additional engines when fewer than this many USABLE results (relevant, non-entry-page) come back (default: server DEFAULT_MIN_RESULTS, usually 5)"
      • addedInput schema / properties / queries
        Added value: +{
        +  "description": "1-4 queries run concurrently and merged (dedup by URL) — for covering multiple intents in one call (e.g. official docs / Chinese community / English)",
        +  "items": {
        +    "maxLength": 500,
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 4,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • addedInput schema / properties / query / description
        Added value: +"Single search query (use queries for multi-intent fan-out)"
      • removedInput schema / required
        Removed value: -[
        -  "query"
        -]
  3. 7 tool updatesv1.0.11
    • First observedfetchCsdnArticle
    • First observedfetchGithubReadme
    • First observedfetchJuejinArticle
    • First observedfetchWebContent
    • First observedqueryDocs
    • First observedresolveLibraryId
    • First observedsearch

TDQS

A4/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a well-defined target: search vs. generic fetching vs. site-specific fetching vs. official docs. The specialized fetchers explicitly delegate non-matching URLs to fetchWebContent, reducing misselection.

Naming Consistency4/5

Most tools follow a verb-first camelCase pattern (fetchX, queryDocs, resolveLibraryId). 'search' is a bare verb without an object, and object specificity varies, but the pattern remains predictable.

Tool Count5/5

Seven tools cover the server's purpose without redundancy or bloat. The count is within the ideal 3–15 range and each tool contributes a distinct capability.

Completeness5/5

The set covers generic web fetching, specialized extraction for CSDN, Juejin, and GitHub READMEs, multi-engine web search, and official documentation lookups. No critical functionality appears missing for a web search and content-fetch server.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    A
    maintenance
    An MCP server for SearXNG that provides web search capabilities with concise model-visible output while preserving full result payloads in metadata. It supports search, parallel fetching, URL extraction, and research workflows through both local stdio and streamable HTTP transports.
    7
    4
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    MCP server providing web search, news search, and X/Twitter search capabilities via HTTP or stdio.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    A multi-function Streamable HTTP MCP tool aggregation server that provides web search via Brave, Exa, and SearXNG with multi-key rotation and cross-provider fallback, and supports extensible tool families (URL fetch, code search, RAG) through a pluggable architecture.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server for multi-engine web search and web page fetching, supporting parallel search, content extraction, and optional LLM-powered search summarization and deep search.
    5
    MIT