Skip to main content
Glama
zouyuanqing

DeepSeek Web Search MCP

by zouyuanqing

DeepSeek Web Search MCP

M8ven Verified

一个独立的 stdio MCP 服务器,向 Codex 等客户端暴露搜索、研究和多轮研究工具:

  • web_search:保留 AnySearch、SearXNG、Tavily 外部源,也可把 DeepSeek 原生搜索 作为第四个源加入同一候选池,再统一去重、融合和排序。

  • web_research:通过 DeepSeek 官方 Anthropic 兼容 Messages API 调用 web_search_20250305,返回回答、结构化来源和引用摘要。

  • research_start / research_followup / research_close:带 TTL 和最大轮数的 内存研究 session,用于多轮追问。

DeepSeek 协议

DeepSeek Responses API 当前会静默忽略 web_search。正确入口是:

https://api.deepseek.com/anthropic/v1/messages

请求使用 web_search_20250305 服务端工具。只有响应中存在 web_search_tool_result 才视为搜索成功;DSML、模型自述或普通文本都不能作为 搜索成功证据。

实现契约参考 DeepSeek 官方 Harness 的 @deepseek-ai/dsh-web-search-deepseek 包。

仓库不包含 API Key、服务器地址、用户名或私钥。所有凭据均从进程环境读取。

Related MCP server: MCP MixSearch

构建

npm install
npm run build
npm test

质量计划的离线 replay 使用无密钥 fixture,不会调用线上 provider:

npm run quality:fixture
npm run quality:replay

fixture 生成 70 个分层 synthetic query(30 development / 40 held-out);结果写入 work/quality/baseline-report.json。 fixture 保存各 provider 的原始顺序、 状态、延迟、来源和可选 rerank 结果,用于比较当前 baseline 与 provider-aware RRF shadow challenger;synthetic 数据只用于工程回归,不替代人工 relevance 标注, shadow 结果不会改变 MCP 生产返回。

npm run quality:summary 打印各方法聚合指标和分类明细;npm run quality:diff 列出清洗后 nDCG@5 下降的具体 case;node scripts/quality-inspect.mjs <caseId> 打印单个 case 的输入、保留和被丢弃的来源。

除 official / current / ambiguous / exploratory 四类外,fixture 还包含两类 专门复现实测反馈的 held-out case:

  • syndication:同一篇文章的 /it/、/ar/、/en/ 多语言副本 + 跨域转载

  • aggregator-dominance:聚合站排在官方文档页之前,且 rerank 给出与线上实测 一致的低分(官方页 0.037 / 0.006)

需要采集真实四路原始运行时数据时执行:

npm run quality:capture -- "your query"

结果写入被 .gitignore 排除的 work/quality/runtime/latest.json,不会把真实 provider 响应或凭据提交到仓库。

配置

复制 .env.example 中的变量到进程环境或 Windows 用户环境。至少需要:

  • DEEPSEEK_API_KEY:web_research、web_search(backend="hybrid") 和 web_search(backend="deepseek-native")

  • WEB_SEARCH_BACKEND:web_search 的默认后端,可选 auto(默认)、 external、hybrid 或 deepseek-native

  • FAST_CLEANING_MODE:fast 清洗模式,可选 shadow(默认,只计算不改变结果)、 on 或 off

  • FAST_DOMAIN_CAP:fast 清洗的同域名保留上限,默认 2

  • DOCUMENTATION_QUERY_MODE:官方文档/API/source code 类 query 的质量路由, 默认 fast,避免低分第三方页面被 deep rerank 压过;可设为 deep opt-in

  • SOURCE_IDENTITY_MODE:同一篇文章的重复来源折叠,默认 on;设为 off 恢复 仅按精确 URL 去重的旧行为

  • SOURCE_IDENTITY_CROSS_HOST:跨域转载检测,默认 on;设为 off 时只折叠 同域的重复 URL 和多语言副本

  • SOURCE_IDENTITY_VERBOSE:默认 off;设为 on 时在 sourceIdentity.groups 中返回每个合并组的保留 URL、被合并 URL 和冲突标记

  • FRESHNESS_INTENT_WARNING:默认 on。检测到时效性 query 但调用方没有主动 使用 strict 时,返回一条建议性 warning(不改变本次请求行为)

  • RESEARCH_SESSION_TTL_MS、RESEARCH_SESSION_MAX_SESSIONS、 RESEARCH_SESSION_MAX_TURNS:研究 session 生命周期和容量

  • ANYSEARCH_API_KEY:可选,匿名模式限额更低

  • TAVILY_API_KEY:全球搜索

  • OPENROUTER_API_KEY:web_search(rerank=true) 使用的可选重排

  • SEARXNG_URL:自建实例地址

DEEPSEEK_SEARCH_BASE_URL 是 Anthropic SDK 的 base URL,SDK 会自动追加 /v1/messages。默认值是 https://api.deepseek.com/anthropic。

freshness 继续支持 any/day/week/month/year。freshness_mode 默认是 soft: 支持原生时间过滤的 provider 会传递时间范围,其他 provider 只提供查询提示并返回 未严格验证 warning。strict 模式只保留有可解析 publishedAt 且满足 cutoff 的来源。

调研快速演变的题目(模型版本、定价、发布说明、今天的状态)时建议显式传 freshness_mode: "strict"。soft 只给 warning,不会把无法证明时间的新来源剔除。

时效性建议

web_search、web_research、research_start 和 research_followup 都会检测 query 的时效性信号(latest、current、changelog、version、最新、当前、 版本、实时 等)。检测到信号但调用方没有使用 strict 时,返回一条明确的 建议 warning,并在 freshness.intent 中给出 recommendedFreshness。

本项目不会静默切换模式:warning 只提示应该传什么,本次请求仍按调用方传入的参数 执行。FRESHNESS_INTENT_WARNING=off 可关闭该提示。

web_search 保留原有的 SearchResult 形状:始终返回 query、scope、 provider、sources 和 warnings。新增的 backend 只决定结果从哪里来:

{
  "query": "DeepSeek web search API",
  "backend": "hybrid",
  "max_results": 5,
  "freshness": "week"
}

可选值:

  • external:旧行为,按 provider 顺序使用 AnySearch、SearXNG、Tavily。

  • hybrid:并行查询外部 provider 和 DeepSeek 原生搜索,把四路结果合并为 同一个候选池;quality: "balanced" / "deep" 时四路一起进入 rank fusion。

  • deepseek-native:调用与 web_research 相同的 DeepSeek 原生搜索, 返回 mode: "native"、模型回答和带 provider: "deepseek-native" 的来源。

  • auto:有 DEEPSEEK_API_KEY 时等价于 hybrid;没有 key 时等价于 external。原生一路失败不会丢掉其他 provider 的结果。

不传 backend 时使用 WEB_SEARCH_BACKEND;未设置该变量时使用 auto,有 DeepSeek key 的客户端默认会把 native 纳入四源候选池。完全保持旧的外部-only 行为时设置 WEB_SEARCH_BACKEND=external 或传 backend: "external"。示例环境 默认设为 hybrid,方便直接体验四源融合。 deepseek-native 不执行 OpenRouter rank fusion;hybrid 只有在 quality: "balanced" / "deep" 或旧 rerank: true 时才执行外部重排。

原生接口没有与本项目 scope / freshness 完全等价的过滤参数。scope 仍按查询语言解析并原样返回;freshness 会作为查询提示传给原生搜索,但不应 被理解为严格的服务端时间过滤。需要完全排除 DeepSeek native 时使用 external; 需要保留 native 参与统一融合时使用 hybrid 或默认 auto。

deep 仍表示多 provider 召回加一次 rerank,不表示多轮对话。多轮研究请使用 research_start 开启 session,再用 research_followup 继续,最后用 research_close 释放内存状态;MCP 进程重启会清空 session。

默认情况下,识别为官方文档/API/source code 的 query 会自动从 deep 路由到 fast, 并在 warnings 中说明。研究 follow-up 不再把上一轮模型答案原文拼回下一次搜索, 只保留来源 URL 作为检索提示;native provider 会被要求优先引用 primary/official source。

引用质量与来源独立性

合成答案里最常见的两个问题不是"排错了",而是"同一篇文章被当成多个来源"和 "多源数字一致其实是互相抄"。1.5.0 针对这两点做了显式处理。

同一篇文章只算一个来源

SOURCE_IDENTITY_MODE=on(默认)会在重排和 rerank 之前折叠重复来源:

  • 规范化 URL:去 tracking 参数、/amp、index.html、尾斜杠

  • 折叠语言前缀:/it/…、/ar/…、/en/… 指向同一篇文章时合并为一条

  • 折叠 ?lang= / ?locale= / ?hl= 之类的语言参数

  • 跨域转载检测:不同域名下 slug 相同且标题高度相似时合并

合并时保留低权威风险的一方(官方/文档站优先于未知域名,未知域名优先于聚合站和 内容农场),并只回填缺失字段,不会覆盖已保留来源自己的标题或日期。日期冲突会 计入 sourceIdentity.conflictingDateCount 并产生 warning,而不是静默选一个。

返回结构新增:

{
  "sourceIdentity": {
    "applied": true,
    "inputCount": 5,
    "mergedCount": 2,
    "languageVariantCount": 1,
    "crossHostCopyCount": 1,
    "conflictingDateCount": 0,
    "independence": {
      "independentDomains": 2,
      "dominantDomainShare": 0.6,
      "lowAuthorityShare": 0.4,
      "level": "low"
    }
  }
}

research_* 系列的 sourceQuality 同时返回 independentDomains 和 independence,用于判断这个答案是"多源印证"还是"单一来源改写"。

聚合站和内容农场

src/hosts.ts 维护显式名单,包含实测中反复出现的内容农场 (ofox.ai、techsy.io、taskade.com 等)。名单内的域名:

  • 永远不获得权威性加分

  • fast 清洗中被乘性降权(系数 0.35)

乘性而非加性是有原因的:实测中排名第 1 的 Reddit 结果得分为 1/1 + 0 − 0.12 = 0.88,而排名第 2 的官方文档页只有 1/2 + 0.7×0.4 = 0.78——有界的加性惩罚永远无法把第一名挤下去。乘性降权 才可以让明确的低质信号压过更好的原始位置。

未知域名保持中立,不做猜测性惩罚。

隐私

本项目不收集也不上传任何遥测数据。查询只会发送给你自己配置的检索/重排提供商, 凭据仅从进程环境读取,缓存只存在于内存。详见 PRIVACY.md。

本地运行

npm run doctor -- --json
node dist/index.js

--doctor 只在终端输出各 provider 的健康状态,不会把健康检查注册成第三个 MCP 工具。

搜索查询会发送给用户配置的搜索提供商;选择 hybrid、deepseek-native 或 auto 时,查询还会发送给 DeepSeek;启用外部重排后,候选结果摘要还会发送给 OpenRouter。项目本身不包含遥测。

质量控制与融合重排

web_search 未配置后端时默认使用 quality: "fast" 和 backend: "auto";有 DeepSeek key 时会把 native 加入多源候选池,没有 key 时退回 external。启用外部 重排后,候选结果摘要才会发送给 OpenRouter。

fast 默认使用 FAST_CLEANING_MODE=shadow:会计算权威性、聚合站降权、同域名 上限和 provider coverage,但不改变旧客户端的返回顺序。完成 held-out 评估后可 切换为 on;off 完全关闭清洗。

offline replay(npm run quality:replay,70 个 synthetic case)当前测得:

方法

mean nDCG@5

officialInTop1

重复来源占比

独立来源数

current(清洗前)

0.9383

0.857

0.029

3.73

productionFast(清洗+去重)

0.9552

0.929

0.000

3.66

收益集中在实测反馈对应的两类 query:

  • aggregator-dominance(聚合站排在官方页之前):nDCG@5 0.7414 → 0.9949, officialInTop1 0 → 1

  • official / current / ambiguous / exploratory:0.9811 → 0.9811,无回退

  • syndication(同一文章多语言 + 跨域转载):重复槽位占比降到 0, 低权威来源占比 0.80 → 0.67

这些数字全部来自 synthetic fixture,只作为工程回归信号,不构成真实相关性 证明,因此 FAST_CLEANING_MODE 默认仍是 shadow;确认真实标注数据后可切到 on。 与之相对,SOURCE_IDENTITY_MODE 默认是 on,因为折叠"同一篇文章的重复 URL"是 正确性修复而不是排序偏好,不存在此消彼长的取舍。

quality

provider 候选

候选目标

行为

fast

每路 10

10

按 provider 顺序降级,不重排

balanced

每路 10

20

并行检索、重排并做 50/50 rank fusion

deep

每路 15

30

更大候选集、重排并做 50/50 rank fusion

在 backend: "hybrid" / 默认 auto 且存在 DeepSeek key 时,上表中的“每路”包含 deepseek-native;候选源采用轮询合并,避免某一 provider 填满前 max_results 而把 native 源挤掉。

重排使用 OpenRouter 的 nvidia/llama-nemotron-rerank-vl-1b-v2:free。融合同时保留原始排名和重排排名, 并按重排置信度缩放重排贡献:低于 0.35 的结果不获得重排分,0.35 到 1 之间线性缩放,避免“噪声但排名靠前”压过可靠候选。原始第一名只有达到 0.50 的重排相关性才会触发 top-1 保护;重排胜者若要凭权威性替换第一名,也 必须达到同一相关性下限。官方域名只获得有限先验,不能凭 URL 形态绕过重排置信度。

{
  "query": "DeepSeek Responses API web_search 是否已经失效",
  "scope": "global",
  "max_results": 8,
  "quality": "balanced",
  "backend": "external"
}

OpenRouter 缺 key、限流、超时或返回异常时,web_search 会回退到多 provider 原始顺序,并在结果中设置 rerank.applied=false 和 warning,不会让整次搜索失败。 旧的 rerank: true 仍作为兼容别名映射到 balanced;显式 quality 优先。

provider 搜索结果在进程内缓存 10 分钟,重排结果缓存 30 分钟,均采用有上限的 TTL/LRU,不写入磁盘。MCP 进程重启后缓存清空。嵌入模型不参与该链路。

Available Tools

2 tools
web_researchWeb ResearchC

Use DeepSeek native web search to research a question and return citeable sources.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
freshnessNoany
max_sourcesNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'native web search' and 'citeable sources' but does not disclose whether the tool performs live web access, how sources are selected, whether results are cached, or any rate limits or failure modes. For a research tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that conveys the core purpose and outcome without waste. It earns its place, though it could add a brief usage note without becoming bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% parameter coverage, the description is too thin. An agent cannot tell how to interpret the response, what 'citeable sources' means structurally, or how freshness and max_sources affect behavior. The sibling web_search also creates a routing ambiguity that is not resolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the three parameters. It does not explain the meaning of 'freshness' or 'max_sources', nor how they affect the research output. The query parameter is obvious from the description, but the other two are left entirely to the schema's enum and default values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('research'), a resource ('DeepSeek native web search'), and an outcome ('return citeable sources'). It is clear enough to distinguish from a generic search tool, though it does not explicitly name the sibling web_search or explain how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for research questions requiring citeable sources, but it does not explicitly state when to use this tool versus web_search, nor does it mention any exclusions or alternatives. The context signal of a sibling tool named web_search makes this gap noticeable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • First observedweb_research
    • First observedweb_search

TDQS

C2.8/5.0

Scored across 2 tools

Disambiguation2/5

Both `web_search` and `web_research` say they search the live web and return citeable sources, so an agent has little basis to choose between them. The descriptions overlap heavily and do not clearly define a simple-search versus deep-research boundary.

Naming Consistency5/5

Both tool names follow the same lowercase `web_<verb>` pattern, which is predictable and consistent. Although `search` and `research` are semantically close, the naming convention itself is uniform.

Tool Count3/5

Two tools is on the thin side for a web search server, and the second tool appears to be a near-duplicate of the first. Still, two tools is a defensible minimal set for a simple query-and-results workflow.

Completeness3/5

The server covers basic live-web searching and question-style research, but there are no tools for fetching specific URLs, filtering results, or managing research sessions. Deeper research workflows would likely need workarounds.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables advanced web search across multiple search engines (Brave, DuckDuckGo, Google, Bing, Yandex) with intelligent backend selection, full content extraction, and advanced filtering by time, language, geography, and content type.
    3
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Enables deep web search across multiple providers including Google, Bing, Brave, DuckDuckGo, and Perplexity, with support for comprehensive AI-powered research using intelligent multi-engine queries.
    2
    231 npm
    9
    MIT