Skip to main content
Glama
Biogod2020
by Biogod2020

dsh-bing-search

简体中文

为 DeepSeek Harness (DSH) 提供 Web 搜索,以一个小型 MCP 服务器实现,并由 curl_cffi 驱动。

search 顺序:

  1. 探测 DuckDuckGo HTML(html.duckduckgo.com)并将可达性缓存约 60 秒。在中国大陆,除非配置了代理,否则此探测通常会失败。

  2. 当 DDG 可达时,使用 DDG。

  3. 当 DDG 不可用、被限流(HTTP 202 / 验证挑战)或结果集为 quality_label=poor 时,回退到 Bing。

  4. 按语言路由 Bing:中文 / zh-* 市场访问 cn.bing.com,否则访问 www.bing.com。

每次搜索响应都包含 quality_score(0–1)和 quality_label(good / weak / poor)。将 poor 视为不可用(词典页面、首词元垃圾结果)。不要引用这些标题。

它为 DSH 智能体提供三个浏览器风格的工具:

  • mcp__web__search — 搜索公共网络并返回规范化的自然结果。

  • mcp__web__open — 打开公共网页并提取可读文本。

  • mcp__web__find — 在长页面中查找文本并返回附近上下文。

DSH agent
  -> @deepseek-ai/dsh-mcp-client
  -> dsh-bing-search (MCP/stdio)
  -> curl_cffi.AsyncSession(impersonate="chrome")
  -> html.duckduckgo.com          (if reachable)
  -> else cn.bing.com / www.bing.com

中国大陆: 如果没有代理或 VPN,DuckDuckGo 通常无法访问。这是预期情况。此时插件会使用 Bing,并将 warnings 设置为 duckduckgo_unreachable。MCP 子进程不会继承你的 shell 中的 HTTP_PROXY / HTTPS_PROXY(trust_env=False)。要强制使用代理,请在插件进程上设置 DSH_WEB_PROXY(例如,在 cordis 的 env: 映射中设置为 http://127.0.0.1:10808)。不要假定 DDG 在典型的中国大陆家庭或校园网络上可用。

社区插件:DeepSeek Harness 要求第三方插件使用 dsh-plugin GitHub 主题以便被发现。

最快安装:将本仓库交给智能体

如果你的编码智能体具有终端和文件系统访问权限(Codex、Claude Code、Pi、OpenCode 等),请粘贴以下内容:

Install this DeepSeek Harness plugin into my current DSH setup:
https://github.com/Biogod2020/dsh-bing-search

Read the repository README and INSTALL.md first. Install it with uv, detect my active
DSH profile, add it through cordis.patch.yml using the required `insert` patch form,
preserve all unrelated config, use the absolute path of the installed dsh-bing-search
executable, then verify that mcp__web__search, mcp__web__open, and mcp__web__find are
registered. Finally run one real web search smoke test and report what changed.

这是推荐的路径。INSTALL.md 包含一份为智能体编写的确定性安装契约。

Related MCP server: webmcp

手动安装

1. 安装可执行文件

需要 Python 3.10+。使用 uv:

uv tool install --force git+https://github.com/Biogod2020/dsh-bing-search.git

找到工具 bin 目录:

uv tool dir --bin

在下面的 DSH 配置中使用 dsh-bing-search(或 Windows 上的 dsh-bing-search.exe)的绝对路径。

如需开发而不是安装工具:

git clone https://github.com/Biogod2020/dsh-bing-search.git
cd dsh-bing-search
uv sync --extra dev

仓库包含 uv.lock,用于可复现的开发安装。

2. 将其添加到 DSH

DSH 配置文件将根 cordis.yml 与补丁层 cordis.patch.yml 结合。通过补丁层添加新插件时,条目必须包装在 insert 中:

- insert:
    - id: mcp-web
      name: '@deepseek-ai/dsh-mcp-client'
      config:
        serverName: web
        transport: stdio
        command: /ABSOLUTE/PATH/TO/dsh-bing-search
        args: []
        toolCallTimeoutMs: 30000
        failOnStartupError: true
        reconnect:
          enabled: true
          initialDelayMs: 500
          maxDelayMs: 30000
          maxAttempts: 10

不要向 cordis.patch.yml 添加裸的 - id: mcp-web 条目:裸条目会修补现有 ID,未知 ID 可能会被跳过。如果你直接编辑根 cordis.yml,普通的裸插件条目是正确的。参见 cordis.example.yml。

3. 验证

DSH 重新加载配置文件后,模型应该能看到:

mcp__web__search
mcp__web__open
mcp__web__find

然后让智能体搜索一些当前内容并打开一个结果。成功的往返验证了搜索访问和 MCP 注册。更改插件代码后,请重启 DSH(或 MCP 子进程);stdio 进程不会热重载 Python。

工具

{
  "query": "DeepSeek Harness GitHub",
  "count": 8,
  "offset": 0,
  "market": "en-US",
  "safe_search": "Moderate"
}

返回:

字段

含义

provider

duckduckgo 或 bing

title / url / snippet / rank

自然搜索结果

source_id

来自规范 URL 的稳定 ID

quality_score

查询与标题/摘要的 0–1 重叠度

quality_label

good / weak / poor

warnings

回退原因和质量说明

对于中文查询,请使用 market=zh-CN。如果查询包含 CJK 字符,即使 market 为 en-US,Bing 回退仍会使用 cn.bing.com。

DuckDuckGo 的 /l/?uddg= 和 Bing 的 /ck/a 重定向在可能的情况下会被解码。常见的跟踪参数会被去除,重复的 URL 会被合并。

对于人物、论文或带插图的博客,请先搜索作者姓名或简短专有名词。如果 quality_label 为 poor,不要再继续加长查询。中文学术元数据应属于专门的语料库(例如 CNKI),而不是这种通用网络搜索。

open

{
  "url": "https://example.com/article",
  "max_chars": 24000
}

使用 curl_cffi 获取公共 HTTP(S) 页面,执行 DNS/IP 检查和安全的重定向,限制响应大小,并提取可读文本 而无需执行 JavaScript。

open 专为类文章 HTML 而生。它不是浏览器。DSH 实际运行表明,天气及其他组件繁重的网站(tianqi.com、weather.com.cn 等)往往会产生导航残留或近乎空白的文本:Trafilatura 找不到主要文章,然后回退逻辑会导出整个 DOM。status 仍然可能是 ok。对于这些页面,请信任搜索的 snippet,或 open 一个更简单的文章 URL。不要期望获得实时温度、地图或其他由 JS 渲染的 UI。

find

{
  "url": "https://example.com/article",
  "pattern": "DeepSeek",
  "max_matches": 5,
  "context_chars": 700
}

返回匹配的区域,而不会将整个页面注入模型上下文。

为什么用三个工具,而不是一个巨大的 search_and_summarize 工具?

插件保持检索的确定性,并让 DSH 模型控制研究循环:

search -> inspect candidates -> open -> find / search again -> synthesize

插件负责 HTTP、解析、清理、缓存、引擎回退、来源追踪和质量标记。智能体决定搜索什么、信任哪些来源、何时重新表述查询,以及何时已收集到足够证据。智能体必须读取 quality_label 和 warnings。

配置

环境变量

默认值

用途

DSH_BING_SEARCH_URL

https://www.bing.com/search

仅当设置为非默认值(测试)时覆盖 Bing HTML 端点。否则根据语言选择主机

DSH_WEB_IMPERSONATE

chrome

curl_cffi 浏览器指纹

DSH_WEB_PROXY

空

HTTP/HTTPS/SOCKS 代理。进程使用 trust_env=False,不会继承 HTTP_PROXY

DSH_WEB_TIMEOUT_SECONDS

20

传输超时

DSH_WEB_CONNECT_TIMEOUT_SECONDS

8

连接超时

DSH_WEB_MAX_BODY_BYTES

5242880

open 的最大响应体大小

DSH_BING_MAX_BODY_BYTES

2097152

搜索页面的最大响应体大小

DSH_WEB_MAX_REDIRECTS

8

最大重定向次数

DSH_WEB_CONCURRENCY

8

进程内最大并发请求数

DSH_BING_CACHE_TTL_SECONDS

90

搜索缓存 TTL

DSH_WEB_CACHE_TTL_SECONDS

600

页面缓存 TTL

测试

离线测试(解析器、质量评分、语言环境路由、DDG 优先 / Bing 回退):

uv run pytest -m "not live"

实时冒烟测试:

RUN_LIVE_BING=1 uv run pytest -m live -s

标记名称仍然是 live / RUN_LIVE_BING。实时运行会先访问 DDG,仅当 DDG 不可用时才使用 Bing。

CI 覆盖 Python 3.10、3.12、3.13 和 3.14。

设计与安全说明

这是一个非官方的 DuckDuckGo HTML + Bing HTML 适配器。它不使用已退役的 Bing Search API。

  • DDG 标记解析代码位于 src/dsh_bing_search/providers/ddg.py。

  • Bing 标记解析代码位于 src/dsh_bing_search/providers/bing_parser.py。

  • 质量评分位于 src/dsh_bing_search/quality.py,并且与引擎无关。

  • 请求使用 curl_cffi.AsyncSession 并进行浏览器模拟。

  • 用户提供的页面 URL 仅限于公共 HTTP(S) 目标,并启用了安全的重定向处理。

  • 响应体大小受限。

  • 验证码 / 验证挑战 / HTTP 202 页面会报告为 status="blocked";插件不会尝试绕过它们。

  • www.bing.com 上的无头 Bing 通常返回结构有效但不相关的卡片。cn.bing.com 对某些热门中文查询有帮助;长尾名称和标题仍然可能坍缩到第一个词元。这就是质量标记的作用。

  • open 不会自动重试缓慢的目标站点;如有需要,请增加超时环境变量。

社区

DeepSeek Harness 目前处于开发者预览阶段,因此插件接口可能仍会演变。如需 DSH 相关的支持和发现:

欢迎贡献和解析器修复。

许可证

MIT

Available Tools

4 tools
findFind in Web PageA

Find a literal phrase in a page and return compact context windows around matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
patternYes
max_matchesNo
context_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYes
errorNo
statusYes
matchesNo
patternYes
source_idNo
total_matchesNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does reveal key behavior: matching is literal rather than regex or semantic, and the response consists of compact context windows around matches. However, it does not mention case sensitivity, failure modes, page loading behavior, or limits, leaving notable gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence contains the core action, the matching mode, and the response shape with no redundant words. It is front-loaded and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple tool, and the output schema likely covers return values. But with no annotations and no parameter documentation, it lacks details about max_matches behavior, exact context window semantics, and when to prefer sibling tools. It is minimally sufficient but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that 'pattern' is a literal phrase and 'context_chars' relates to compact context windows, but it does not explain 'max_matches', 'url', defaults, or the exact relationship between parameters and output. This is only partial compensation for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: finding a literal phrase in a page and returning compact context windows around matches. The word 'literal' helps distinguish it from the sibling 'search' tool, which implies broader or semantic search. This is a clear, specific purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this tool when an exact literal phrase is needed within a page. However, it does not explicitly say when not to use it or mention alternatives like 'search' or 'search_images'. The usage guidance is present only by implication.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openOpen Web PageA

Fetch a public HTTP(S) page with curl_cffi and return cleaned readable text.

Use after search when result snippets are insufficient. Private/local addresses are rejected, redirect targets use curl_cffi safe-follow mode, and response bytes are capped.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
textNo
errorNo
titleNo
statusYes
final_urlNo
source_idNo
truncatedNo
elapsed_msNo
content_typeNo
fetched_bytesNo
requested_urlYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and discloses several useful traits: public-only access, rejection of private/local addresses, safe-follow redirect mode, and a response byte cap. It could also mention error behavior or timeout handling, but the provided constraints are substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences front-load the purpose, then add usage context and behavioral constraints. No filler, every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the description covers URL type, output format, redirect behavior, and a cap. The main gap is max_chars semantics, which matters because there is no schema-level documentation and no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds that URL must be public HTTP(S), but it never explains the max_chars parameter or how the response cap relates to it. An agent cannot confidently tune max_chars based on this text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch'), resource ('public HTTP(S) page'), and output ('cleaned readable text'). This distinguishes it from siblings like search and search_images: it retrieves page content rather than result snippets or images.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use after search when result snippets are insufficient,' giving a clear trigger condition and relationship to the primary sibling. It also states a when-not: private/local addresses are rejected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_imagesSearch ImagesA

Search image indexes and rank results with pure text so vision is not required.

auto (default) tries Bing Images first and falls back to Wikimedia Commons when the top text score is below ~40, so one call yields a ranked set. bing_images parses Bing Images metadata (original URL / thumbnail / source page / title). commons queries Wikimedia Commons, a curated and licence-clear platform. Every result carries a 0-100 text score, a domain hint and explainable signals; pick the highest score, treat scores below ~40 as unverified, and optionally verify with find/open on the source page before downloading.

Args: query: What the image should depict. Compact concrete nouns plus the qualifier that uniquely identifies the subject (e.g. "复旦光华楼", "台州城墙"). "复旦光华楼" is better than "光华楼". Do not write whole sentences. If a compact query is still ambiguous or hits the wrong entity, write more (place, institution, year, type). count: Number of ranked image results to return, from 1 to 20. market: Locale such as en-US or zh-CN (Bing Images; Commons is language-neutral). provider: auto (default), bing_images, or commons.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNo
queryYes
marketNoen-US
providerNoauto

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
queryYes
marketNo
statusYes
resultsNo
providerNo
warningsNo
elapsed_msNo
returned_countNo
requested_countNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does so thoroughly. It discloses the ranking mechanism, the auto fallback threshold, what each provider does, and the exact result signals: 0-100 text score, domain hint, and explainable signals. It even tells the agent how to assess confidence and when verification is needed, which goes well beyond a minimal tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and behavior, and the Args section is logically organized. It is longer than typical descriptions, but that length is justified by the zero-coverage schema and the need to explain provider behavior and scoring. Minor redundancy exists because provider defaults and enum values are repeated from the schema, but the added context still earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's provider-switching complexity, fallback threshold, scoring semantics, and four parameters, the description provides everything needed to select and invoke it correctly. It explains query formulation, ranking confidence, provider differences, and optional verification workflow. The output schema covers return structure, so the description does not need to detail the exact JSON response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully compensate for the schema's lack of parameter documentation. It does: `query` has concrete examples and wording advice ('复旦光华楼' is better than '光华楼'), `count` is bounded 1-20, `market` is explained as locale-specific to Bing while Commons is language-neutral, and `provider` enumerates the options. This is excellent parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Search image indexes and rank results with pure text so vision is not required.' This clearly distinguishes the tool from the sibling `search`, `open`, and `find` by emphasizing image indexes and text-based ranking. The provider variants (bing_images, commons) further specify exactly what kind of image search this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives usable routing guidance: `auto` is the default, it falls back to Commons below ~40 text score, and results below ~40 should be treated as unverified. It also recommends verifying with `find`/`open` before downloading, which indirectly differentiates this search tool from sibling file/URL tools. It lacks an explicit 'when not to use this tool' statement, but the behavioral and provider guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedfind
    • First observedopen
    • First observedsearch
    • First observedsearch_images

TDQS

A4.3/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a clearly distinct action: web search, image search, page retrieval, and in-page phrase matching. Search and search_images are separated by media type, while open and find both operate on pages but serve complementary pre- and post-retrieval needs, so an agent can select without confusion.

Naming Consistency5/5

All tool names are short imperative verbs in snake_case: search, search_images, open, find. The only compound name, search_images, naturally follows a verb_noun pattern, and the overall naming is predictable and consistent.

Tool Count5/5

Four tools form a tightly scoped search-and-browse toolset. Each tool earns its place: web search, image search, full-page reading, and targeted phrase lookup. The count is neither thin nor bloated for the server's stated purpose.

Completeness5/5

The server covers the full core workflow: discovering content via web or image search, opening pages when snippets are insufficient, and locating specific phrases within pages. Pagination, locale, safesearch, and provider fallback options also cover important search variations, leaving no obvious dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for web search and content extraction using DuckDuckGo or SearXNG, with Playwright-based fetching and LLM-powered data extraction.
    140
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.
    MIT