Skip to main content
Glama
Biogod2020
by Biogod2020

dsh-bing-search

简体中文

DeepSeek Harness (DSH) 用のウェブ検索。小さな MCP サーバーとして実装され、curl_cffi を利用しています。

search の順序:

  1. DuckDuckGo HTML(html.duckduckgo.com)をプローブし、到達可能性を約60秒間キャッシュします。中国本土からは、プロキシが設定されていない限り、このプローブは頻繁に失敗します。

  2. DDG が到達可能な場合はそれを使用します。

  3. DDG がダウンしている、レート制限されている(HTTP 202 / チャレンジ)、または結果セットが quality_label=poor の場合は、Bing にフォールバックします。

  4. Bing を言語でルーティングします: 中国語 / zh-* マーケットは cn.bing.com、それ以外は www.bing.com です。

すべての検索レスポンスには quality_score (0–1) と quality_label (good / weak / poor) が含まれます。poor は使用不可(辞書ページ、先頭トークンのジャンク)として扱ってください。それらのタイトルを引用しないでください。

DSH エージェントに3つのブラウザ風ツールを提供します:

  • mcp__web__search — 公開ウェブを検索し、正規化されたオーガニック結果を返します。

  • mcp__web__open — 公開ウェブページを開き、読み取り可能なテキストを抽出します。

  • mcp__web__find — 長いページ内のテキストを検索し、周辺のコンテキストを返します。

DSH agent
  -> @deepseek-ai/dsh-mcp-client
  -> dsh-bing-search (MCP/stdio)
  -> curl_cffi.AsyncSession(impersonate="chrome")
  -> html.duckduckgo.com          (if reachable)
  -> else cn.bing.com / www.bing.com

中国本土: DuckDuckGo はプロキシや VPN なしでは到達できないことがよくあります。これは想定内です。プラグインはその場合 Bing を使用し、warnings に duckduckgo_unreachable を設定します。MCP 子プロセスはシェルの HTTP_PROXY / HTTPS_PROXY を継承しません(trust_env=False)。プロキシを強制するには、プラグインプロセスで DSH_WEB_PROXY を設定してください(例: cordis の env: マップ内で http://127.0.0.1:10808)。一般的な中国本土の家庭やキャンパスネットワークで DDG が機能するとは想定しないでください。

コミュニティプラグイン: DeepSeek Harness はサードパーティプラグインに対して、発見のために dsh-plugin GitHub トピックを使用するよう求めています。

最速のインストール: このリポジトリをエージェントに渡す

コーディングエージェントがターミナルとファイルシステムへのアクセス権を持っている場合(Codex、Claude Code、Pi、OpenCode など)、以下を貼り付けてください:

Install this DeepSeek Harness plugin into my current DSH setup:
https://github.com/Biogod2020/dsh-bing-search

Read the repository README and INSTALL.md first. Install it with uv, detect my active
DSH profile, add it through cordis.patch.yml using the required `insert` patch form,
preserve all unrelated config, use the absolute path of the installed dsh-bing-search
executable, then verify that mcp__web__search, mcp__web__open, and mcp__web__find are
registered. Finally run one real web search smoke test and report what changed.

これが推奨パスです。INSTALL.md には、エージェント向けに書かれた決定的なインストール契約が含まれています。

Related MCP server: webmcp

手動インストール

1. 実行可能ファイルをインストール

Python 3.10+ が必要です。uv を使用する場合:

uv tool install --force git+https://github.com/Biogod2020/dsh-bing-search.git

ツールの bin ディレクトリを見つけます:

uv tool dir --bin

以下の DSH 設定で dsh-bing-search(Windows では dsh-bing-search.exe)への絶対パスを使用してください。

ツールインストールの代わりに開発用の場合:

git clone https://github.com/Biogod2020/dsh-bing-search.git
cd dsh-bing-search
uv sync --extra dev

リポジトリには再現可能な開発インストール用の uv.lock が含まれています。

2. DSH に追加

DSH プロファイルはルート cordis.yml とパッチレイヤー cordis.patch.yml を組み合わせます。パッチレイヤー経由で新しいプラグインを追加する場合、エントリは insert でラップする必要があります:

- insert:
    - id: mcp-web
      name: '@deepseek-ai/dsh-mcp-client'
      config:
        serverName: web
        transport: stdio
        command: /ABSOLUTE/PATH/TO/dsh-bing-search
        args: []
        toolCallTimeoutMs: 30000
        failOnStartupError: true
        reconnect:
          enabled: true
          initialDelayMs: 500
          maxDelayMs: 30000
          maxAttempts: 10

cordis.patch.yml に裸の - id: mcp-web エントリを追加しないでください: 裸のエントリは既存の ID をパッチし、未知の ID はスキップされる可能性があります。ルート cordis.yml を直接編集している場合は、通常の裸のプラグインエントリで正しいです。cordis.example.yml を参照してください。

3. 確認

DSH がプロファイルをリロードした後、モデルは以下を表示するはずです:

mcp__web__search
mcp__web__open
mcp__web__find

次にエージェントに最新の何かを検索して1件の結果を開くよう依頼してください。ラウンドトリップが成功すれば、検索アクセスと MCP 登録の両方が検証されます。プラグインコードを変更した後は DSH(または MCP 子プロセス)を再起動してください。stdio プロセスは Python をホットリロードしません。

ツール

{
  "query": "DeepSeek Harness GitHub",
  "count": 8,
  "offset": 0,
  "market": "en-US",
  "safe_search": "Moderate"
}

戻り値:

フィールド

意味

provider

duckduckgo または bing

title / url / snippet / rank

オーガニック結果

source_id

正規URLからの安定ID

quality_score

クエリとタイトル/スニペットの0–1の重なり

quality_label

good / weak / poor

warnings

フォールバック理由と品質メモ

中国語のクエリには market=zh-CN を使用してください。クエリに CJK が含まれる場合、market が en-US でも Bing フォールバックは cn.bing.com を使用します。

DuckDuckGo の /l/?uddg= と Bing の /ck/a リダイレクトは、可能な場合にデコードされます。一般的なトラッキングパラメータは除去され、重複 URL はマージされます。

人物、論文、イラスト付きブログの場合、まず著者名または短い固有名詞を検索してください。quality_label が poor の場合は、クエリを長くし続けないでください。中国の学術メタデータは、この一般的なウェブ検索ではなく、専門コーパス(例: CNKI)に属します。

open

{
  "url": "https://example.com/article",
  "max_chars": 24000
}

curl_cffi を使用して公開 HTTP(S) ページを取得し、DNS/IP チェックと安全なリダイレクトを適用し、レスポンスサイズを制限し、JavaScript を実行せずに読み取り可能なテキストを抽出します。

open は記事のような HTML 向けに作られています。ブラウザではありません。実際の DSH 実行では、天気やその他のウィジェットが多いサイト(tianqi.com、weather.com.cn など)がナビゲーションチャームやほぼ空のテキストを返すことがよくありました。Trafilatura がメイン記事を見つけられず、フォールバックが DOM 全体をダンプします。status は依然として ok になる場合があります。そのようなページでは、検索 snippet を信頼するか、よりシンプルな記事 URL を open してください。ライブの気温、地図、その他の JS レンダリング UI を期待しないでください。

find

{
  "url": "https://example.com/article",
  "pattern": "DeepSeek",
  "max_matches": 5,
  "context_chars": 700
}

ページ全体をモデルコンテキストに注入せずに、一致する領域を返します。

なぜ1つの巨大な search_and_summarize ツールではなく3つのツールなのか?

プラグインは取得を決定論的に保ち、DSH モデルが調査ループを制御できるようにします:

search -> inspect candidates -> open -> find / search again -> synthesize

プラグインは HTTP、解析、クリーニング、キャッシュ、エンジンのフォールバック、来歴、品質マークを処理します。エージェントは何を検索するか、どのソースを信頼するか、いつクエリを再構成するか、いつ十分な証拠が収集されたかを決定します。エージェントは quality_label と warnings を読む必要があります。

設定

環境変数

デフォルト

目的

DSH_BING_SEARCH_URL

https://www.bing.com/search

デフォルト以外の値に設定された場合のみ Bing HTML エンドポイントを上書きします(テスト用)。それ以外の場合、ホストは言語によって選択されます

DSH_WEB_IMPERSONATE

chrome

curl_cffi ブラウザフィンガープリント

DSH_WEB_PROXY

empty

HTTP/HTTPS/SOCKS プロキシ。プロセスは trust_env=False を使用し、HTTP_PROXY を継承しません

DSH_WEB_TIMEOUT_SECONDS

20

転送タイムアウト

DSH_WEB_CONNECT_TIMEOUT_SECONDS

8

接続タイムアウト

DSH_WEB_MAX_BODY_BYTES

5242880

open の最大ボディサイズ

DSH_BING_MAX_BODY_BYTES

2097152

検索ページの最大ボディサイズ

DSH_WEB_MAX_REDIRECTS

8

最大リダイレクト数

DSH_WEB_CONCURRENCY

8

プロセス内の最大同時リクエスト数

DSH_BING_CACHE_TTL_SECONDS

90

検索キャッシュ TTL

DSH_WEB_CACHE_TTL_SECONDS

600

ページキャッシュ TTL

テスト

オフラインテスト(パーサー、品質スコア、ロケールルーティング、DDG 優先 / Bing フォールバック):

uv run pytest -m "not live"

ライブスモークテスト:

RUN_LIVE_BING=1 uv run pytest -m live -s

マーカー名は依然として live / RUN_LIVE_BING です。ライブ実行はまず DDG にアクセスし、DDG が利用できない場合のみ Bing を使用します。

CI は Python 3.10、3.12、3.13、3.14 をカバーしています。

設計と安全上の注意

これは非公式の DuckDuckGo HTML + Bing HTML アダプターです。廃止された Bing Search API は使用していません。

  • DDG のマークアップは src/dsh_bing_search/providers/ddg.py にあります。

  • Bing のマークアップは src/dsh_bing_search/providers/bing_parser.py にあります。

  • 品質スコアリングは src/dsh_bing_search/quality.py にあり、エンジンに依存しません。

  • リクエストはブラウザの偽装を伴う curl_cffi.AsyncSession を使用します。

  • ユーザー指定のページ URL は公開 HTTP(S) ターゲットに制限され、安全なリダイレクト処理が有効です。

  • レスポンスボディはサイズ制限があります。

  • CAPTCHA / チャレンジ / HTTP 202 ページは status="blocked" として報告され、プラグインはそれらをバイパスしようとはしません。

  • www.bing.com のヘッドレス Bing は、構造的には有効だが無関係なカードを返すことがよくあります。cn.bing.com は一部の人気のある中国語クエリに役立ちますが、ロングテールの名前やタイトルは依然として最初のトークンに潰れることがあります。それが品質マークの目的です。

  • open は遅いターゲットサイトを自動的に再試行しません。必要に応じてタイムアウトの環境変数を増やしてください。

コミュニティ

DeepSeek Harness は現在開発者プレビュー中です。そのため、プラグインインターフェースはまだ進化する可能性があります。DSH 固有のサポートと発見については:

  • dsh-plugin トピックを参照してください。

  • DeepSeek Harness リポジトリ を参照してください。

  • 公式リポジトリからリンクされている DSH コミュニティチャンネルに参加してください。

貢献とパーサーの修正を歓迎します。

ライセンス

MIT

Available Tools

4 tools
findFind in Web PageA

Find a literal phrase in a page and return compact context windows around matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
patternYes
max_matchesNo
context_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYes
errorNo
statusYes
matchesNo
patternYes
source_idNo
total_matchesNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does reveal key behavior: matching is literal rather than regex or semantic, and the response consists of compact context windows around matches. However, it does not mention case sensitivity, failure modes, page loading behavior, or limits, leaving notable gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence contains the core action, the matching mode, and the response shape with no redundant words. It is front-loaded and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple tool, and the output schema likely covers return values. But with no annotations and no parameter documentation, it lacks details about max_matches behavior, exact context window semantics, and when to prefer sibling tools. It is minimally sufficient but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that 'pattern' is a literal phrase and 'context_chars' relates to compact context windows, but it does not explain 'max_matches', 'url', defaults, or the exact relationship between parameters and output. This is only partial compensation for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: finding a literal phrase in a page and returning compact context windows around matches. The word 'literal' helps distinguish it from the sibling 'search' tool, which implies broader or semantic search. This is a clear, specific purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this tool when an exact literal phrase is needed within a page. However, it does not explicitly say when not to use it or mention alternatives like 'search' or 'search_images'. The usage guidance is present only by implication.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openOpen Web PageA

Fetch a public HTTP(S) page with curl_cffi and return cleaned readable text.

Use after search when result snippets are insufficient. Private/local addresses are rejected, redirect targets use curl_cffi safe-follow mode, and response bytes are capped.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
textNo
errorNo
titleNo
statusYes
final_urlNo
source_idNo
truncatedNo
elapsed_msNo
content_typeNo
fetched_bytesNo
requested_urlYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and discloses several useful traits: public-only access, rejection of private/local addresses, safe-follow redirect mode, and a response byte cap. It could also mention error behavior or timeout handling, but the provided constraints are substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences front-load the purpose, then add usage context and behavioral constraints. No filler, every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the description covers URL type, output format, redirect behavior, and a cap. The main gap is max_chars semantics, which matters because there is no schema-level documentation and no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds that URL must be public HTTP(S), but it never explains the max_chars parameter or how the response cap relates to it. An agent cannot confidently tune max_chars based on this text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch'), resource ('public HTTP(S) page'), and output ('cleaned readable text'). This distinguishes it from siblings like search and search_images: it retrieves page content rather than result snippets or images.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use after search when result snippets are insufficient,' giving a clear trigger condition and relationship to the primary sibling. It also states a when-not: private/local addresses are rejected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_imagesSearch ImagesA

Search image indexes and rank results with pure text so vision is not required.

auto (default) tries Bing Images first and falls back to Wikimedia Commons when the top text score is below ~40, so one call yields a ranked set. bing_images parses Bing Images metadata (original URL / thumbnail / source page / title). commons queries Wikimedia Commons, a curated and licence-clear platform. Every result carries a 0-100 text score, a domain hint and explainable signals; pick the highest score, treat scores below ~40 as unverified, and optionally verify with find/open on the source page before downloading.

Args: query: What the image should depict. Compact concrete nouns plus the qualifier that uniquely identifies the subject (e.g. "复旦光华楼", "台州城墙"). "复旦光华楼" is better than "光华楼". Do not write whole sentences. If a compact query is still ambiguous or hits the wrong entity, write more (place, institution, year, type). count: Number of ranked image results to return, from 1 to 20. market: Locale such as en-US or zh-CN (Bing Images; Commons is language-neutral). provider: auto (default), bing_images, or commons.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNo
queryYes
marketNoen-US
providerNoauto

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
queryYes
marketNo
statusYes
resultsNo
providerNo
warningsNo
elapsed_msNo
returned_countNo
requested_countNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does so thoroughly. It discloses the ranking mechanism, the auto fallback threshold, what each provider does, and the exact result signals: 0-100 text score, domain hint, and explainable signals. It even tells the agent how to assess confidence and when verification is needed, which goes well beyond a minimal tool description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and behavior, and the Args section is logically organized. It is longer than typical descriptions, but that length is justified by the zero-coverage schema and the need to explain provider behavior and scoring. Minor redundancy exists because provider defaults and enum values are repeated from the schema, but the added context still earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's provider-switching complexity, fallback threshold, scoring semantics, and four parameters, the description provides everything needed to select and invoke it correctly. It explains query formulation, ranking confidence, provider differences, and optional verification workflow. The output schema covers return structure, so the description does not need to detail the exact JSON response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully compensate for the schema's lack of parameter documentation. It does: `query` has concrete examples and wording advice ('复旦光华楼' is better than '光华楼'), `count` is bounded 1-20, `market` is explained as locale-specific to Bing while Commons is language-neutral, and `provider` enumerates the options. This is excellent parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Search image indexes and rank results with pure text so vision is not required.' This clearly distinguishes the tool from the sibling `search`, `open`, and `find` by emphasizing image indexes and text-based ranking. The provider variants (bing_images, commons) further specify exactly what kind of image search this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives usable routing guidance: `auto` is the default, it falls back to Commons below ~40 text score, and results below ~40 should be treated as unverified. It also recommends verifying with `find`/`open` before downloading, which indirectly differentiates this search tool from sibling file/URL tools. It lacks an explicit 'when not to use this tool' statement, but the behavioral and provider guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedfind
    • First observedopen
    • First observedsearch
    • First observedsearch_images

TDQS

A4.3/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a clearly distinct action: web search, image search, page retrieval, and in-page phrase matching. Search and search_images are separated by media type, while open and find both operate on pages but serve complementary pre- and post-retrieval needs, so an agent can select without confusion.

Naming Consistency5/5

All tool names are short imperative verbs in snake_case: search, search_images, open, find. The only compound name, search_images, naturally follows a verb_noun pattern, and the overall naming is predictable and consistent.

Tool Count5/5

Four tools form a tightly scoped search-and-browse toolset. Each tool earns its place: web search, image search, full-page reading, and targeted phrase lookup. The count is neither thin nor bloated for the server's stated purpose.

Completeness5/5

The server covers the full core workflow: discovering content via web or image search, opening pages when snippets are insufficient, and locating specific phrases within pages. Pagination, locale, safesearch, and provider fallback options also cover important search variations, leaving no obvious dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for web search and content extraction using DuckDuckGo or SearXNG, with Playwright-based fetching and LLM-powered data extraction.
    140
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.
    MIT