dsh-bing-search
dsh-bing-search
**DeepSeek Harness (DSH)**용 웹 검색으로, curl_cffi로 구동되는 소형 MCP 서버로 구현되었습니다.
search 순서:
DuckDuckGo HTML(
html.duckduckgo.com)을 프로브하고 약 60초 동안 연결 가능성을 캐시합니다. 중국 본토에서는 프록시가 구성되지 않으면 이 프로브가 자주 실패합니다.연결 가능하면 DDG를 사용합니다.
DDG가 다운되었거나, 속도 제한(HTTP 202 / 챌린지)을 받거나, 결과 집합이
quality_label=poor인 경우 Bing으로 폴백합니다.Bing을 언어별로 라우팅합니다: 중국어 /
zh-*시장은cn.bing.com으로, 그 외에는www.bing.com으로 이동합니다.
모든 검색 응답에는 quality_score(0–1)와 quality_label(good / weak / poor)이 포함됩니다. poor는 사용할 수 없는 것으로 취급하세요(사전 페이지, 첫 토큰 쓰레기). 해당 제목은 인용하지 마세요.
DSH 에이전트에 세 가지 브라우저 스타일 도구를 제공합니다:
mcp__web__search— 공개 웹을 검색하고 정규화된 자연 결과를 반환합니다.mcp__web__open— 공개 웹 페이지를 열고 읽을 수 있는 텍스트를 추출합니다.mcp__web__find— 긴 페이지 내에서 텍스트를 찾고 주변 컨텍스트를 반환합니다.
DSH agent
-> @deepseek-ai/dsh-mcp-client
-> dsh-bing-search (MCP/stdio)
-> curl_cffi.AsyncSession(impersonate="chrome")
-> html.duckduckgo.com (if reachable)
-> else cn.bing.com / www.bing.com중국 본토: 프록시나 VPN이 없으면 DuckDuckGo에 연결할 수 없는 경우가 많습니다. 이는 예상된 동작입니다. 플러그인은 이때 Bing을 사용하고 warnings를 duckduckgo_unreachable로 설정합니다. MCP 자식 프로세스는 셸의 HTTP_PROXY / HTTPS_PROXY를 상속하지 않습니다(trust_env=False). 프록시를 강제하려면 플러그인 프로세스에 DSH_WEB_PROXY를 설정하세요(예: cordis env: 맵에서 http://127.0.0.1:10808). 일반적인 중국 본토 가정용 또는 캠퍼스 네트워크에서 DDG가 작동할 것이라고 가정하지 마세요.
커뮤니티 플러그인: DeepSeek Harness는 서드파티 플러그인이
dsh-pluginGitHub 토픽을 사용하여 발견될 수 있도록 요청합니다.
가장 빠른 설치: 이 저장소를 에이전트에 넘기세요
코딩 에이전트에 터미널과 파일 시스템 접근 권한이 있다면(Codex, Claude Code, Pi, OpenCode 등), 다음을 붙여넣으세요:
Install this DeepSeek Harness plugin into my current DSH setup:
https://github.com/Biogod2020/dsh-bing-search
Read the repository README and INSTALL.md first. Install it with uv, detect my active
DSH profile, add it through cordis.patch.yml using the required `insert` patch form,
preserve all unrelated config, use the absolute path of the installed dsh-bing-search
executable, then verify that mcp__web__search, mcp__web__open, and mcp__web__find are
registered. Finally run one real web search smoke test and report what changed.이것이 권장 경로입니다. INSTALL.md에는 에이전트를 위해 작성된 결정적 설치 계약이 포함되어 있습니다.
Related MCP server: webmcp
수동 설치
1. 실행 파일 설치
Python 3.10+가 필요합니다. uv 사용 시:
uv tool install --force git+https://github.com/Biogod2020/dsh-bing-search.git도구 bin 디렉터리를 찾으세요:
uv tool dir --bin아래 DSH 구성에서 dsh-bing-search(또는 Windows에서는 dsh-bing-search.exe)의 절대 경로를 사용하세요.
도구 설치 대신 개발용으로는:
git clone https://github.com/Biogod2020/dsh-bing-search.git
cd dsh-bing-search
uv sync --extra dev저장소에는 재현 가능한 개발 설치를 위한 uv.lock이 포함되어 있습니다.
2. DSH에 추가
DSH 프로필은 루트 cordis.yml과 패치 레이어 cordis.patch.yml을 결합합니다. 패치 레이어를 통해 새 플러그인을 추가할 때는 항목을 반드시 insert로 감싸야 합니다:
- insert:
- id: mcp-web
name: '@deepseek-ai/dsh-mcp-client'
config:
serverName: web
transport: stdio
command: /ABSOLUTE/PATH/TO/dsh-bing-search
args: []
toolCallTimeoutMs: 30000
failOnStartupError: true
reconnect:
enabled: true
initialDelayMs: 500
maxDelayMs: 30000
maxAttempts: 10cordis.patch.yml에 - id: mcp-web 항목을 단독으로 추가하지 마세요: 단독 항목은 기존 ID를 패치하며, 알 수 없는 ID는 건너뛸 수 있습니다. 루트 cordis.yml을 직접 편집하는 경우 일반적인 단독 플러그인 항목이 올바릅니다. cordis.example.yml을 참조하세요.
3. 확인
DSH가 프로필을 다시 로드한 후 모델은 다음을 볼 수 있어야 합니다:
mcp__web__search
mcp__web__open
mcp__web__find그런 다음 에이전트에게 최신 주제를 검색하고 결과 하나를 열도록 요청하세요. 성공적인 왕복은 검색 접근과 MCP 등록을 모두 검증합니다. 플러그인 코드를 변경한 후에는 DSH(또는 MCP 자식 프로세스)를 다시 시작하세요. stdio 프로세스는 Python을 핫 리로드하지 않습니다.
도구
search
{
"query": "DeepSeek Harness GitHub",
"count": 8,
"offset": 0,
"market": "en-US",
"safe_search": "Moderate"
}반환:
필드 | 의미 |
|
|
| 자연 검색 결과 |
| 정식 URL의 안정적인 ID |
| 쿼리와 제목/스니펫의 0–1 중첩 정도 |
|
|
| 폴백 사유 및 품질 메모 |
중국어 쿼리에는 market=zh-CN을 사용하세요. 쿼리에 CJK가 포함된 경우 market이 en-US여도 Bing 폴백은 여전히 cn.bing.com을 사용합니다.
DuckDuckGo /l/?uddg= 및 Bing /ck/a 리디렉션은 가능한 경우 디코딩됩니다. 일반적인 추적 매개변수는 제거되고 중복 URL은 병합됩니다.
사람, 논문, 일러스트 블로그의 경우 먼저 저자 이름이나 짧은 고유 명사를 검색하세요. quality_label이 poor이면 쿼리를 계속 늘리지 마세요. 중국 학술 메타데이터는 일반 웹 검색이 아닌 특수 코퍼스(예: CNKI)에 속합니다.
open
{
"url": "https://example.com/article",
"max_chars": 24000
}curl_cffi로 공개 HTTP(S) 페이지를 가져오고 DNS/IP 검사와 안전한 리디렉션을 적용하며 응답 크기를 제한하고 자바스크립트를 실행하지 않고 읽을 수 있는 텍스트를 추출합니다.
open은 기사형 HTML용으로 제작되었습니다. 브라우저가 아닙니다. 실제 DSH 실행에서 날씨 및 기타 위젯이 많은 사이트(tianqi.com, weather.com.cn 등)는 탐색 크롬이나 거의 빈 텍스트를 생성하는 경우가 많았습니다: Trafilatura가 본문 기사를 찾지 못하면 폴백이 전체 DOM을 덤프합니다. status는 여전히 ok일 수 있습니다. 이러한 페이지에서는 검색 snippet을 신뢰하거나 더 간단한 기사 URL을 open하세요. 실시간 온도, 지도 또는 기타 JS 렌더링 UI를 기대하지 마세요.
find
{
"url": "https://example.com/article",
"pattern": "DeepSeek",
"max_matches": 5,
"context_chars": 700
}전체 페이지를 모델 컨텍스트에 주입하지 않고 일치하는 영역을 반환합니다.
하나의 거대한 search_and_summarize 도구 대신 세 개의 도구가 필요한 이유는 무엇인가요?
플러그인은 검색을 결정적으로 유지하고 DSH 모델이 리서치 루프를 제어하도록 합니다:
search -> inspect candidates -> open -> find / search again -> synthesize플러그인은 HTTP, 파싱, 정리, 캐싱, 엔진 폴백, 출처 추적, 품질 표시를 처리합니다. 에이전트는 무엇을 검색할지, 어떤 출처를 신뢰할지, 언제 쿼리를 다시 구성할지, 언제 충분한 증거가 수집되었는지를 결정합니다. 에이전트는 quality_label과 warnings를 읽어야 합니다.
구성
환경 변수 | 기본값 | 목적 |
|
| 기본값이 아닌 값으로 설정된 경우에만 Bing HTML 엔드포인트를 재정의합니다(테스트). 그 외에는 호스트가 언어별로 선택됩니다 |
|
|
|
| empty | HTTP/HTTPS/SOCKS 프록시. 프로세스는 |
|
| 전송 타임아웃 |
|
| 연결 타임아웃 |
|
|
|
|
| 검색 페이지 최대 본문 크기 |
|
| 최대 리디렉션 수 |
|
| 프로세스 내 최대 동시 요청 수 |
|
| 검색 캐시 TTL |
|
| 페이지 캐시 TTL |
테스트
오프라인 테스트(파서, 품질 점수, 로케일 라우팅, DDG 우선 / Bing 폴백):
uv run pytest -m "not live"라이브 스모크 테스트:
RUN_LIVE_BING=1 uv run pytest -m live -s마커 이름은 여전히 live / RUN_LIVE_BING입니다. 라이브 실행은 DDG를 먼저 시도하고 DDG를 사용할 수 없는 경우에만 Bing을 사용합니다.
CI는 Python 3.10, 3.12, 3.13 및 3.14를 포함합니다.
설계 및 안전 참고 사항
이것은 비공식 DuckDuckGo HTML + Bing HTML 어댑터입니다. 폐지된 Bing Search API를 사용하지 않습니다.
DDG 마크업은
src/dsh_bing_search/providers/ddg.py에 있습니다.Bing 마크업은
src/dsh_bing_search/providers/bing_parser.py에 있습니다.품질 점수는
src/dsh_bing_search/quality.py에 있으며 엔진에 독립적입니다.요청은 브라우저 임의화(impersonation)와 함께
curl_cffi.AsyncSession을 사용합니다.사용자가 제공한 페이지 URL은 공개 HTTP(S) 대상으로 제한되며 안전한 리디렉션 처리가 활성화됩니다.
응답 본문은 크기가 제한됩니다.
CAPTCHA / 챌린지 / HTTP 202 페이지는
status="blocked"로 보고됩니다. 플러그인은 이를 우회하려 시도하지 않습니다.www.bing.com의 헤드리스 Bing은 구조적으로 유효하지만 관련 없는 카드를 자주 반환합니다.cn.bing.com은 일부 인기 중국어 쿼리에 도움이 되지만, 롱테일 이름과 제목은 여전히 첫 번째 토큰으로 축소될 수 있습니다. 이것이 품질 표시의 목적입니다.open은 느린 대상 사이트를 자동으로 재시도하지 않습니다. 필요한 경우 타임아웃 환경 변수를 늘리세요.
커뮤니티
DeepSeek Harness는 현재 개발자 미리보기 상태이므로 플러그인 인터페이스는 계속 변경될 수 있습니다. DSH 관련 지원 및 검색을 원한다면:
dsh-plugin토픽을 둘러보세요.DeepSeek Harness 저장소를 확인하세요.
공식 저장소에서 연결된 DSH 커뮤니티 채널에 참여하세요.
기여와 파서 수정을 환영합니다.
라이선스
MIT
Available Tools
4 toolsfindFind in Web PageA
Find a literal phrase in a page and return compact context windows around matches.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| pattern | Yes | ||
| max_matches | No | ||
| context_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| error | No | |
| status | Yes | |
| matches | No | |
| pattern | Yes | |
| source_id | No | |
| total_matches | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does reveal key behavior: matching is literal rather than regex or semantic, and the response consists of compact context windows around matches. However, it does not mention case sensitivity, failure modes, page loading behavior, or limits, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence contains the core action, the matching mode, and the response shape with no redundant words. It is front-loaded and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple tool, and the output schema likely covers return values. But with no annotations and no parameter documentation, it lacks details about max_matches behavior, exact context window semantics, and when to prefer sibling tools. It is minimally sufficient but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that 'pattern' is a literal phrase and 'context_chars' relates to compact context windows, but it does not explain 'max_matches', 'url', defaults, or the exact relationship between parameters and output. This is only partial compensation for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: finding a literal phrase in a page and returning compact context windows around matches. The word 'literal' helps distinguish it from the sibling 'search' tool, which implies broader or semantic search. This is a clear, specific purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this tool when an exact literal phrase is needed within a page. However, it does not explicitly say when not to use it or mention alternatives like 'search' or 'search_images'. The usage guidance is present only by implication.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
openOpen Web PageA
Fetch a public HTTP(S) page with curl_cffi and return cleaned readable text.
Use after search when result snippets are insufficient. Private/local addresses are rejected, redirect targets use curl_cffi safe-follow mode, and response bytes are capped.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | |
| error | No | |
| title | No | |
| status | Yes | |
| final_url | No | |
| source_id | No | |
| truncated | No | |
| elapsed_ms | No | |
| content_type | No | |
| fetched_bytes | No | |
| requested_url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and discloses several useful traits: public-only access, rejection of private/local addresses, safe-follow redirect mode, and a response byte cap. It could also mention error behavior or timeout handling, but the provided constraints are substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences front-load the purpose, then add usage context and behavioral constraints. No filler, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers URL type, output format, redirect behavior, and a cap. The main gap is max_chars semantics, which matters because there is no schema-level documentation and no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds that URL must be public HTTP(S), but it never explains the max_chars parameter or how the response cap relates to it. An agent cannot confidently tune max_chars based on this text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch'), resource ('public HTTP(S) page'), and output ('cleaned readable text'). This distinguishes it from siblings like search and search_images: it retrieves page content rather than result snippets or images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use after search when result snippets are insufficient,' giving a clear trigger condition and relationship to the primary sibling. It also states a when-not: private/local addresses are rejected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchSearch the WebA
Search the public web. DuckDuckGo is tried first when reachable; Bing is the fallback.
Chinese queries / zh-* markets use cn.bing.com. Read quality_label: poor means the titles are unrelated or first-token junk — do not treat them as answers.
Args: query: Compact concrete nouns plus the qualifier that uniquely identifies the subject. "复旦光华楼" is better than "光华楼" — the extra place/institution is necessary, not padding. Do not write whole sentences. If a compact query is still ambiguous or hits the wrong entity, write more (place, institution, year, type). For a person plus a paper, search the author name first. count: Number of organic results to return, from 1 to 20. offset: Result offset for pagination, from 0 to 100. market: Locale such as en-US or zh-CN. Chinese text should use zh-CN. safe_search: SafeSearch level (used when Bing is the engine).
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| query | Yes | ||
| market | No | en-US | |
| offset | No | ||
| safe_search | No | Moderate |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| query | Yes | |
| market | No | |
| offset | No | |
| status | Yes | |
| results | No | |
| provider | No | |
| warnings | No | |
| elapsed_ms | No | |
| safe_search | No | |
| quality_label | No | good / weak / poor. If poor, do not treat results as answers. |
| quality_score | No | 0-1 overlap of the query with titles/snippets. Below 0.3 is not trustworthy. |
| returned_count | No | |
| requested_count | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and handles it well: it discloses the DuckDuckGo/Bing fallback order, the cn.bing.com behavior for Chinese markets, and the meaning of quality_label=poor. This gives agents useful execution expectations beyond what the schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but the length is justified by the need to explain query construction and engine quirks. The opening behavior is front-loaded, and the Args section is clearly organized. A small amount of redundancy exists, but every major sentence adds practical value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations are absent, the description covers the key operational context: engine fallback, locale behavior, quality_label handling, and parameter semantics. It lacks explicit when-to-use versus search_images/open/find guidance, but the other information is sufficient for an agent to call and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does thoroughly. Each parameter is explained: query receives detailed formulation rules, count is bounded to 1–20, offset to 0–100, market is tied to locale, and safe_search enum values are named. This is far more helpful than the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence clearly states the action and scope: 'Search the public web.' The description goes beyond the title by specifying engine behavior (DuckDuckGo first, Bing fallback) and the Chinese-market variant, which distinguishes this tool from image or document navigation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong guidance on how to construct queries, including concrete examples and disambiguation advice (e.g., '复旦光华楼' is better than '光华楼'). It also warns when not to trust results via quality_label. It does not explicitly name alternative tools like search_images, so some sibling differentiation is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_imagesSearch ImagesA
Search image indexes and rank results with pure text so vision is not required.
auto (default) tries Bing Images first and falls back to Wikimedia Commons
when the top text score is below ~40, so one call yields a ranked set.
bing_images parses Bing Images metadata (original URL / thumbnail / source
page / title). commons queries Wikimedia Commons, a curated and
licence-clear platform. Every result carries a 0-100 text score, a domain
hint and explainable signals; pick the highest score, treat scores below
~40 as unverified, and optionally verify with find/open on the source
page before downloading.
Args: query: What the image should depict. Compact concrete nouns plus the qualifier that uniquely identifies the subject (e.g. "复旦光华楼", "台州城墙"). "复旦光华楼" is better than "光华楼". Do not write whole sentences. If a compact query is still ambiguous or hits the wrong entity, write more (place, institution, year, type). count: Number of ranked image results to return, from 1 to 20. market: Locale such as en-US or zh-CN (Bing Images; Commons is language-neutral). provider: auto (default), bing_images, or commons.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| query | Yes | ||
| market | No | en-US | |
| provider | No | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| query | Yes | |
| market | No | |
| status | Yes | |
| results | No | |
| provider | No | |
| warnings | No | |
| elapsed_ms | No | |
| returned_count | No | |
| requested_count | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so thoroughly. It discloses the ranking mechanism, the auto fallback threshold, what each provider does, and the exact result signals: 0-100 text score, domain hint, and explainable signals. It even tells the agent how to assess confidence and when verification is needed, which goes well beyond a minimal tool description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and behavior, and the Args section is logically organized. It is longer than typical descriptions, but that length is justified by the zero-coverage schema and the need to explain provider behavior and scoring. Minor redundancy exists because provider defaults and enum values are repeated from the schema, but the added context still earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's provider-switching complexity, fallback threshold, scoring semantics, and four parameters, the description provides everything needed to select and invoke it correctly. It explains query formulation, ranking confidence, provider differences, and optional verification workflow. The output schema covers return structure, so the description does not need to detail the exact JSON response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate for the schema's lack of parameter documentation. It does: `query` has concrete examples and wording advice ('复旦光华楼' is better than '光华楼'), `count` is bounded 1-20, `market` is explained as locale-specific to Bing while Commons is language-neutral, and `provider` enumerates the options. This is excellent parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search image indexes and rank results with pure text so vision is not required.' This clearly distinguishes the tool from the sibling `search`, `open`, and `find` by emphasizing image indexes and text-based ranking. The provider variants (bing_images, commons) further specify exactly what kind of image search this is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives usable routing guidance: `auto` is the default, it falls back to Commons below ~40 text score, and results below ~40 should be treated as unverified. It also recommends verifying with `find`/`open` before downloading, which indirectly differentiates this search tool from sibling file/URL tools. It lacks an explicit 'when not to use this tool' statement, but the behavioral and provider guidance is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
find - First observed
open - First observed
search - First observed
search_images
TDQS
Scored across 4 tools
Each tool targets a clearly distinct action: web search, image search, page retrieval, and in-page phrase matching. Search and search_images are separated by media type, while open and find both operate on pages but serve complementary pre- and post-retrieval needs, so an agent can select without confusion.
All tool names are short imperative verbs in snake_case: search, search_images, open, find. The only compound name, search_images, naturally follows a verb_noun pattern, and the overall naming is predictable and consistent.
Four tools form a tightly scoped search-and-browse toolset. Each tool earns its place: web search, image search, full-page reading, and targeted phrase lookup. The count is neither thin nor bloated for the server's stated purpose.
The server covers the full core workflow: discovering content via web or image search, opening pages when snippets are insufficient, and locating specific phrases within pages. Pagination, locale, safesearch, and provider fallback options also cover important search variations, leaving no obvious dead ends.
Maintenance
Related MCP Connectors
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Free web search for AI agents. No API key required. Hosted MCP in active development.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceThis MCP server provides tools for AI agents to search the web, fetch page content, and query specific elements from pages using DuckDuckGo.13 npm-
- AlicenseNot gradedqualityCmaintenanceMCP server for web search and content extraction using DuckDuckGo or SearXNG, with Playwright-based fetching and LLM-powered data extraction.140MIT
- AlicenseAqualityDmaintenanceA general MCP server providing web search capabilities using DeepSeek's native online search API.167 npm31MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.MIT