dsh-bing-search
dsh-bing-search
Websuche für DeepSeek Harness (DSH), implementiert als kleiner MCP-Server und unterstützt von curl_cffi.
search-Reihenfolge:
DuckDuckGo-HTML (
html.duckduckgo.com) prüfen und die Erreichbarkeit für etwa 60 Sekunden zwischenspeichern. Vom chinesischen Festland aus schlägt diese Prüfung oft fehl, sofern kein Proxy konfiguriert ist.DDG verwenden, wenn es erreichbar ist.
Auf Bing zurückfallen, wenn DDG nicht verfügbar ist, rate-limitiert (HTTP 202 / Challenge) ist oder das Ergebnisset
quality_label=pooraufweist.Bing je nach Sprache weiterleiten: Chinesische /
zh-*-Märkte gehen ancn.bing.com, andernfalls anwww.bing.com.
Jede Suchantwort enthält quality_score (0–1) und quality_label (good / weak / poor). Behandle poor als unbrauchbar (Wörterbuchseiten, Erst-Token-Müll). Zitiere diese Titel nicht.
Es gibt einem DSH-Agenten drei browserartige Werkzeuge:
mcp__web__search— durchsuche das öffentliche Web und gib normalisierte organische Ergebnisse zurück.mcp__web__open— öffne eine öffentliche Webseite und extrahiere lesbaren Text.mcp__web__find— finde Text in einer langen Seite und gib umliegenden Kontext zurück.
DSH agent
-> @deepseek-ai/dsh-mcp-client
-> dsh-bing-search (MCP/stdio)
-> curl_cffi.AsyncSession(impersonate="chrome")
-> html.duckduckgo.com (if reachable)
-> else cn.bing.com / www.bing.comFestlandchina: DuckDuckGo ist ohne Proxy oder VPN oft nicht erreichbar. Das ist zu erwarten. Das Plugin verwendet dann Bing und setzt warnings auf duckduckgo_unreachable. Der MCP-Kindprozess erbt nicht die HTTP_PROXY/HTTPS_PROXY-Einstellungen Ihrer Shell (trust_env=False). Um einen Proxy zu erzwingen, setzen Sie DSH_WEB_PROXY im Plugin-Prozess (zum Beispiel http://127.0.0.1:10808 in der env:-Map von cordis). Gehen Sie nicht davon aus, dass DDG in einem typischen Festland-Heim- oder Campusnetzwerk funktioniert.
Community-Plugin: DeepSeek Harness bittet Drittanbieter-Plugins, das GitHub-Thema
dsh-pluginzur Auffindbarkeit zu verwenden.
Schnellste Installation: Dieses Repository einem Agenten übergeben
Wenn Ihr Coding-Agent Terminal- und Dateisystemzugriff hat (Codex, Claude Code, Pi, OpenCode usw.), fügen Sie Folgendes ein:
Install this DeepSeek Harness plugin into my current DSH setup:
https://github.com/Biogod2020/dsh-bing-search
Read the repository README and INSTALL.md first. Install it with uv, detect my active
DSH profile, add it through cordis.patch.yml using the required `insert` patch form,
preserve all unrelated config, use the absolute path of the installed dsh-bing-search
executable, then verify that mcp__web__search, mcp__web__open, and mcp__web__find are
registered. Finally run one real web search smoke test and report what changed.Das ist der empfohlene Weg. INSTALL.md enthält einen deterministischen Installationsvertrag, der für Agenten geschrieben wurde.
Related MCP server: webmcp
Manuelle Installation
1. Die ausführbare Datei installieren
Python 3.10+ ist erforderlich. Mit uv:
uv tool install --force git+https://github.com/Biogod2020/dsh-bing-search.gitFinden Sie das Tool-Bin-Verzeichnis:
uv tool dir --binVerwenden Sie den absoluten Pfad zu dsh-bing-search (oder dsh-bing-search.exe unter Windows) in der DSH-Konfiguration unten.
Für die Entwicklung anstelle einer Tool-Installation:
git clone https://github.com/Biogod2020/dsh-bing-search.git
cd dsh-bing-search
uv sync --extra devDas Repository enthält uv.lock für reproduzierbare Entwicklungsinstallationen.
2. Zu DSH hinzufügen
DSH-Profile kombinieren eine Stamm-cordis.yml mit einer Patch-Schicht cordis.patch.yml. Wenn Sie ein neues Plugin über die Patch-Schicht hinzufügen, muss der Eintrag in insert eingebettet werden:
- insert:
- id: mcp-web
name: '@deepseek-ai/dsh-mcp-client'
config:
serverName: web
transport: stdio
command: /ABSOLUTE/PATH/TO/dsh-bing-search
args: []
toolCallTimeoutMs: 30000
failOnStartupError: true
reconnect:
enabled: true
initialDelayMs: 500
maxDelayMs: 30000
maxAttempts: 10Fügen Sie keinen nackten Eintrag - id: mcp-web zu cordis.patch.yml hinzu: Nackte Einträge patchen vorhandene IDs, und eine unbekannte ID kann übersprungen werden. Wenn Sie die Stamm-cordis.yml direkt bearbeiten, ist ein normaler nackter Plugin-Eintrag korrekt. Siehe cordis.example.yml.
3. Überprüfen
Nachdem DSH das Profil neu geladen hat, sollte das Modell Folgendes sehen:
mcp__web__search
mcp__web__open
mcp__web__findBitten Sie den Agenten dann, nach etwas Aktuellem zu suchen und ein Ergebnis zu öffnen. Ein erfolgreicher Roundtrip bestätigt sowohl den Suchzugriff als auch die MCP-Registrierung. Starten Sie DSH (oder den MCP-Kindprozess) neu, nachdem Sie Plugin-Code geändert haben; der Stdio-Prozess lädt Python nicht neu.
Werkzeuge
search
{
"query": "DeepSeek Harness GitHub",
"count": 8,
"offset": 0,
"market": "en-US",
"safe_search": "Moderate"
}Gibt zurück:
Feld | Bedeutung |
|
|
| Organisches Ergebnis |
| Stabile ID aus der kanonischen URL |
| 0–1 Überlappung der Abfrage mit Titeln/Textauszügen |
|
|
| Fallback-Grund und Qualitätshinweise |
Verwenden Sie market=zh-CN für chinesische Abfragen. Wenn die Abfrage CJK enthält, verwendet der Bing-Fallback auch bei market=en-US weiterhin cn.bing.com.
DuckDuckGo-/l/?uddg=- und Bing-/ck/a-Weiterleitungen werden nach Möglichkeit dekodiert. Häufige Tracking-Parameter werden entfernt und doppelte URLs werden zusammengeführt.
Suchen Sie bei Personen, Papieren oder illustrierten Blogs zuerst nach dem Autorennamen oder einem kurzen Eigennamen. Wenn quality_label poor ist, verlängern Sie die Abfrage nicht weiter. Chinesische akademische Metadaten gehören in ein spezialisiertes Korpus (z. B. CNKI), nicht in diese allgemeine Websuche.
open
{
"url": "https://example.com/article",
"max_chars": 24000
}Ruft öffentliche HTTP(S)-Seiten mit curl_cffi ab, wendet DNS/IP-Prüfungen und sichere Weiterleitungen an, begrenzt die Antwortgröße und extrahiert lesbaren Text ohne JavaScript auszuführen.
open ist für artikelähnliches HTML gebaut. Es ist kein Browser. Live-DSH-Läufe zeigten, dass Wetter- und andere Widget-lastige Websites (tianqi.com, weather.com.cn und ähnliche) oft Navigations-Chrome oder fast leeren Text liefern: Trafilatura findet keinen Hauptartikel, und der Fallback gibt das gesamte DOM aus. status kann trotzdem ok sein. Vertrauen Sie bei diesen Seiten dem Such-snippet oder öffnen Sie eine einfachere Artikel-URL. Erwarten Sie keine Live-Temperatur, Karten oder andere JS-gerenderte Benutzeroberfläche.
find
{
"url": "https://example.com/article",
"pattern": "DeepSeek",
"max_matches": 5,
"context_chars": 700
}Gibt passende Bereiche zurück, ohne die gesamte Seite in den Modellkontext einzufügen.
Warum drei Werkzeuge statt eines riesigen search_and_summarize-Werkzeugs?
Das Plugin hält die Abfrage deterministisch und lässt das DSH-Modell die Recherche-Schleife steuern:
search -> inspect candidates -> open -> find / search again -> synthesizeDas Plugin übernimmt HTTP, Parsing, Bereinigung, Caching, Engine-Fallback, Herkunft und eine Qualitätskennzeichnung. Der Agent entscheidet, wonach gesucht wird, welchen Quellen er vertraut, wann die Abfrage umformuliert wird und wann genügend Beweise gesammelt wurden. Der Agent muss quality_label und warnings lesen.
Konfiguration
Umgebungsvariable | Standard | Zweck |
|
| Bing-HTML-Endpunkt nur überschreiben, wenn ein Nicht-Standardwert gesetzt ist (Tests). Andernfalls wird der Host anhand der Sprache gewählt |
|
|
|
| leer | HTTP/HTTPS/SOCKS-Proxy. Der Prozess verwendet |
|
| Übertragungs-Timeout |
|
| Verbindungs-Timeout |
|
| Maximale Größe des Inhalts für |
|
| Maximale Größe des Suchseiteninhalts |
|
| Maximale Anzahl an Weiterleitungen |
|
| Maximale gleichzeitige Anforderungen pro Prozess |
|
| Suchcache-TTL |
|
| Seiten-Cache-TTL |
Tests
Offline-Tests (Parser, Qualitätswert, Locale-Routing, DDG-zuerst / Bing-Fallback):
uv run pytest -m "not live"Live-Smoke-Test:
RUN_LIVE_BING=1 uv run pytest -m live -sDer Markername ist weiterhin live / RUN_LIVE_BING. Ein Live-Lauf greift zuerst auf DDG zu und verwendet Bing nur, wenn DDG nicht verfügbar ist.
CI umfasst Python 3.10, 3.12, 3.13 und 3.14.
Design- und Sicherheitshinweise
Dies ist ein inoffizieller DuckDuckGo-HTML- und Bing-HTML-Adapter. Er verwendet die eingestellte Bing-Such-API nicht.
DDG-Markup befindet sich in
src/dsh_bing_search/providers/ddg.py.Bing-Markup befindet sich in
src/dsh_bing_search/providers/bing_parser.py.Die Qualitätsbewertung befindet sich in
src/dsh_bing_search/quality.pyund ist engine-agnostisch.Anforderungen verwenden
curl_cffi.AsyncSessionmit Browser-Imitation.Benutzer bereitgestellte Seiten-URLs sind auf öffentliche HTTP(S)-Ziele beschränkt, und eine sichere Weiterleitungsbehandlung ist aktiviert.
Antwortkörper sind größenbegrenzt.
CAPTCHA-/Challenge-/HTTP-202-Seiten werden als
status="blocked"gemeldet; das Plugin versucht nicht, sie zu umgehen.Headless Bing auf
www.bing.comliefert oft strukturierte, aber unpassende Karten.cn.bing.comhilft bei einigen heißen chinesischen Abfragen; Long-Tail-Namen und Titel können weiterhin auf das erste Token kollabieren. Dafür ist die Qualitätskennzeichnung da.openwiederholt langsame Zielseiten nicht automatisch; erhöhen Sie bei Bedarf die Timeout-Umgebungsvariablen.
Community
DeepSeek Harness befindet sich derzeit in der Entwicklervorschau, daher können sich Plugin-Schnittstellen noch ändern. Für DSH-spezifischen Support und Auffindbarkeit:
Durchsuchen Sie das Thema
dsh-plugin.Siehe das DeepSeek Harness Repository.
Treten Sie den DSH-Community-Kanälen bei, die im offiziellen Repository verlinkt sind.
Beiträge und Parser-Fixes sind willkommen.
Lizenz
MIT
Available Tools
4 toolsfindFind in Web PageA
Find a literal phrase in a page and return compact context windows around matches.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| pattern | Yes | ||
| max_matches | No | ||
| context_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| error | No | |
| status | Yes | |
| matches | No | |
| pattern | Yes | |
| source_id | No | |
| total_matches | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does reveal key behavior: matching is literal rather than regex or semantic, and the response consists of compact context windows around matches. However, it does not mention case sensitivity, failure modes, page loading behavior, or limits, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence contains the core action, the matching mode, and the response shape with no redundant words. It is front-loaded and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple tool, and the output schema likely covers return values. But with no annotations and no parameter documentation, it lacks details about max_matches behavior, exact context window semantics, and when to prefer sibling tools. It is minimally sufficient but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that 'pattern' is a literal phrase and 'context_chars' relates to compact context windows, but it does not explain 'max_matches', 'url', defaults, or the exact relationship between parameters and output. This is only partial compensation for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: finding a literal phrase in a page and returning compact context windows around matches. The word 'literal' helps distinguish it from the sibling 'search' tool, which implies broader or semantic search. This is a clear, specific purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this tool when an exact literal phrase is needed within a page. However, it does not explicitly say when not to use it or mention alternatives like 'search' or 'search_images'. The usage guidance is present only by implication.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
openOpen Web PageA
Fetch a public HTTP(S) page with curl_cffi and return cleaned readable text.
Use after search when result snippets are insufficient. Private/local addresses are rejected, redirect targets use curl_cffi safe-follow mode, and response bytes are capped.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| text | No | |
| error | No | |
| title | No | |
| status | Yes | |
| final_url | No | |
| source_id | No | |
| truncated | No | |
| elapsed_ms | No | |
| content_type | No | |
| fetched_bytes | No | |
| requested_url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and discloses several useful traits: public-only access, rejection of private/local addresses, safe-follow redirect mode, and a response byte cap. It could also mention error behavior or timeout handling, but the provided constraints are substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences front-load the purpose, then add usage context and behavioral constraints. No filler, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers URL type, output format, redirect behavior, and a cap. The main gap is max_chars semantics, which matters because there is no schema-level documentation and no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds that URL must be public HTTP(S), but it never explains the max_chars parameter or how the response cap relates to it. An agent cannot confidently tune max_chars based on this text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch'), resource ('public HTTP(S) page'), and output ('cleaned readable text'). This distinguishes it from siblings like search and search_images: it retrieves page content rather than result snippets or images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use after search when result snippets are insufficient,' giving a clear trigger condition and relationship to the primary sibling. It also states a when-not: private/local addresses are rejected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchSearch the WebA
Search the public web. DuckDuckGo is tried first when reachable; Bing is the fallback.
Chinese queries / zh-* markets use cn.bing.com. Read quality_label: poor means the titles are unrelated or first-token junk — do not treat them as answers.
Args: query: Compact concrete nouns plus the qualifier that uniquely identifies the subject. "复旦光华楼" is better than "光华楼" — the extra place/institution is necessary, not padding. Do not write whole sentences. If a compact query is still ambiguous or hits the wrong entity, write more (place, institution, year, type). For a person plus a paper, search the author name first. count: Number of organic results to return, from 1 to 20. offset: Result offset for pagination, from 0 to 100. market: Locale such as en-US or zh-CN. Chinese text should use zh-CN. safe_search: SafeSearch level (used when Bing is the engine).
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| query | Yes | ||
| market | No | en-US | |
| offset | No | ||
| safe_search | No | Moderate |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| query | Yes | |
| market | No | |
| offset | No | |
| status | Yes | |
| results | No | |
| provider | No | |
| warnings | No | |
| elapsed_ms | No | |
| safe_search | No | |
| quality_label | No | good / weak / poor. If poor, do not treat results as answers. |
| quality_score | No | 0-1 overlap of the query with titles/snippets. Below 0.3 is not trustworthy. |
| returned_count | No | |
| requested_count | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and handles it well: it discloses the DuckDuckGo/Bing fallback order, the cn.bing.com behavior for Chinese markets, and the meaning of quality_label=poor. This gives agents useful execution expectations beyond what the schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but the length is justified by the need to explain query construction and engine quirks. The opening behavior is front-loaded, and the Args section is clearly organized. A small amount of redundancy exists, but every major sentence adds practical value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations are absent, the description covers the key operational context: engine fallback, locale behavior, quality_label handling, and parameter semantics. It lacks explicit when-to-use versus search_images/open/find guidance, but the other information is sufficient for an agent to call and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does thoroughly. Each parameter is explained: query receives detailed formulation rules, count is bounded to 1–20, offset to 0–100, market is tied to locale, and safe_search enum values are named. This is far more helpful than the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence clearly states the action and scope: 'Search the public web.' The description goes beyond the title by specifying engine behavior (DuckDuckGo first, Bing fallback) and the Chinese-market variant, which distinguishes this tool from image or document navigation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong guidance on how to construct queries, including concrete examples and disambiguation advice (e.g., '复旦光华楼' is better than '光华楼'). It also warns when not to trust results via quality_label. It does not explicitly name alternative tools like search_images, so some sibling differentiation is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_imagesSearch ImagesA
Search image indexes and rank results with pure text so vision is not required.
auto (default) tries Bing Images first and falls back to Wikimedia Commons
when the top text score is below ~40, so one call yields a ranked set.
bing_images parses Bing Images metadata (original URL / thumbnail / source
page / title). commons queries Wikimedia Commons, a curated and
licence-clear platform. Every result carries a 0-100 text score, a domain
hint and explainable signals; pick the highest score, treat scores below
~40 as unverified, and optionally verify with find/open on the source
page before downloading.
Args: query: What the image should depict. Compact concrete nouns plus the qualifier that uniquely identifies the subject (e.g. "复旦光华楼", "台州城墙"). "复旦光华楼" is better than "光华楼". Do not write whole sentences. If a compact query is still ambiguous or hits the wrong entity, write more (place, institution, year, type). count: Number of ranked image results to return, from 1 to 20. market: Locale such as en-US or zh-CN (Bing Images; Commons is language-neutral). provider: auto (default), bing_images, or commons.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | ||
| query | Yes | ||
| market | No | en-US | |
| provider | No | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| query | Yes | |
| market | No | |
| status | Yes | |
| results | No | |
| provider | No | |
| warnings | No | |
| elapsed_ms | No | |
| returned_count | No | |
| requested_count | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so thoroughly. It discloses the ranking mechanism, the auto fallback threshold, what each provider does, and the exact result signals: 0-100 text score, domain hint, and explainable signals. It even tells the agent how to assess confidence and when verification is needed, which goes well beyond a minimal tool description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and behavior, and the Args section is logically organized. It is longer than typical descriptions, but that length is justified by the zero-coverage schema and the need to explain provider behavior and scoring. Minor redundancy exists because provider defaults and enum values are repeated from the schema, but the added context still earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's provider-switching complexity, fallback threshold, scoring semantics, and four parameters, the description provides everything needed to select and invoke it correctly. It explains query formulation, ranking confidence, provider differences, and optional verification workflow. The output schema covers return structure, so the description does not need to detail the exact JSON response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate for the schema's lack of parameter documentation. It does: `query` has concrete examples and wording advice ('复旦光华楼' is better than '光华楼'), `count` is bounded 1-20, `market` is explained as locale-specific to Bing while Commons is language-neutral, and `provider` enumerates the options. This is excellent parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search image indexes and rank results with pure text so vision is not required.' This clearly distinguishes the tool from the sibling `search`, `open`, and `find` by emphasizing image indexes and text-based ranking. The provider variants (bing_images, commons) further specify exactly what kind of image search this is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives usable routing guidance: `auto` is the default, it falls back to Commons below ~40 text score, and results below ~40 should be treated as unverified. It also recommends verifying with `find`/`open` before downloading, which indirectly differentiates this search tool from sibling file/URL tools. It lacks an explicit 'when not to use this tool' statement, but the behavioral and provider guidance is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
find - First observed
open - First observed
search - First observed
search_images
TDQS
Scored across 4 tools
Each tool targets a clearly distinct action: web search, image search, page retrieval, and in-page phrase matching. Search and search_images are separated by media type, while open and find both operate on pages but serve complementary pre- and post-retrieval needs, so an agent can select without confusion.
All tool names are short imperative verbs in snake_case: search, search_images, open, find. The only compound name, search_images, naturally follows a verb_noun pattern, and the overall naming is predictable and consistent.
Four tools form a tightly scoped search-and-browse toolset. Each tool earns its place: web search, image search, full-page reading, and targeted phrase lookup. The count is neither thin nor bloated for the server's stated purpose.
The server covers the full core workflow: discovering content via web or image search, opening pages when snippets are insufficient, and locating specific phrases within pages. Pagination, locale, safesearch, and provider fallback options also cover important search variations, leaving no obvious dead ends.
Maintenance
Related MCP Connectors
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Free web search for AI agents. No API key required. Hosted MCP in active development.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceThis MCP server provides tools for AI agents to search the web, fetch page content, and query specific elements from pages using DuckDuckGo.13 npm-
- AlicenseNot gradedqualityCmaintenanceMCP server for web search and content extraction using DuckDuckGo or SearXNG, with Playwright-based fetching and LLM-powered data extraction.140MIT
- AlicenseAqualityDmaintenanceA general MCP server providing web search capabilities using DeepSeek's native online search API.167 npm31MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.MIT