mcp-omnisearch
mcp-omnisearch
Ein Model Context Protocol (MCP)-Server, der einheitlichen Zugriff auf mehrere Suchanbieter und KI-Tools bietet. Dieser Server kombiniert die Funktionen von Tavily, Perplexity, Kagi, Jina AI, Brave und Firecrawl und bietet umfassende Such-, KI-Antwort-, Inhaltsverarbeitungs- und Erweiterungsfunktionen über eine einzige Schnittstelle.
Merkmale
🔍 Suchwerkzeuge
Tavily Search : Optimiert für sachliche Informationen mit starker Zitationsunterstützung. Unterstützt Domänenfilterung über API-Parameter (include_domains/exclude_domains).
Brave Search : Datenschutzorientierte Suche mit guter technischer Inhaltsabdeckung. Bietet native Unterstützung für Suchoperatoren (site:, -site:, filetype:, intitle:, inurl:, before:, after: und exakte Ausdrücke).
Kagi Search : Hochwertige Suchergebnisse mit minimalem Werbeeinfluss, konzentriert auf maßgebliche Quellen. Unterstützt Suchoperatoren in Abfragezeichenfolgen (site:, -site:, filetype:, intitle:, inurl:, before:, after: und exakte Ausdrücke).
🎯 Suchoperatoren
MCP Omnisearch bietet leistungsstarke Suchfunktionen über Operatoren und Parameter:
Allgemeine Suchfunktionen
Domänenfilterung: Anbieterübergreifend verfügbar
Tavily: Über API-Parameter (include_domains/exclude_domains)
Brave & Kagi: Durch site: und -site: Operatoren
Dateitypfilterung: Verfügbar in Brave und Kagi (Dateityp:)
Titel- und URL-Filterung: Verfügbar in Brave und Kagi (intitle:, inurl:)
Datumsfilterung: Verfügbar in Brave und Kagi (vorher:, nachher:)
Genaue Phrasenübereinstimmung: Verfügbar in Brave und Kagi ("Phrase")
Beispielverwendung
// Using Brave or Kagi with query string operators
{
"query": "filetype:pdf site:microsoft.com typescript guide"
}
// Using Tavily with API parameters
{
"query": "typescript guide",
"include_domains": ["microsoft.com"],
"exclude_domains": ["github.com"]
}Anbieterfunktionen
Brave Search : Vollständige native Operatorunterstützung in der Abfragezeichenfolge
Kagi Search : Vollständige Operatorunterstützung in der Abfragezeichenfolge
Tavily Search : Domänenfilterung durch API-Parameter
🤖 KI-Reaktionstools
Perplexity AI : Erweiterte Antwortgenerierung durch Kombination der Echtzeit-Websuche mit GPT-4 Omni und Claude 3
Kagi FastGPT : Schnelle, KI-generierte Antworten mit Zitaten (typische Antwortzeit 900 ms)
📄 Tools zur Inhaltsverarbeitung
Jina AI Reader : Saubere Inhaltsextraktion mit Bildunterschriften und PDF-Unterstützung
Kagi Universal Summarizer : Inhaltszusammenfassung für Seiten, Videos und Podcasts
Tavily Extract : Extrahieren Sie Rohinhalte von einzelnen oder mehreren Webseiten mit konfigurierbarer Extraktionstiefe (einfach oder erweitert). Gibt sowohl kombinierte Inhalte als auch einzelne URL-Inhalte zurück, mit Metadaten wie Wortanzahl und Extraktionsstatistiken.
Firecrawl Scrape : Extrahieren Sie saubere, LLM-fähige Daten aus einzelnen URLs mit erweiterten Formatierungsoptionen
Firecrawl Crawl : Deep Crawling aller erreichbaren Unterseiten einer Website mit konfigurierbaren Tiefenlimits
Firecrawl Map : Schnelle URL-Sammlung von Websites für umfassendes Site-Mapping
Firecrawl Extract : Strukturierte Datenextraktion mit KI unter Verwendung natürlicher Sprachanweisungen
Firecrawl-Aktionen : Unterstützung für Seiteninteraktionen (Klicken, Scrollen usw.) vor der Extraktion dynamischer Inhalte
🔄 Verbesserungstools
Kagi Enrichment API : Ergänzende Inhalte aus spezialisierten Indizes (Teclis, TinyGem)
Jina AI Grounding : Echtzeit-Faktenüberprüfung anhand von Webdaten
Related MCP server: MCP Search Server
Flexible API-Schlüsselanforderungen
MCP Omnisearch ist für die Verwendung mit den verfügbaren API-Schlüsseln konzipiert. Sie benötigen nicht für alle Anbieter Schlüssel – der Server erkennt automatisch, welche API-Schlüssel verfügbar sind und aktiviert nur diese Anbieter.
Zum Beispiel:
Wenn Sie nur einen Tavily- und Perplexity-API-Schlüssel haben, sind nur diese Anbieter verfügbar
Wenn Sie keinen Kagi-API-Schlüssel haben, sind Kagi-basierte Dienste nicht verfügbar, aber alle anderen Anbieter funktionieren normal
Der Server protokolliert, welche Anbieter basierend auf den von Ihnen konfigurierten API-Schlüsseln verfügbar sind
Diese Flexibilität macht es einfach, mit nur einem oder zwei Anbietern zu beginnen und bei Bedarf weitere hinzuzufügen.
Konfiguration
Dieser Server muss über Ihren MCP-Client konfiguriert werden. Hier sind Beispiele für verschiedene Umgebungen:
Cline-Konfiguration
Fügen Sie dies zu Ihren Cline MCP-Einstellungen hinzu:
{
"mcpServers": {
"mcp-omnisearch": {
"command": "node",
"args": ["/path/to/mcp-omnisearch/dist/index.js"],
"env": {
"TAVILY_API_KEY": "your-tavily-key",
"PERPLEXITY_API_KEY": "your-perplexity-key",
"KAGI_API_KEY": "your-kagi-key",
"JINA_AI_API_KEY": "your-jina-key",
"BRAVE_API_KEY": "your-brave-key",
"FIRECRAWL_API_KEY": "your-firecrawl-key"
},
"disabled": false,
"autoApprove": []
}
}
}Claude Desktop mit WSL-Konfiguration
Fügen Sie für WSL-Umgebungen Folgendes zu Ihrer Claude Desktop-Konfiguration hinzu:
{
"mcpServers": {
"mcp-omnisearch": {
"command": "wsl.exe",
"args": [
"bash",
"-c",
"TAVILY_API_KEY=key1 PERPLEXITY_API_KEY=key2 KAGI_API_KEY=key3 JINA_AI_API_KEY=key4 BRAVE_API_KEY=key5 FIRECRAWL_API_KEY=key6 node /path/to/mcp-omnisearch/dist/index.js"
]
}
}
}Umgebungsvariablen
Der Server verwendet API-Schlüssel für jeden Anbieter. Sie benötigen nicht für alle Anbieter Schlüssel – es werden nur die Anbieter aktiviert, die Ihren verfügbaren API-Schlüsseln entsprechen:
TAVILY_API_KEY: Für die Tavily-SuchePERPLEXITY_API_KEY: Für Perplexity AIKAGI_API_KEY: Für Kagi-Dienste (FastGPT, Summarizer, Enrichment)JINA_AI_API_KEY: Für Jina AI-Dienste (Reader, Grounding)BRAVE_API_KEY: Für Brave SearchFIRECRAWL_API_KEY: Für Firecrawl-Dienste (Scrape, Crawl, Map, Extract, Actions)
Sie können mit nur einem oder zwei API-Schlüsseln beginnen und später bei Bedarf weitere hinzufügen. Der Server protokolliert beim Start, welche Anbieter verfügbar sind.
API
Der Server implementiert nach Kategorien organisierte MCP-Tools:
Suchwerkzeuge
search_tavily
Durchsuchen Sie das Internet mit der Tavily Search API. Am besten geeignet für sachliche Abfragen, die zuverlässige Quellen und Zitate erfordern.
Parameter:
query(Zeichenfolge, erforderlich): Suchanfrage
Beispiel:
{
"query": "latest developments in quantum computing"
}search_brave
Datenschutzorientierte Websuche mit guter Abdeckung technischer Themen.
Parameter:
query(Zeichenfolge, erforderlich): Suchanfrage
Beispiel:
{
"query": "rust programming language features"
}search_kagi
Hochwertige Suchergebnisse mit minimalem Werbeeinfluss. Ideal für die Suche nach maßgeblichen Quellen und Forschungsmaterialien.
Parameter:
query(Zeichenfolge, erforderlich): Suchanfragelanguage(Zeichenfolge, optional): Sprachfilter (z. B. „en“)no_cache(boolesch, optional): Cache umgehen für aktuelle Ergebnisse
Beispiel:
{
"query": "latest research in machine learning",
"language": "en"
}KI-Reaktionstools
ai_perplexity
KI-gestützte Antwortgenerierung mit Echtzeit-Websuchintegration.
Parameter:
query(Zeichenfolge, erforderlich): Frage oder Thema für die KI-Antwort
Beispiel:
{
"query": "Explain the differences between REST and GraphQL"
}ai_kagi_fastgpt
Schnelle, KI-generierte Antworten mit Zitaten.
Parameter:
query(Zeichenfolge, erforderlich): Frage für eine schnelle KI-Antwort
Beispiel:
{
"query": "What are the main features of TypeScript?"
}Tools zur Inhaltsverarbeitung
process_jina_reader
Konvertieren Sie URLs in sauberen, LLM-freundlichen Text mit Bildunterschriften.
Parameter:
url(Zeichenfolge, erforderlich): Zu verarbeitende URL
Beispiel:
{
"url": "https://example.com/article"
}process_kagi_summarizer
Inhalte aus URLs zusammenfassen.
Parameter:
url(Zeichenfolge, erforderlich): URL zur Zusammenfassung
Beispiel:
{
"url": "https://example.com/long-article"
}process_tavily_extract
Extrahieren Sie mit Tavily Extract Rohinhalte aus Webseiten.
Parameter:
url(Zeichenfolge | Zeichenfolge[], erforderlich): Einzelne URL oder Array von URLs, aus denen Inhalte extrahiert werden sollenextract_depth(Zeichenfolge, optional): Extraktionstiefe – „basic“ (Standard) oder „advanced“
Beispiel:
{
"url": [
"https://example.com/article1",
"https://example.com/article2"
],
"extract_depth": "advanced"
}Die Antwort umfasst:
Kombinierter Inhalt aller URLs
Individueller Rohinhalt für jede URL
Metadaten mit Wortanzahl, erfolgreichen Extraktionen und allen fehlgeschlagenen URLs
Firecrawl_Scrape_Prozess
Extrahieren Sie saubere, LLM-fähige Daten aus einzelnen URLs mit erweiterten Formatierungsoptionen.
Parameter:
url(Zeichenfolge | Zeichenfolge[], erforderlich): Einzelne URL oder Array von URLs, aus denen Inhalte extrahiert werden sollenextract_depth(Zeichenfolge, optional): Extraktionstiefe – „basic“ (Standard) oder „advanced“
Beispiel:
{
"url": "https://example.com/article",
"extract_depth": "basic"
}Die Antwort umfasst:
Sauberer, Markdown-formatierter Inhalt
Metadaten, einschließlich Titel, Wortanzahl und Extraktionsstatistiken
firecrawl_crawl_prozess
Deep Crawling aller erreichbaren Unterseiten einer Website mit konfigurierbaren Tiefenlimits.
Parameter:
url(string | string[], erforderlich): Start-URL für das Crawlenextract_depth(Zeichenfolge, optional): Extraktionstiefe – „basic“ (Standard) oder „advanced“ (steuert Crawl-Tiefe und -Grenzen)
Beispiel:
{
"url": "https://example.com",
"extract_depth": "advanced"
}Die Antwort umfasst:
Kombinierter Inhalt aller gecrawlten Seiten
Individueller Inhalt für jede Seite
Metadaten, einschließlich Titel, Wortanzahl und Crawl-Statistiken
Firecrawl_Map-Prozess
Schnelle URL-Sammlung von Websites für umfassendes Site-Mapping.
Parameter:
url(Zeichenfolge | Zeichenfolge[], erforderlich): URL zur Karteextract_depth(Zeichenfolge, optional): Extraktionstiefe – „basic“ (Standard) oder „advanced“ (steuert die Kartentiefe)
Beispiel:
{
"url": "https://example.com",
"extract_depth": "basic"
}Die Antwort umfasst:
Liste aller gefundenen URLs
Metadaten, einschließlich Site-Titel und URL-Anzahl
Firecrawl_Extract-Prozess
Strukturierte Datenextraktion mit KI unter Verwendung natürlicher Sprachanweisungen.
Parameter:
url(Zeichenfolge | Zeichenfolge[], erforderlich): URL zum Extrahieren strukturierter Datenextract_depth(Zeichenfolge, optional): Extraktionstiefe – „basic“ (Standard) oder „advanced“
Beispiel:
{
"url": "https://example.com",
"extract_depth": "basic"
}Die Antwort umfasst:
Strukturierte Daten, die aus der Seite extrahiert wurden
Metadaten einschließlich Titel, Extraktionsstatistiken
Firecrawl_Aktionsprozess
Unterstützung für Seiteninteraktionen (Klicken, Scrollen usw.) vor der Extraktion dynamischer Inhalte.
Parameter:
url(Zeichenfolge | Zeichenfolge[], erforderlich): URL zur Interaktion und zum Extrahieren von Inhaltenextract_depth(Zeichenfolge, optional): Extraktionstiefe – „basic“ (Standard) oder „advanced“ (steuert die Komplexität der Interaktionen)
Beispiel:
{
"url": "https://news.ycombinator.com",
"extract_depth": "basic"
}Die Antwort umfasst:
Nach der Durchführung von Interaktionen extrahierter Inhalt
Beschreibung der durchgeführten Aktionen
Screenshot der Seite (falls verfügbar)
Metadaten einschließlich Titel und Extraktionsstatistiken
Verbesserungstools
Kagi-Anreicherung verbessern
Erhalten Sie ergänzende Inhalte aus spezialisierten Indizes.
Parameter:
query(Zeichenfolge, erforderlich): Abfrage zur Anreicherung
Beispiel:
{
"query": "emerging web technologies"
}verbessern_jina_grounding
Überprüfen Sie die Aussagen anhand von Web-Wissen.
Parameter:
statement(Zeichenfolge, erforderlich): Zu überprüfende Anweisung
Beispiel:
{
"statement": "TypeScript adds static typing to JavaScript"
}Entwicklung
Aufstellen
Klonen Sie das Repository
Installieren Sie Abhängigkeiten:
pnpm installErstellen Sie das Projekt:
pnpm run buildIm Entwicklungsmodus ausführen:
pnpm run devVeröffentlichen
Version in package.json aktualisieren
Erstellen Sie das Projekt:
pnpm run buildAuf npm veröffentlichen:
pnpm publishFehlerbehebung
API-Schlüssel und Zugriff
Jeder Anbieter benötigt einen eigenen API-Schlüssel und kann unterschiedliche Zugriffsvoraussetzungen haben:
Tavily : Erfordert einen API-Schlüssel von ihrem Entwicklerportal
Perplexity : API-Zugriff über ihr Entwicklerprogramm
Kagi : Einige Funktionen sind auf Benutzer des Business-Plans (Team) beschränkt
Jina AI : API-Schlüssel für alle Dienste erforderlich
Brave : API-Schlüssel von ihrem Entwicklerportal
Firecrawl : API-Schlüssel von ihrem Entwicklerportal erforderlich
Ratenbegrenzungen
Jeder Anbieter hat seine eigenen Ratenbegrenzungen. Der Server verarbeitet Fehler bei der Ratenbegrenzung ordnungsgemäß und gibt entsprechende Fehlermeldungen zurück.
Beitragen
Beiträge sind willkommen! Senden Sie gerne einen Pull Request.
Lizenz
MIT-Lizenz – Einzelheiten finden Sie in der Datei LICENSE .
Danksagung
Aufbauend auf:
Available Tools
3 toolsai_searchGet AI-powered answers with citations and reasoning. Use when you need synthesized answers rather than raw search results. Providers: kagi_fastgpt (fast answers), exa_answer (semantic AI), linkup (deep agentic search), tavily_research (asynchronous multi-search reports; resubmit its research_id to retrieve results).BRead-onlyIdempotent
Get AI-powered answers with citations and reasoning. Use when you need synthesized answers rather than raw search results. Providers: kagi_fastgpt (fast answers), exa_answer (semantic AI), linkup (deep agentic search), tavily_research (asynchronous multi-search reports; resubmit its research_id to retrieve results).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | Yes | Search query | |
| provider | Yes | AI search provider to use | |
| research_id | No | Existing asynchronous research task ID to retrieve. Supported by Tavily Research. | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent safety, so the description adds value by disclosing asynchronous retrieval behavior for tavily_research ('resubmit its research_id'), provider-specific behaviors, and the output nature (citations and reasoning). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose front-loaded and provider details compactly listed. Every clause contributes meaning without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers when-to-use, provider differences, and the async resubmission pattern, and annotations cover safety. However, it omits response structure and contains a provider/enum inconsistency, leaving an agent with an ambiguous picture for a 5-parameter tool without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the description need not repeat parameter details, but it adds provider characteristics that conflict with the enum by naming exa_answer and linkup which are not valid values. It does usefully explain research_id for Tavily, but the misinformation undermines reliability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use when you need synthesized answers rather than raw search results', giving a clear when-to-use signal. However, it lists exa_answer and linkup as providers even though the schema enum only allows kagi_fastgpt and tavily_research, making the provider-selection guidance partially misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_extractExtract, process, or summarize web content from URLs. Use when you need to read page content, summarize articles, crawl sites, or extract structured data. Providers: tavily (content extraction), kagi (summarization of pages/videos/podcasts), firecrawl (scraping/crawling/mapping/structured extraction/interactive), exa (content retrieval/similar pages).BRead-onlyIdempotent
Extract, process, or summarize web content from URLs. Use when you need to read page content, summarize articles, crawl sites, or extract structured data. Providers: tavily (content extraction), kagi (summarization of pages/videos/podcasts), firecrawl (scraping/crawling/mapping/structured extraction/interactive), exa (content retrieval/similar pages).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL or array of URLs to process | |
| mode | No | Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract/crawl/map. Kagi: summarize. Defaults to provider default. | |
| query | No | Focus extracted content on information relevant to this query. | |
| format | No | Extracted page format (default: markdown). | |
| provider | Yes | Processing provider to use | |
| extract_depth | No | Extraction depth (default: basic) | |
| chunks_per_source | No | Maximum relevant content chunks per source when a query is provided. | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. | |
| include_raw_contents | No | Whether extraction responses should include per-URL raw_contents alongside combined content (default: true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds provider capability context (e.g., firecrawl for scraping/crawling/interactive, kagi for summarization). However, the mention of 'exa' as a provider is misleading because the input schema's provider enum omits exa, creating uncertainty about available behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loads the purpose, and packs useful provider information into a short list. Minor redundancy ('process' with 'extract') and the misleading exa reference are the only blemishes; overall it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with 5 enums and provider-specific modes, the description is too thin. It does not explain how to choose a provider for a given task, what the different modes do relative to providers, or how to handle edge cases like exa's absence from the provider enum. There is no output schema, and the description gives no hint about return value shapes, so an agent would likely need to infer provider-mode compatibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are fully documented in the schema. The description adds value by loosely mapping providers to capabilities (e.g., kagi for summarization, firecrawl for scraping), which helps select provider and mode. But it introduces a conflict by listing exa although the provider enum does not include it, and it does not explain the relationship between provider and mode beyond the parentheticals.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when...' conditions covering the main scenarios, and the provider list gives a starting point for mode selection. It does not state when not to use this tool or point to alternatives like web_search or ai_search, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchSearch the web for information. Use when you need to find web pages, articles, or data. Providers: tavily (factual/citations and search controls), brave (privacy/operators), kagi (quality/operators), exa (AI-semantic), kagi_enrichment (specialized indexes). Search depth, topic, time range, safe search, raw content, and automatic parameters apply when supported by the provider.BRead-onlyIdempotent
Search the web for information. Use when you need to find web pages, articles, or data. Providers: tavily (factual/citations and search controls), brave (privacy/operators), kagi (quality/operators), exa (AI-semantic), kagi_enrichment (specialized indexes). Search depth, topic, time range, safe search, raw content, and automatic parameters apply when supported by the provider.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | Yes | Search query | |
| topic | No | Search topic category. | |
| provider | Yes | Search provider to use | |
| time_range | No | Only return results from this recent time range. | |
| safe_search | No | Enable provider safe-search filtering. | |
| search_depth | No | Search depth. Providers may use this to balance speed, relevance, and cost. | |
| auto_parameters | No | Let supported providers select search settings from the query. This can change cost. | |
| exclude_domains | No | Exclude results from these domains | |
| include_domains | No | Only return results from these domains | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. | |
| include_raw_content | No | Include full page content when the selected provider supports it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false), lowering the burden on the description. The description adds useful context that search depth, topic, time range, safe search, raw content, and auto_parameters behave conditionally 'when supported by the provider,' but it does not disclose result-format, citation, pagination, or cost behavior, which would add meaningful transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose appears in the first sentence, usage in the second, and the provider rundown is dense but informative. The title is a verbatim duplicate of the description, which is mildly redundant, but no other space is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, 5 enums, and no output schema, the description carries a heavy burden. It covers purpose, usage context, and provider-specific parameter behavior, but it does not describe the result format or return expectations, and the provider list conflicts with the schema enum. For such a configurable tool, these are meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description does add meta-information: several parameters (search_depth, topic, time_range, safe_search, raw_content, auto_parameters) are provider-dependent, which is not in the schema. However, the description lists 'exa' as a valid provider while the schema enum omits it, which could mislead an agent into sending an invalid provider value and offset the added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-to-use directive ('Use when you need to find web pages, articles, or data') and adds provider-selection guidance by use case (e.g., tavily for factual/citations, brave for privacy/operators). It does not mention sibling alternatives or state when not to use this tool, stopping short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- Changed
ai_search2 fields changed- changed
Input schema / properties / provider / enumPrevious value: -[ - "kagi_fastgpt" -]New value: +[ + "kagi_fastgpt", + "tavily_research" +] - added
Input schema / properties / research_idAdded value: +{ + "description": "Existing asynchronous research task ID to retrieve. Supported by Tavily Research.", + "minLength": 1, + "type": "string" +}
- Changed
web_extract5 fields changed- added
Input schema / properties / chunks_per_sourceAdded value: +{ + "description": "Maximum relevant content chunks per source when a query is provided.", + "maximum": 5, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / formatAdded value: +{ + "description": "Extracted page format (default: markdown).", + "enum": [ + "markdown", + "text" + ], + "type": "string" +} - changed
Input schema / properties / mode / descriptionPrevious value: -"Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract. Kagi: summarize. Defaults to provider default."New value: +"Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract/crawl/map. Kagi: summarize. Defaults to provider default." - changed
Input schema / properties / mode / enumPrevious value: -[ - "extract", - "summarize", - "scrape", - "crawl", - "map", - "actions", - "contents", - "similar" -]New value: +[ + "extract", + "crawl", + "map", + "summarize", + "scrape", + "actions", + "contents", + "similar" +] - added
Input schema / properties / queryAdded value: +{ + "description": "Focus extracted content on information relevant to this query.", + "minLength": 1, + "pattern": "\\S", + "type": "string" +}
- Changed
web_search6 fields changed- added
Input schema / properties / auto_parametersAdded value: +{ + "description": "Let supported providers select search settings from the query. This can change cost.", + "type": "boolean" +} - added
Input schema / properties / include_raw_contentAdded value: +{ + "description": "Include full page content when the selected provider supports it.", + "type": "boolean" +} - added
Input schema / properties / safe_searchAdded value: +{ + "description": "Enable provider safe-search filtering.", + "type": "boolean" +} - added
Input schema / properties / search_depthAdded value: +{ + "description": "Search depth. Providers may use this to balance speed, relevance, and cost.", + "enum": [ + "basic", + "advanced", + "fast", + "ultra-fast" + ], + "type": "string" +} - added
Input schema / properties / time_rangeAdded value: +{ + "description": "Only return results from this recent time range.", + "enum": [ + "day", + "week", + "month", + "year" + ], + "type": "string" +} - added
Input schema / properties / topicAdded value: +{ + "description": "Search topic category.", + "enum": [ + "general", + "news", + "finance" + ], + "type": "string" +}
3 tool updates
v0.0.29- Changed
ai_search7 fields changed- added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - added
Input schema / properties / limit / maximumAdded value: +50 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - changed
Input schema / properties / query / descriptionPrevious value: -"Question or search query"New value: +"Search query" - added
Input schema / properties / query / minLengthAdded value: +1 - added
Input schema / properties / query / patternAdded value: +"\\S"
- Changed
web_extract3 fields changed- added
Input schema / properties / include_raw_contentsAdded value: +{ + "description": "Whether extraction responses should include per-URL raw_contents alongside combined content (default: true).", + "type": "boolean" +} - added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - changed
Input schema / properties / url / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "type": "string" - }, - "type": "array" - } -]New value: +[ + { + "format": "uri", + "pattern": "^https?:\\/\\/", + "type": "string" + }, + { + "items": { + "format": "uri", + "pattern": "^https?:\\/\\/", + "type": "string" + }, + "maxItems": 10, + "minItems": 1, + "type": "array" + } +]
- Changed
web_search10 fields changed- added
Input schema / properties / exclude_domains / items / patternAdded value: +"^(?:\\*\\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\\.)+[a-zA-Z]{2,63}$" - added
Input schema / properties / exclude_domains / maxItemsAdded value: +20 - added
Input schema / properties / include_domains / items / patternAdded value: +"^(?:\\*\\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\\.)+[a-zA-Z]{2,63}$" - added
Input schema / properties / include_domains / maxItemsAdded value: +20 - added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - added
Input schema / properties / limit / maximumAdded value: +50 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / query / minLengthAdded value: +1 - added
Input schema / properties / query / patternAdded value: +"\\S"
8 tool updates
v0.0.4- Changed
ai_search6 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Result limit"New value: +"Maximum number of results (default: 10)" - removed
Input schema / properties / provider / anyOfRemoved value: -[ - { - "const": "perplexity" - }, - { - "const": "kagi_fastgpt" - }, - { - "const": "exa_answer" - } -] - changed
Input schema / properties / provider / descriptionPrevious value: -"AI provider"New value: +"AI search provider to use" - added
Input schema / properties / provider / enumAdded value: +[ + "kagi_fastgpt" +] - added
Input schema / properties / provider / typeAdded value: +"string" - changed
Input schema / properties / query / descriptionPrevious value: -"Query"New value: +"Question or search query"
- Removed
firecrawl_process - Removed
jina_grounding_enhance - Removed
kagi_enrichment_enhance - Removed
kagi_summarizer_process - Removed
tavily_extract_process - Added
web_extract - Changed
web_search8 fields changed- changed
Input schema / properties / exclude_domains / descriptionPrevious value: -"Domains to exclude"New value: +"Exclude results from these domains" - changed
Input schema / properties / include_domains / descriptionPrevious value: -"Domains to include"New value: +"Only return results from these domains" - changed
Input schema / properties / limit / descriptionPrevious value: -"Result limit"New value: +"Maximum number of results (default: 10)" - removed
Input schema / properties / provider / anyOfRemoved value: -[ - { - "const": "tavily" - }, - { - "const": "brave" - }, - { - "const": "kagi" - }, - { - "const": "exa" - } -] - changed
Input schema / properties / provider / descriptionPrevious value: -"Search provider"New value: +"Search provider to use" - added
Input schema / properties / provider / enumAdded value: +[ + "tavily", + "brave", + "kagi", + "kagi_enrichment" +] - added
Input schema / properties / provider / typeAdded value: +"string" - changed
Input schema / properties / query / descriptionPrevious value: -"Query"New value: +"Search query"
18 tool updates
v1.0.0- Added
ai_search - Removed
brave_search - Removed
firecrawl_actions_process - Removed
firecrawl_crawl_process - Removed
firecrawl_extract_process - Removed
firecrawl_map_process - Added
firecrawl_process - Removed
firecrawl_scrape_process - Changed
jina_grounding_enhance2 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / content / descriptionPrevious value: -"Content to enhance"New value: +"Content"
- Removed
jina_reader_process - Changed
kagi_enrichment_enhance2 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / content / descriptionPrevious value: -"Content to enhance"New value: +"Content"
- Removed
kagi_fastgpt_search - Removed
kagi_search - Changed
kagi_summarizer_process9 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / extract_depth / anyOfAdded value: +[ + { + "const": "basic" + }, + { + "const": "advanced" + } +] - removed
Input schema / properties / extract_depth / defaultRemoved value: -"basic" - changed
Input schema / properties / extract_depth / descriptionPrevious value: -"The depth of the extraction process. \"advanced\" retrieves more data but costs more credits."New value: +"Extraction depth" - removed
Input schema / properties / extract_depth / enumRemoved value: -[ - "basic", - "advanced" -] - removed
Input schema / properties / extract_depth / typeRemoved value: -"string" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + } +] - added
Input schema / properties / url / descriptionAdded value: +"URL(s)" - removed
Input schema / properties / url / oneOfRemoved value: -[ - { - "description": "Single URL to process", - "type": "string" - }, - { - "description": "Multiple URLs to process", - "items": { - "type": "string" - }, - "type": "array" - } -]
- Removed
perplexity_search - Changed
tavily_extract_process9 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / extract_depth / anyOfAdded value: +[ + { + "const": "basic" + }, + { + "const": "advanced" + } +] - removed
Input schema / properties / extract_depth / defaultRemoved value: -"basic" - changed
Input schema / properties / extract_depth / descriptionPrevious value: -"The depth of the extraction process. \"advanced\" retrieves more data but costs more credits."New value: +"Extraction depth" - removed
Input schema / properties / extract_depth / enumRemoved value: -[ - "basic", - "advanced" -] - removed
Input schema / properties / extract_depth / typeRemoved value: -"string" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + } +] - added
Input schema / properties / url / descriptionAdded value: +"URL(s)" - removed
Input schema / properties / url / oneOfRemoved value: -[ - { - "description": "Single URL to process", - "type": "string" - }, - { - "description": "Multiple URLs to process", - "items": { - "type": "string" - }, - "type": "array" - } -]
- Removed
tavily_search - Added
web_search
15 tool updates
- First observed
brave_search - First observed
firecrawl_actions_process - First observed
firecrawl_crawl_process - First observed
firecrawl_extract_process - First observed
firecrawl_map_process - First observed
firecrawl_scrape_process - First observed
jina_grounding_enhance - First observed
jina_reader_process - First observed
kagi_enrichment_enhance - First observed
kagi_fastgpt_search - First observed
kagi_search - First observed
kagi_summarizer_process - First observed
perplexity_search - First observed
tavily_extract_process - First observed
tavily_search
TDQS
Scored across 3 tools
Each tool has a clearly distinct job: web_search returns raw search results, ai_search returns synthesized answers with citations, and web_extract processes specific URLs. There is no meaningful overlap or ambiguity between them.
All tool names follow a consistent snake_case pattern combining a domain prefix with an action: web_search, ai_search, web_extract. The naming style is uniform and predictable.
Three tools is well-scoped for an omnisearch server covering the core needs of searching, getting AI answers, and extracting web content. Each tool earns its place without redundancy.
The toolset covers the full search-to-insight workflow: finding sources, getting synthesized answers, and extracting or summarizing content from URLs. There are no obvious dead ends or missing core operations for the stated purpose.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Jina AI Reader/Search MCP — turn any URL into clean LLM-ready markdown, plus web search.
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Related MCP Servers
- AlicenseAqualityCmaintenanceA Model Context Protocol server that enables web search, scraping, crawling, and content extraction through multiple engines including SearXNG, Firecrawl, and Tavily.4338 npm143MIT
- -licenseNot gradedqualityNot gradedmaintenanceA unified Model Context Protocol server that integrates multiple search providers including Brave, Tavily, Exa, Semantic Scholar, and arXiv. It enables users to perform web, news, and image searches alongside academic research and citation analysis through a single interface.-
- AlicenseBqualityAmaintenanceA Model Context Protocol server that gives AI assistants access to 7 search providers with intelligent auto-routing. Analyzes query intent and picks the best provider automatically — no manual switching needed. Install, configure your keys, and go.2417 PyPI5MIT
- AlicenseNot gradedqualityDmaintenanceA unified MCP server aggregating 15 web search and extraction tools across 5 providers (Jina, Tavily, Exa, Firecrawl, Bocha) with automatic API key validation and plugin architecture.MIT