mcp-omnisearch
mcp-omnisearch
複数の検索プロバイダーとAIツールへの統合アクセスを提供するモデルコンテキストプロトコル(MCP)サーバー。Tavily、Perplexity、Kagi、Jina AI、Brave、Firecrawlの機能を統合し、包括的な検索、AIレスポンス、コンテンツ処理、拡張機能を単一のインターフェースで提供します。
特徴
🔍 検索ツール
Tavily Search :強力な引用サポートを備え、事実情報に最適化されています。APIパラメータ(include_domains/exclude_domains)によるドメインフィルタリングをサポートします。
Brave Search :プライバシーを重視し、幅広い技術コンテンツを網羅した検索機能。検索演算子(site:、-site:、filetype:、intitle:、inurl:、before:、after:、exact phrases)をネイティブサポート。
Kagi Search :広告の影響を最小限に抑え、信頼性の高い情報源に重点を置いた高品質な検索結果を提供します。クエリ文字列での検索演算子(site:、-site:、filetype:、intitle:、inurl:、before:、after:、完全一致フレーズ)をサポートします。
🎯 検索演算子
MCP Omnisearch は、演算子とパラメータを通じて強力な検索機能を提供します。
一般的な検索機能
ドメインフィルタリング: すべてのプロバイダーで利用可能
Tavily: APIパラメータ(include_domains/exclude_domains)を通じて
Brave & Kagi: site: と -site: 演算子を通じて
ファイルタイプのフィルタリング: Brave と Kagi で利用可能 (filetype:)
タイトルとURLフィルタリング:BraveとKagiで利用可能(intitle:、inurl:)
日付フィルタリング: Brave と Kagi で利用可能 (before:、after:)
完全フレーズ一致: Brave と Kagi で利用可能 (「フレーズ」)
使用例
// Using Brave or Kagi with query string operators
{
"query": "filetype:pdf site:microsoft.com typescript guide"
}
// Using Tavily with API parameters
{
"query": "typescript guide",
"include_domains": ["microsoft.com"],
"exclude_domains": ["github.com"]
}プロバイダーの機能
Brave Search : クエリ文字列でのネイティブ演算子の完全サポート
Kagi Search : クエリ文字列での完全な演算子サポート
Tavily Search : APIパラメータによるドメインフィルタリング
🤖 AI対応ツール
Perplexity AI : リアルタイムウェブ検索とGPT-4 Omni、Claude 3を組み合わせた高度な応答生成
Kagi FastGPT : 引用付きの AI 生成の迅速な回答 (通常応答時間は 900 ミリ秒)
📄 コンテンツ処理ツール
Jina AI Reader :画像キャプションとPDFサポートによるクリーンなコンテンツ抽出
Kagi Universal Summarizer :ページ、ビデオ、ポッドキャストのコンテンツ要約
Tavily Extract :単一または複数のウェブページから、設定可能な抽出深度(「基本」または「詳細」)で生のコンテンツを抽出します。単語数や抽出統計などのメタデータとともに、結合コンテンツと個々のURLコンテンツの両方を返します。
Firecrawl Scrape : 強化されたフォーマットオプションを使用して、単一の URL からクリーンで LLM 対応のデータを抽出します。
Firecrawl クロール: 設定可能な深度制限を使用して、ウェブサイト上のアクセス可能なすべてのサブページをディープクロールします。
Firecrawl Map : ウェブサイトから URL を高速収集し、包括的なサイト マッピングを実現します
Firecrawl Extract : 自然言語プロンプトを使用した AI による構造化データ抽出
Firecrawl アクション: 動的コンテンツの抽出前のページ操作 (クリック、スクロールなど) のサポート
🔄 強化ツール
Kagi Enrichment API : 専門インデックス (Teclis、TinyGem) からの補足コンテンツ
Jina AI Grounding :Web知識に基づくリアルタイムの事実検証
Related MCP server: MCP Search Server
柔軟なAPIキー要件
MCP Omnisearchは、利用可能なAPIキーで動作するように設計されています。すべてのプロバイダーのAPIキーを用意する必要はありません。サーバーが利用可能なAPIキーを自動的に検出し、それらのプロバイダーのみを有効にします。
例えば:
TavilyとPerplexityのAPIキーのみをお持ちの場合は、それらのプロバイダーのみが利用可能になります。
Kagi APIキーをお持ちでない場合、Kagiベースのサービスは利用できませんが、他のすべてのプロバイダーは正常に動作します。
サーバーは、設定したAPIキーに基づいて利用可能なプロバイダーを記録します。
この柔軟性により、1 つまたは 2 つのプロバイダーから簡単に開始し、必要に応じてさらに追加することができます。
構成
このサーバーはMCPクライアント経由で設定する必要があります。以下に、様々な環境における設定例を示します。
傾斜構成
Cline MCP 設定に以下を追加します:
{
"mcpServers": {
"mcp-omnisearch": {
"command": "node",
"args": ["/path/to/mcp-omnisearch/dist/index.js"],
"env": {
"TAVILY_API_KEY": "your-tavily-key",
"PERPLEXITY_API_KEY": "your-perplexity-key",
"KAGI_API_KEY": "your-kagi-key",
"JINA_AI_API_KEY": "your-jina-key",
"BRAVE_API_KEY": "your-brave-key",
"FIRECRAWL_API_KEY": "your-firecrawl-key"
},
"disabled": false,
"autoApprove": []
}
}
}WSL 構成の Claude デスクトップ
WSL 環境の場合は、Claude Desktop 構成に以下を追加します。
{
"mcpServers": {
"mcp-omnisearch": {
"command": "wsl.exe",
"args": [
"bash",
"-c",
"TAVILY_API_KEY=key1 PERPLEXITY_API_KEY=key2 KAGI_API_KEY=key3 JINA_AI_API_KEY=key4 BRAVE_API_KEY=key5 FIRECRAWL_API_KEY=key6 node /path/to/mcp-omnisearch/dist/index.js"
]
}
}
}環境変数
サーバーは各プロバイダーのAPIキーを使用します。すべてのプロバイダーのキーは必要ありません。利用可能なAPIキーに対応するプロバイダーのみが有効化されます。
TAVILY_API_KEY: Tavily Search用PERPLEXITY_API_KEY: Perplexity AI用KAGI_API_KEY: Kagi サービス (FastGPT、Summarizer、Enrichment) 用JINA_AI_API_KEY: Jina AIサービス(Reader、Grounding)用BRAVE_API_KEY: Brave Search用FIRECRAWL_API_KEY: Firecrawl サービス (スクレイプ、クロール、マップ、抽出、アクション) 用
まずは1つか2つのAPIキーから始めて、必要に応じて後から追加できます。サーバーは起動時に利用可能なプロバイダーを記録します。
API
サーバーは、カテゴリ別に整理された MCP ツールを実装します。
検索ツール
search_tavily
Tavily Search APIを使ってウェブを検索します。信頼できる情報源や引用を必要とする事実検索に最適です。
パラメータ:
query(文字列、必須): 検索クエリ
例:
{
"query": "latest developments in quantum computing"
}検索_勇敢
技術的なトピックを幅広くカバーした、プライバシー重視の Web 検索。
パラメータ:
query(文字列、必須): 検索クエリ
例:
{
"query": "rust programming language features"
}検索カギ
広告の影響を最小限に抑えた高品質な検索結果。信頼できる情報源や研究資料を見つけるのに最適です。
パラメータ:
query(文字列、必須): 検索クエリlanguage(文字列、オプション): 言語フィルター (例: "en")no_cache(ブール値、オプション): 最新の結果を得るためにキャッシュをバイパスする
例:
{
"query": "latest research in machine learning",
"language": "en"
}AI対応ツール
ai_perplexity
リアルタイムの Web 検索統合による AI を活用した応答生成。
パラメータ:
query(文字列、必須): AI 応答の質問またはトピック
例:
{
"query": "Explain the differences between REST and GraphQL"
}ai_kagi_fastgpt
引用付きの AI 生成の迅速な回答。
パラメータ:
query(文字列、必須): 迅速な AI 応答を求める質問
例:
{
"query": "What are the main features of TypeScript?"
}コンテンツ処理ツール
プロセス_jina_reader
URL を、画像キャプション付きのクリーンな LLM 対応テキストに変換します。
パラメータ:
url(文字列、必須): 処理するURL
例:
{
"url": "https://example.com/article"
}プロセスカギサマライザー
URL からコンテンツを要約します。
パラメータ:
url(文字列、必須): 要約するURL
例:
{
"url": "https://example.com/long-article"
}プロセス_tavily_extract
Tavily Extract を使用して、Web ページから生のコンテンツを抽出します。
パラメータ:
url(文字列 | 文字列[], 必須): コンテンツを抽出する単一のURLまたはURLの配列extract_depth(文字列、オプション): 抽出の深さ - 'basic' (デフォルト) または 'advanced'
例:
{
"url": [
"https://example.com/article1",
"https://example.com/article2"
],
"extract_depth": "advanced"
}回答には以下が含まれます:
すべての URL の結合されたコンテンツ
各 URL の個別の生のコンテンツ
単語数、成功した抽出、失敗した URL を含むメタデータ
firecrawl_scrape_process
強化されたフォーマット オプションを使用して、単一の URL からクリーンな LLM 対応データを抽出します。
パラメータ:
url(文字列 | 文字列[], 必須): コンテンツを抽出する単一のURLまたはURLの配列extract_depth(文字列、オプション): 抽出の深さ - 'basic' (デフォルト) または 'advanced'
例:
{
"url": "https://example.com/article",
"extract_depth": "basic"
}回答には以下が含まれます:
クリーンなマークダウン形式のコンテンツ
タイトル、単語数、抽出統計などのメタデータ
firecrawl_crawl_process
設定可能な深度制限を使用して、Web サイト上のアクセス可能なすべてのサブページをディープ クロールします。
パラメータ:
url(文字列 | 文字列[], 必須): クロールの開始URLextract_depth(文字列、オプション): 抽出深度 - 'basic' (デフォルト) または 'advanced' (クロール深度と制限を制御)
例:
{
"url": "https://example.com",
"extract_depth": "advanced"
}回答には以下が含まれます:
クロールされたすべてのページの結合されたコンテンツ
各ページの個別のコンテンツ
タイトル、単語数、クロール統計などのメタデータ
firecrawl_map_process
包括的なサイト マッピングのために、Web サイトから URL を高速に収集します。
パラメータ:
url(文字列 | 文字列[], 必須): マップするURLextract_depth(文字列、オプション): 抽出深度 - 'basic' (デフォルト) または 'advanced' (マップ深度を制御)
例:
{
"url": "https://example.com",
"extract_depth": "basic"
}回答には以下が含まれます:
検出されたすべてのURLのリスト
サイトのタイトルやURL数などのメタデータ
firecrawl_extract_process
自然言語プロンプトを使用した AI による構造化データ抽出。
パラメータ:
url(文字列 | 文字列[], 必須): 構造化データを抽出するURLextract_depth(文字列、オプション): 抽出の深さ - 'basic' (デフォルト) または 'advanced'
例:
{
"url": "https://example.com",
"extract_depth": "basic"
}回答には以下が含まれます:
ページから抽出された構造化データ
タイトル、抽出統計などのメタデータ
firecrawl_actions_process
動的コンテンツの抽出前のページ操作 (クリック、スクロールなど) をサポートします。
パラメータ:
url(文字列 | 文字列[], 必須): 対話してコンテンツを抽出するためのURLextract_depth(文字列、オプション): 抽出の深さ - 'basic' (デフォルト) または 'advanced' (インタラクションの複雑さを制御)
例:
{
"url": "https://news.ycombinator.com",
"extract_depth": "basic"
}回答には以下が含まれます:
インタラクション実行後に抽出されたコンテンツ
実行されたアクションの説明
ページのスクリーンショット(可能な場合)
タイトルや抽出統計などのメタデータ
強化ツール
鍵強化を強化する
専門のインデックスから補足コンテンツを取得します。
パラメータ:
query(文字列、必須): エンリッチメントのクエリ
例:
{
"query": "emerging web technologies"
}強化されたジナの接地
ウェブの知識に照らしてステートメントを検証します。
パラメータ:
statement(文字列、必須): 検証するステートメント
例:
{
"statement": "TypeScript adds static typing to JavaScript"
}発達
設定
リポジトリをクローンする
依存関係をインストールします:
pnpm installプロジェクトをビルドします。
pnpm run build開発モードで実行:
pnpm run dev出版
package.json のバージョンを更新する
プロジェクトをビルドします。
pnpm run buildnpm に公開:
pnpm publishトラブルシューティング
APIキーとアクセス
各プロバイダーには独自の API キーが必要であり、アクセス要件が異なる場合があります。
Tavily : 開発者ポータルからのAPIキーが必要
困惑:開発者プログラムを通じたAPIアクセス
Kagi : 一部の機能はビジネス(チーム)プランユーザーに限定されます
Jina AI : すべてのサービスにAPIキーが必要
Brave : 開発者ポータルからのAPIキー
Firecrawl : 開発者ポータルからAPIキーが必要
レート制限
各プロバイダーには独自のレート制限があります。サーバーはレート制限エラーを適切に処理し、適切なエラーメッセージを返します。
貢献
貢献を歓迎します!お気軽にプルリクエストを送信してください。
ライセンス
MIT ライセンス - 詳細についてはLICENSEファイルを参照してください。
謝辞
構築されたもの:
Available Tools
3 toolsai_searchGet AI-powered answers with citations and reasoning. Use when you need synthesized answers rather than raw search results. Providers: kagi_fastgpt (fast answers), exa_answer (semantic AI), linkup (deep agentic search), tavily_research (asynchronous multi-search reports; resubmit its research_id to retrieve results).BRead-onlyIdempotent
Get AI-powered answers with citations and reasoning. Use when you need synthesized answers rather than raw search results. Providers: kagi_fastgpt (fast answers), exa_answer (semantic AI), linkup (deep agentic search), tavily_research (asynchronous multi-search reports; resubmit its research_id to retrieve results).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | Yes | Search query | |
| provider | Yes | AI search provider to use | |
| research_id | No | Existing asynchronous research task ID to retrieve. Supported by Tavily Research. | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent safety, so the description adds value by disclosing asynchronous retrieval behavior for tavily_research ('resubmit its research_id'), provider-specific behaviors, and the output nature (citations and reasoning). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose front-loaded and provider details compactly listed. Every clause contributes meaning without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers when-to-use, provider differences, and the async resubmission pattern, and annotations cover safety. However, it omits response structure and contains a provider/enum inconsistency, leaving an agent with an ambiguous picture for a 5-parameter tool without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the description need not repeat parameter details, but it adds provider characteristics that conflict with the enum by naming exa_answer and linkup which are not valid values. It does usefully explain research_id for Tavily, but the misinformation undermines reliability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use when you need synthesized answers rather than raw search results', giving a clear when-to-use signal. However, it lists exa_answer and linkup as providers even though the schema enum only allows kagi_fastgpt and tavily_research, making the provider-selection guidance partially misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_extractExtract, process, or summarize web content from URLs. Use when you need to read page content, summarize articles, crawl sites, or extract structured data. Providers: tavily (content extraction), kagi (summarization of pages/videos/podcasts), firecrawl (scraping/crawling/mapping/structured extraction/interactive), exa (content retrieval/similar pages).BRead-onlyIdempotent
Extract, process, or summarize web content from URLs. Use when you need to read page content, summarize articles, crawl sites, or extract structured data. Providers: tavily (content extraction), kagi (summarization of pages/videos/podcasts), firecrawl (scraping/crawling/mapping/structured extraction/interactive), exa (content retrieval/similar pages).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL or array of URLs to process | |
| mode | No | Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract/crawl/map. Kagi: summarize. Defaults to provider default. | |
| query | No | Focus extracted content on information relevant to this query. | |
| format | No | Extracted page format (default: markdown). | |
| provider | Yes | Processing provider to use | |
| extract_depth | No | Extraction depth (default: basic) | |
| chunks_per_source | No | Maximum relevant content chunks per source when a query is provided. | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. | |
| include_raw_contents | No | Whether extraction responses should include per-URL raw_contents alongside combined content (default: true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds provider capability context (e.g., firecrawl for scraping/crawling/interactive, kagi for summarization). However, the mention of 'exa' as a provider is misleading because the input schema's provider enum omits exa, creating uncertainty about available behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loads the purpose, and packs useful provider information into a short list. Minor redundancy ('process' with 'extract') and the misleading exa reference are the only blemishes; overall it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with 5 enums and provider-specific modes, the description is too thin. It does not explain how to choose a provider for a given task, what the different modes do relative to providers, or how to handle edge cases like exa's absence from the provider enum. There is no output schema, and the description gives no hint about return value shapes, so an agent would likely need to infer provider-mode compatibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are fully documented in the schema. The description adds value by loosely mapping providers to capabilities (e.g., kagi for summarization, firecrawl for scraping), which helps select provider and mode. But it introduces a conflict by listing exa although the provider enum does not include it, and it does not explain the relationship between provider and mode beyond the parentheticals.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when...' conditions covering the main scenarios, and the provider list gives a starting point for mode selection. It does not state when not to use this tool or point to alternatives like web_search or ai_search, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchSearch the web for information. Use when you need to find web pages, articles, or data. Providers: tavily (factual/citations and search controls), brave (privacy/operators), kagi (quality/operators), exa (AI-semantic), kagi_enrichment (specialized indexes). Search depth, topic, time range, safe search, raw content, and automatic parameters apply when supported by the provider.BRead-onlyIdempotent
Search the web for information. Use when you need to find web pages, articles, or data. Providers: tavily (factual/citations and search controls), brave (privacy/operators), kagi (quality/operators), exa (AI-semantic), kagi_enrichment (specialized indexes). Search depth, topic, time range, safe search, raw content, and automatic parameters apply when supported by the provider.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | Yes | Search query | |
| topic | No | Search topic category. | |
| provider | Yes | Search provider to use | |
| time_range | No | Only return results from this recent time range. | |
| safe_search | No | Enable provider safe-search filtering. | |
| search_depth | No | Search depth. Providers may use this to balance speed, relevance, and cost. | |
| auto_parameters | No | Let supported providers select search settings from the query. This can change cost. | |
| exclude_domains | No | Exclude results from these domains | |
| include_domains | No | Only return results from these domains | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. | |
| include_raw_content | No | Include full page content when the selected provider supports it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false), lowering the burden on the description. The description adds useful context that search depth, topic, time range, safe search, raw content, and auto_parameters behave conditionally 'when supported by the provider,' but it does not disclose result-format, citation, pagination, or cost behavior, which would add meaningful transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose appears in the first sentence, usage in the second, and the provider rundown is dense but informative. The title is a verbatim duplicate of the description, which is mildly redundant, but no other space is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, 5 enums, and no output schema, the description carries a heavy burden. It covers purpose, usage context, and provider-specific parameter behavior, but it does not describe the result format or return expectations, and the provider list conflicts with the schema enum. For such a configurable tool, these are meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description does add meta-information: several parameters (search_depth, topic, time_range, safe_search, raw_content, auto_parameters) are provider-dependent, which is not in the schema. However, the description lists 'exa' as a valid provider while the schema enum omits it, which could mislead an agent into sending an invalid provider value and offset the added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-to-use directive ('Use when you need to find web pages, articles, or data') and adds provider-selection guidance by use case (e.g., tavily for factual/citations, brave for privacy/operators). It does not mention sibling alternatives or state when not to use this tool, stopping short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- Changed
ai_search2 fields changed- changed
Input schema / properties / provider / enumPrevious value: -[ - "kagi_fastgpt" -]New value: +[ + "kagi_fastgpt", + "tavily_research" +] - added
Input schema / properties / research_idAdded value: +{ + "description": "Existing asynchronous research task ID to retrieve. Supported by Tavily Research.", + "minLength": 1, + "type": "string" +}
- Changed
web_extract5 fields changed- added
Input schema / properties / chunks_per_sourceAdded value: +{ + "description": "Maximum relevant content chunks per source when a query is provided.", + "maximum": 5, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / formatAdded value: +{ + "description": "Extracted page format (default: markdown).", + "enum": [ + "markdown", + "text" + ], + "type": "string" +} - changed
Input schema / properties / mode / descriptionPrevious value: -"Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract. Kagi: summarize. Defaults to provider default."New value: +"Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract/crawl/map. Kagi: summarize. Defaults to provider default." - changed
Input schema / properties / mode / enumPrevious value: -[ - "extract", - "summarize", - "scrape", - "crawl", - "map", - "actions", - "contents", - "similar" -]New value: +[ + "extract", + "crawl", + "map", + "summarize", + "scrape", + "actions", + "contents", + "similar" +] - added
Input schema / properties / queryAdded value: +{ + "description": "Focus extracted content on information relevant to this query.", + "minLength": 1, + "pattern": "\\S", + "type": "string" +}
- Changed
web_search6 fields changed- added
Input schema / properties / auto_parametersAdded value: +{ + "description": "Let supported providers select search settings from the query. This can change cost.", + "type": "boolean" +} - added
Input schema / properties / include_raw_contentAdded value: +{ + "description": "Include full page content when the selected provider supports it.", + "type": "boolean" +} - added
Input schema / properties / safe_searchAdded value: +{ + "description": "Enable provider safe-search filtering.", + "type": "boolean" +} - added
Input schema / properties / search_depthAdded value: +{ + "description": "Search depth. Providers may use this to balance speed, relevance, and cost.", + "enum": [ + "basic", + "advanced", + "fast", + "ultra-fast" + ], + "type": "string" +} - added
Input schema / properties / time_rangeAdded value: +{ + "description": "Only return results from this recent time range.", + "enum": [ + "day", + "week", + "month", + "year" + ], + "type": "string" +} - added
Input schema / properties / topicAdded value: +{ + "description": "Search topic category.", + "enum": [ + "general", + "news", + "finance" + ], + "type": "string" +}
3 tool updates
v0.0.29- Changed
ai_search7 fields changed- added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - added
Input schema / properties / limit / maximumAdded value: +50 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - changed
Input schema / properties / query / descriptionPrevious value: -"Question or search query"New value: +"Search query" - added
Input schema / properties / query / minLengthAdded value: +1 - added
Input schema / properties / query / patternAdded value: +"\\S"
- Changed
web_extract3 fields changed- added
Input schema / properties / include_raw_contentsAdded value: +{ + "description": "Whether extraction responses should include per-URL raw_contents alongside combined content (default: true).", + "type": "boolean" +} - added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - changed
Input schema / properties / url / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "type": "string" - }, - "type": "array" - } -]New value: +[ + { + "format": "uri", + "pattern": "^https?:\\/\\/", + "type": "string" + }, + { + "items": { + "format": "uri", + "pattern": "^https?:\\/\\/", + "type": "string" + }, + "maxItems": 10, + "minItems": 1, + "type": "array" + } +]
- Changed
web_search10 fields changed- added
Input schema / properties / exclude_domains / items / patternAdded value: +"^(?:\\*\\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\\.)+[a-zA-Z]{2,63}$" - added
Input schema / properties / exclude_domains / maxItemsAdded value: +20 - added
Input schema / properties / include_domains / items / patternAdded value: +"^(?:\\*\\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\\.)+[a-zA-Z]{2,63}$" - added
Input schema / properties / include_domains / maxItemsAdded value: +20 - added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - added
Input schema / properties / limit / maximumAdded value: +50 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / query / minLengthAdded value: +1 - added
Input schema / properties / query / patternAdded value: +"\\S"
8 tool updates
v0.0.4- Changed
ai_search6 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Result limit"New value: +"Maximum number of results (default: 10)" - removed
Input schema / properties / provider / anyOfRemoved value: -[ - { - "const": "perplexity" - }, - { - "const": "kagi_fastgpt" - }, - { - "const": "exa_answer" - } -] - changed
Input schema / properties / provider / descriptionPrevious value: -"AI provider"New value: +"AI search provider to use" - added
Input schema / properties / provider / enumAdded value: +[ + "kagi_fastgpt" +] - added
Input schema / properties / provider / typeAdded value: +"string" - changed
Input schema / properties / query / descriptionPrevious value: -"Query"New value: +"Question or search query"
- Removed
firecrawl_process - Removed
jina_grounding_enhance - Removed
kagi_enrichment_enhance - Removed
kagi_summarizer_process - Removed
tavily_extract_process - Added
web_extract - Changed
web_search8 fields changed- changed
Input schema / properties / exclude_domains / descriptionPrevious value: -"Domains to exclude"New value: +"Exclude results from these domains" - changed
Input schema / properties / include_domains / descriptionPrevious value: -"Domains to include"New value: +"Only return results from these domains" - changed
Input schema / properties / limit / descriptionPrevious value: -"Result limit"New value: +"Maximum number of results (default: 10)" - removed
Input schema / properties / provider / anyOfRemoved value: -[ - { - "const": "tavily" - }, - { - "const": "brave" - }, - { - "const": "kagi" - }, - { - "const": "exa" - } -] - changed
Input schema / properties / provider / descriptionPrevious value: -"Search provider"New value: +"Search provider to use" - added
Input schema / properties / provider / enumAdded value: +[ + "tavily", + "brave", + "kagi", + "kagi_enrichment" +] - added
Input schema / properties / provider / typeAdded value: +"string" - changed
Input schema / properties / query / descriptionPrevious value: -"Query"New value: +"Search query"
18 tool updates
v1.0.0- Added
ai_search - Removed
brave_search - Removed
firecrawl_actions_process - Removed
firecrawl_crawl_process - Removed
firecrawl_extract_process - Removed
firecrawl_map_process - Added
firecrawl_process - Removed
firecrawl_scrape_process - Changed
jina_grounding_enhance2 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / content / descriptionPrevious value: -"Content to enhance"New value: +"Content"
- Removed
jina_reader_process - Changed
kagi_enrichment_enhance2 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / content / descriptionPrevious value: -"Content to enhance"New value: +"Content"
- Removed
kagi_fastgpt_search - Removed
kagi_search - Changed
kagi_summarizer_process9 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / extract_depth / anyOfAdded value: +[ + { + "const": "basic" + }, + { + "const": "advanced" + } +] - removed
Input schema / properties / extract_depth / defaultRemoved value: -"basic" - changed
Input schema / properties / extract_depth / descriptionPrevious value: -"The depth of the extraction process. \"advanced\" retrieves more data but costs more credits."New value: +"Extraction depth" - removed
Input schema / properties / extract_depth / enumRemoved value: -[ - "basic", - "advanced" -] - removed
Input schema / properties / extract_depth / typeRemoved value: -"string" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + } +] - added
Input schema / properties / url / descriptionAdded value: +"URL(s)" - removed
Input schema / properties / url / oneOfRemoved value: -[ - { - "description": "Single URL to process", - "type": "string" - }, - { - "description": "Multiple URLs to process", - "items": { - "type": "string" - }, - "type": "array" - } -]
- Removed
perplexity_search - Changed
tavily_extract_process9 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / extract_depth / anyOfAdded value: +[ + { + "const": "basic" + }, + { + "const": "advanced" + } +] - removed
Input schema / properties / extract_depth / defaultRemoved value: -"basic" - changed
Input schema / properties / extract_depth / descriptionPrevious value: -"The depth of the extraction process. \"advanced\" retrieves more data but costs more credits."New value: +"Extraction depth" - removed
Input schema / properties / extract_depth / enumRemoved value: -[ - "basic", - "advanced" -] - removed
Input schema / properties / extract_depth / typeRemoved value: -"string" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + } +] - added
Input schema / properties / url / descriptionAdded value: +"URL(s)" - removed
Input schema / properties / url / oneOfRemoved value: -[ - { - "description": "Single URL to process", - "type": "string" - }, - { - "description": "Multiple URLs to process", - "items": { - "type": "string" - }, - "type": "array" - } -]
- Removed
tavily_search - Added
web_search
15 tool updates
- First observed
brave_search - First observed
firecrawl_actions_process - First observed
firecrawl_crawl_process - First observed
firecrawl_extract_process - First observed
firecrawl_map_process - First observed
firecrawl_scrape_process - First observed
jina_grounding_enhance - First observed
jina_reader_process - First observed
kagi_enrichment_enhance - First observed
kagi_fastgpt_search - First observed
kagi_search - First observed
kagi_summarizer_process - First observed
perplexity_search - First observed
tavily_extract_process - First observed
tavily_search
TDQS
Scored across 3 tools
Each tool has a clearly distinct job: web_search returns raw search results, ai_search returns synthesized answers with citations, and web_extract processes specific URLs. There is no meaningful overlap or ambiguity between them.
All tool names follow a consistent snake_case pattern combining a domain prefix with an action: web_search, ai_search, web_extract. The naming style is uniform and predictable.
Three tools is well-scoped for an omnisearch server covering the core needs of searching, getting AI answers, and extracting web content. Each tool earns its place without redundancy.
The toolset covers the full search-to-insight workflow: finding sources, getting synthesized answers, and extracting or summarizing content from URLs. There are no obvious dead ends or missing core operations for the stated purpose.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Jina AI Reader/Search MCP — turn any URL into clean LLM-ready markdown, plus web search.
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Related MCP Servers
- AlicenseAqualityCmaintenanceA Model Context Protocol server that enables web search, scraping, crawling, and content extraction through multiple engines including SearXNG, Firecrawl, and Tavily.4338 npm143MIT
- -licenseNot gradedqualityNot gradedmaintenanceA unified Model Context Protocol server that integrates multiple search providers including Brave, Tavily, Exa, Semantic Scholar, and arXiv. It enables users to perform web, news, and image searches alongside academic research and citation analysis through a single interface.-
- AlicenseBqualityAmaintenanceA Model Context Protocol server that gives AI assistants access to 7 search providers with intelligent auto-routing. Analyzes query intent and picks the best provider automatically — no manual switching needed. Install, configure your keys, and go.2417 PyPI5MIT
- AlicenseNot gradedqualityDmaintenanceA unified MCP server aggregating 15 web search and extraction tools across 5 providers (Jina, Tavily, Exa, Firecrawl, Bocha) with automatic API key validation and plugin architecture.MIT