mcp-omnisearch
mcp-omnisearch
Un servidor de Protocolo de Contexto de Modelo (MCP) que proporciona acceso unificado a múltiples proveedores de búsqueda y herramientas de IA. Este servidor combina las capacidades de Tavily, Perplexity, Kagi, Jina AI, Brave y Firecrawl para ofrecer funciones integrales de búsqueda, respuestas de IA, procesamiento de contenido y mejora a través de una única interfaz.
Características
🔍 Herramientas de búsqueda
Búsqueda Tavily : Optimizada para obtener información objetiva con un sólido sistema de citas. Admite el filtrado de dominios mediante parámetros de la API (include_domains/exclude_domains).
Brave Search : Búsqueda centrada en la privacidad con una buena cobertura de contenido técnico. Ofrece compatibilidad nativa con operadores de búsqueda (site:, -site:, filetype:, intitle:, inurl:, before:, after: y frases exactas).
Kagi Search : Resultados de búsqueda de alta calidad con mínima influencia publicitaria, enfocados en fuentes confiables. Admite operadores de búsqueda en cadenas de consulta (site:, -site:, filetype:, intitle:, inurl:, before:, after: y frases exactas).
🎯 Operadores de búsqueda
MCP Omnisearch proporciona potentes capacidades de búsqueda a través de operadores y parámetros:
Funciones de búsqueda comunes
Filtrado de dominios: disponible en todos los proveedores
Tavily: A través de parámetros API (include_domains/exclude_domains)
Brave & Kagi: A través de los operadores site: y -site:
Filtrado de tipo de archivo: disponible en Brave y Kagi (tipo de archivo:)
Filtrado de títulos y URL: disponible en Brave y Kagi (intitle:, inurl:)
Filtrado de fechas: Disponible en Brave y Kagi (antes:, después:)
Coincidencia de frases exactas: Disponible en Brave y Kagi ("frase")
Ejemplo de uso
// Using Brave or Kagi with query string operators
{
"query": "filetype:pdf site:microsoft.com typescript guide"
}
// Using Tavily with API parameters
{
"query": "typescript guide",
"include_domains": ["microsoft.com"],
"exclude_domains": ["github.com"]
}Capacidades del proveedor
Brave Search : Compatibilidad total con operadores nativos en la cadena de consulta
Kagi Search : Soporte completo de operadores en la cadena de consulta
Tavily Search : Filtrado de dominios mediante parámetros API
🤖 Herramientas de respuesta de IA
Perplexity AI : Generación de respuestas avanzadas que combina la búsqueda web en tiempo real con GPT-4 Omni y Claude 3
Kagi FastGPT : Respuestas rápidas generadas por IA con citas (tiempo de respuesta típico: 900 ms)
📄 Herramientas de procesamiento de contenido
Jina AI Reader : Extracción de contenido limpio con subtítulos de imágenes y compatibilidad con PDF
Kagi Universal Summarizer : resumen de contenido para páginas, vídeos y podcasts
Tavily Extract : Extrae contenido sin procesar de una o varias páginas web con un nivel de extracción configurable (básico o avanzado). Devuelve contenido combinado y contenido de URL individuales, con metadatos que incluyen recuento de palabras y estadísticas de extracción.
Firecrawl Scrape : extrae datos limpios y listos para LLM de URL individuales con opciones de formato mejoradas
Firecrawl Crawl : rastreo profundo de todas las subpáginas accesibles en un sitio web con límites de profundidad configurables
Mapa de Firecrawl : recopilación rápida de URL de sitios web para un mapeo completo del sitio
Firecrawl Extract : Extracción de datos estructurados con IA mediante indicaciones en lenguaje natural
Acciones de Firecrawl : Compatibilidad con interacciones de página (clics, desplazamientos, etc.) antes de la extracción de contenido dinámico
🔄 Herramientas de mejora
API de enriquecimiento de Kagi : contenido complementario de índices especializados (Teclis, TinyGem)
Jina AI Grounding : Verificación de hechos en tiempo real con base en el conocimiento web
Related MCP server: MCP Search Server
Requisitos de clave API flexibles
MCP Omnisearch está diseñado para funcionar con las claves API disponibles. No es necesario tener claves para todos los proveedores: el servidor detectará automáticamente las claves API disponibles y solo las habilitará.
Por ejemplo:
Si solo tiene una clave API de Tavily y Perplexity, solo esos proveedores estarán disponibles
Si no tiene una clave API de Kagi, los servicios basados en Kagi no estarán disponibles, pero todos los demás proveedores funcionarán normalmente.
El servidor registrará qué proveedores están disponibles según las claves API que haya configurado
Esta flexibilidad permite comenzar fácilmente con uno o dos proveedores y agregar más según sea necesario.
Configuración
Este servidor requiere configuración a través de su cliente MCP. A continuación, se muestran ejemplos para diferentes entornos:
Configuración de Cline
Agregue esto a su configuración de Cline MCP:
{
"mcpServers": {
"mcp-omnisearch": {
"command": "node",
"args": ["/path/to/mcp-omnisearch/dist/index.js"],
"env": {
"TAVILY_API_KEY": "your-tavily-key",
"PERPLEXITY_API_KEY": "your-perplexity-key",
"KAGI_API_KEY": "your-kagi-key",
"JINA_AI_API_KEY": "your-jina-key",
"BRAVE_API_KEY": "your-brave-key",
"FIRECRAWL_API_KEY": "your-firecrawl-key"
},
"disabled": false,
"autoApprove": []
}
}
}Escritorio Claude con configuración WSL
Para entornos WSL, agregue esto a su configuración de Claude Desktop:
{
"mcpServers": {
"mcp-omnisearch": {
"command": "wsl.exe",
"args": [
"bash",
"-c",
"TAVILY_API_KEY=key1 PERPLEXITY_API_KEY=key2 KAGI_API_KEY=key3 JINA_AI_API_KEY=key4 BRAVE_API_KEY=key5 FIRECRAWL_API_KEY=key6 node /path/to/mcp-omnisearch/dist/index.js"
]
}
}
}Variables de entorno
El servidor utiliza claves API para cada proveedor. No necesita claves para todos los proveedores ; solo se activarán los que correspondan a sus claves API disponibles.
TAVILY_API_KEY: Para la búsqueda de TavilyPERPLEXITY_API_KEY: Para Perplexity AIKAGI_API_KEY: Para servicios Kagi (FastGPT, Summarizer, Enriquecimiento)JINA_AI_API_KEY: Para servicios de inteligencia artificial de Jina (Lector, Conexión a tierra)BRAVE_API_KEY: Para Brave SearchFIRECRAWL_API_KEY: Para servicios de Firecrawl (Rastreo, Mapeo, Extracción, Acciones)
Puedes empezar con una o dos claves API y añadir más según sea necesario. El servidor registrará los proveedores disponibles al iniciar.
API
El servidor implementa herramientas MCP organizadas por categoría:
Herramientas de búsqueda
búsqueda_tavily
Busca en la web con la API de búsqueda de Tavily. Ideal para consultas basadas en hechos que requieren fuentes y citas fiables.
Parámetros:
query(cadena, obligatoria): consulta de búsqueda
Ejemplo:
{
"query": "latest developments in quantum computing"
}búsqueda_valiente
Búsqueda web centrada en la privacidad con buena cobertura de temas técnicos.
Parámetros:
query(cadena, obligatoria): Consulta de búsqueda
Ejemplo:
{
"query": "rust programming language features"
}búsqueda_kagi
Resultados de búsqueda de alta calidad con mínima influencia publicitaria. Ideal para encontrar fuentes confiables y materiales de investigación.
Parámetros:
query(cadena, obligatoria): Consulta de búsquedalanguage(cadena, opcional): filtro de idioma (p. ej., "en")no_cache(booleano, opcional): omite la caché para resultados nuevos
Ejemplo:
{
"query": "latest research in machine learning",
"language": "en"
}Herramientas de respuesta de IA
ai_perplejidad
Generación de respuestas impulsada por IA con integración de búsqueda web en tiempo real.
Parámetros:
query(cadena, obligatoria): Pregunta o tema para la respuesta de IA
Ejemplo:
{
"query": "Explain the differences between REST and GraphQL"
}ai_kagi_fastgpt
Respuestas rápidas generadas por IA con citas.
Parámetros:
query(cadena, obligatoria): Pregunta para una respuesta rápida de IA
Ejemplo:
{
"query": "What are the main features of TypeScript?"
}Herramientas de procesamiento de contenido
lector de procesos de Jina
Convierta URL en texto limpio y compatible con LLM con subtítulos de imagen.
Parámetros:
url(cadena, obligatoria): URL a procesar
Ejemplo:
{
"url": "https://example.com/article"
}resumen_de_proceso_kagi
Resumir el contenido de las URL.
Parámetros:
url(cadena, obligatoria): URL para resumir
Ejemplo:
{
"url": "https://example.com/long-article"
}proceso_tavily_extracto
Extraiga contenido sin procesar de páginas web con Tavily Extract.
Parámetros:
url(cadena | cadena[], obligatorio): URL única o matriz de URL para extraer contenidoextract_depth(cadena, opcional): Profundidad de extracción: 'básica' (predeterminada) o 'avanzada'
Ejemplo:
{
"url": [
"https://example.com/article1",
"https://example.com/article2"
],
"extract_depth": "advanced"
}La respuesta incluye:
Contenido combinado de todas las URL
Contenido sin procesar individual para cada URL
Metadatos con recuento de palabras, extracciones exitosas y URL fallidas
proceso de raspado de firecrawl
Extraiga datos limpios y listos para LLM de URL individuales con opciones de formato mejoradas.
Parámetros:
url(cadena | cadena[], obligatorio): URL única o matriz de URL para extraer contenidoextract_depth(cadena, opcional): Profundidad de extracción: 'básica' (predeterminada) o 'avanzada'
Ejemplo:
{
"url": "https://example.com/article",
"extract_depth": "basic"
}La respuesta incluye:
Contenido limpio y con formato Markdown
Metadatos que incluyen título, recuento de palabras y estadísticas de extracción
proceso de rastreo de fuego
Rastreo profundo de todas las subpáginas accesibles en un sitio web con límites de profundidad configurables.
Parámetros:
url(cadena | cadena[], obligatorio): URL de inicio para el rastreoextract_depth(cadena, opcional): Profundidad de extracción: 'básica' (predeterminada) o 'avanzada' (controla la profundidad y los límites del rastreo)
Ejemplo:
{
"url": "https://example.com",
"extract_depth": "advanced"
}La respuesta incluye:
Contenido combinado de todas las páginas rastreadas
Contenido individual para cada página
Metadatos que incluyen título, recuento de palabras y estadísticas de rastreo
proceso del mapa de rastreo de fuego
Recopilación rápida de URL de sitios web para un mapeo completo del sitio.
Parámetros:
url(cadena | cadena[], obligatorio): URL a mapearextract_depth(cadena, opcional): Profundidad de extracción: 'básica' (predeterminada) o 'avanzada' (controla la profundidad del mapa)
Ejemplo:
{
"url": "https://example.com",
"extract_depth": "basic"
}La respuesta incluye:
Lista de todas las URL descubiertas
Metadatos que incluyen el título del sitio y el número de URL
proceso de extracción de firecrawl
Extracción de datos estructurados con IA utilizando indicaciones en lenguaje natural.
Parámetros:
url(cadena | cadena[], obligatorio): URL de la que extraer datos estructuradosextract_depth(cadena, opcional): Profundidad de extracción: 'básica' (predeterminada) o 'avanzada'
Ejemplo:
{
"url": "https://example.com",
"extract_depth": "basic"
}La respuesta incluye:
Datos estructurados extraídos de la página
Metadatos que incluyen título y estadísticas de extracción
proceso de acciones de firecrawl
Soporte para interacciones de página (clics, desplazamiento, etc.) antes de la extracción de contenido dinámico.
Parámetros:
url(cadena | cadena[], obligatorio): URL para interactuar y extraer contenido de ellaextract_depth(cadena, opcional): Profundidad de extracción: 'básica' (predeterminada) o 'avanzada' (controla la complejidad de las interacciones)
Ejemplo:
{
"url": "https://news.ycombinator.com",
"extract_depth": "basic"
}La respuesta incluye:
Contenido extraído después de realizar interacciones
Descripción de las acciones realizadas
Captura de pantalla de la página (si está disponible)
Metadatos que incluyen título y estadísticas de extracción
Herramientas de mejora
mejorar_el_enriquecimiento_kagi
Obtenga contenido complementario de índices especializados.
Parámetros:
query(cadena, obligatoria): Consulta para enriquecimiento
Ejemplo:
{
"query": "emerging web technologies"
}mejorar_jina_grounding
Verificar las afirmaciones frente al conocimiento web.
Parámetros:
statement(cadena, obligatoria): Declaración a verificar
Ejemplo:
{
"statement": "TypeScript adds static typing to JavaScript"
}Desarrollo
Configuración
Clonar el repositorio
Instalar dependencias:
pnpm installConstruir el proyecto:
pnpm run buildEjecutar en modo de desarrollo:
pnpm run devPublicación
Actualizar la versión en package.json
Construir el proyecto:
pnpm run buildPublicar en npm:
pnpm publishSolución de problemas
Claves API y acceso
Cada proveedor requiere su propia clave API y puede tener diferentes requisitos de acceso:
Tavily : Requiere una clave API de su portal para desarrolladores
Perplejidad : acceso a la API a través de su programa de desarrollador
Kagi : Algunas funciones están limitadas a los usuarios del plan Business (Team)
Jina AI : Se requiere una clave API para todos los servicios
Brave : clave API de su portal para desarrolladores
Firecrawl : Se requiere una clave API desde su portal para desarrolladores
Límites de velocidad
Cada proveedor tiene sus propios límites de velocidad. El servidor gestionará los errores de límite de velocidad correctamente y devolverá los mensajes de error correspondientes.
Contribuyendo
¡Agradecemos sus contribuciones! No dude en enviar una solicitud de incorporación de cambios.
Licencia
Licencia MIT: consulte el archivo LICENCIA para obtener más detalles.
Expresiones de gratitud
Construido sobre:
Available Tools
3 toolsai_searchGet AI-powered answers with citations and reasoning. Use when you need synthesized answers rather than raw search results. Providers: kagi_fastgpt (fast answers), exa_answer (semantic AI), linkup (deep agentic search), tavily_research (asynchronous multi-search reports; resubmit its research_id to retrieve results).BRead-onlyIdempotent
Get AI-powered answers with citations and reasoning. Use when you need synthesized answers rather than raw search results. Providers: kagi_fastgpt (fast answers), exa_answer (semantic AI), linkup (deep agentic search), tavily_research (asynchronous multi-search reports; resubmit its research_id to retrieve results).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | Yes | Search query | |
| provider | Yes | AI search provider to use | |
| research_id | No | Existing asynchronous research task ID to retrieve. Supported by Tavily Research. | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent safety, so the description adds value by disclosing asynchronous retrieval behavior for tavily_research ('resubmit its research_id'), provider-specific behaviors, and the output nature (citations and reasoning). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose front-loaded and provider details compactly listed. Every clause contributes meaning without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers when-to-use, provider differences, and the async resubmission pattern, and annotations cover safety. However, it omits response structure and contains a provider/enum inconsistency, leaving an agent with an ambiguous picture for a 5-parameter tool without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the description need not repeat parameter details, but it adds provider characteristics that conflict with the enum by naming exa_answer and linkup which are not valid values. It does usefully explain research_id for Tavily, but the misinformation undermines reliability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use when you need synthesized answers rather than raw search results', giving a clear when-to-use signal. However, it lists exa_answer and linkup as providers even though the schema enum only allows kagi_fastgpt and tavily_research, making the provider-selection guidance partially misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_extractExtract, process, or summarize web content from URLs. Use when you need to read page content, summarize articles, crawl sites, or extract structured data. Providers: tavily (content extraction), kagi (summarization of pages/videos/podcasts), firecrawl (scraping/crawling/mapping/structured extraction/interactive), exa (content retrieval/similar pages).BRead-onlyIdempotent
Extract, process, or summarize web content from URLs. Use when you need to read page content, summarize articles, crawl sites, or extract structured data. Providers: tavily (content extraction), kagi (summarization of pages/videos/podcasts), firecrawl (scraping/crawling/mapping/structured extraction/interactive), exa (content retrieval/similar pages).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL or array of URLs to process | |
| mode | No | Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract/crawl/map. Kagi: summarize. Defaults to provider default. | |
| query | No | Focus extracted content on information relevant to this query. | |
| format | No | Extracted page format (default: markdown). | |
| provider | Yes | Processing provider to use | |
| extract_depth | No | Extraction depth (default: basic) | |
| chunks_per_source | No | Maximum relevant content chunks per source when a query is provided. | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. | |
| include_raw_contents | No | Whether extraction responses should include per-URL raw_contents alongside combined content (default: true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds provider capability context (e.g., firecrawl for scraping/crawling/interactive, kagi for summarization). However, the mention of 'exa' as a provider is misleading because the input schema's provider enum omits exa, creating uncertainty about available behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loads the purpose, and packs useful provider information into a short list. Minor redundancy ('process' with 'extract') and the misleading exa reference are the only blemishes; overall it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with 5 enums and provider-specific modes, the description is too thin. It does not explain how to choose a provider for a given task, what the different modes do relative to providers, or how to handle edge cases like exa's absence from the provider enum. There is no output schema, and the description gives no hint about return value shapes, so an agent would likely need to infer provider-mode compatibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are fully documented in the schema. The description adds value by loosely mapping providers to capabilities (e.g., kagi for summarization, firecrawl for scraping), which helps select provider and mode. But it introduces a conflict by listing exa although the provider enum does not include it, and it does not explain the relationship between provider and mode beyond the parentheticals.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when...' conditions covering the main scenarios, and the provider list gives a starting point for mode selection. It does not state when not to use this tool or point to alternatives like web_search or ai_search, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchSearch the web for information. Use when you need to find web pages, articles, or data. Providers: tavily (factual/citations and search controls), brave (privacy/operators), kagi (quality/operators), exa (AI-semantic), kagi_enrichment (specialized indexes). Search depth, topic, time range, safe search, raw content, and automatic parameters apply when supported by the provider.BRead-onlyIdempotent
Search the web for information. Use when you need to find web pages, articles, or data. Providers: tavily (factual/citations and search controls), brave (privacy/operators), kagi (quality/operators), exa (AI-semantic), kagi_enrichment (specialized indexes). Search depth, topic, time range, safe search, raw content, and automatic parameters apply when supported by the provider.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (default: 10) | |
| query | Yes | Search query | |
| topic | No | Search topic category. | |
| provider | Yes | Search provider to use | |
| time_range | No | Only return results from this recent time range. | |
| safe_search | No | Enable provider safe-search filtering. | |
| search_depth | No | Search depth. Providers may use this to balance speed, relevance, and cost. | |
| auto_parameters | No | Let supported providers select search settings from the query. This can change cost. | |
| exclude_domains | No | Exclude results from these domains | |
| include_domains | No | Only return results from these domains | |
| large_result_mode | No | How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file. | |
| include_raw_content | No | Include full page content when the selected provider supports it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false), lowering the burden on the description. The description adds useful context that search depth, topic, time range, safe search, raw content, and auto_parameters behave conditionally 'when supported by the provider,' but it does not disclose result-format, citation, pagination, or cost behavior, which would add meaningful transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose appears in the first sentence, usage in the second, and the provider rundown is dense but informative. The title is a verbatim duplicate of the description, which is mildly redundant, but no other space is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, 5 enums, and no output schema, the description carries a heavy burden. It covers purpose, usage context, and provider-specific parameter behavior, but it does not describe the result format or return expectations, and the provider list conflicts with the schema enum. For such a configurable tool, these are meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description does add meta-information: several parameters (search_depth, topic, time_range, safe_search, raw_content, auto_parameters) are provider-dependent, which is not in the schema. However, the description lists 'exa' as a valid provider while the schema enum omits it, which could mislead an agent into sending an invalid provider value and offset the added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tautological: description restates name/title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-to-use directive ('Use when you need to find web pages, articles, or data') and adds provider-selection guidance by use case (e.g., tavily for factual/citations, brave for privacy/operators). It does not mention sibling alternatives or state when not to use this tool, stopping short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- Changed
ai_search2 fields changed- changed
Input schema / properties / provider / enumPrevious value: -[ - "kagi_fastgpt" -]New value: +[ + "kagi_fastgpt", + "tavily_research" +] - added
Input schema / properties / research_idAdded value: +{ + "description": "Existing asynchronous research task ID to retrieve. Supported by Tavily Research.", + "minLength": 1, + "type": "string" +}
- Changed
web_extract5 fields changed- added
Input schema / properties / chunks_per_sourceAdded value: +{ + "description": "Maximum relevant content chunks per source when a query is provided.", + "maximum": 5, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / formatAdded value: +{ + "description": "Extracted page format (default: markdown).", + "enum": [ + "markdown", + "text" + ], + "type": "string" +} - changed
Input schema / properties / mode / descriptionPrevious value: -"Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract. Kagi: summarize. Defaults to provider default."New value: +"Processing mode. Firecrawl: scrape/crawl/map/extract/actions. Exa: contents/similar. Tavily: extract/crawl/map. Kagi: summarize. Defaults to provider default." - changed
Input schema / properties / mode / enumPrevious value: -[ - "extract", - "summarize", - "scrape", - "crawl", - "map", - "actions", - "contents", - "similar" -]New value: +[ + "extract", + "crawl", + "map", + "summarize", + "scrape", + "actions", + "contents", + "similar" +] - added
Input schema / properties / queryAdded value: +{ + "description": "Focus extracted content on information relevant to this query.", + "minLength": 1, + "pattern": "\\S", + "type": "string" +}
- Changed
web_search6 fields changed- added
Input schema / properties / auto_parametersAdded value: +{ + "description": "Let supported providers select search settings from the query. This can change cost.", + "type": "boolean" +} - added
Input schema / properties / include_raw_contentAdded value: +{ + "description": "Include full page content when the selected provider supports it.", + "type": "boolean" +} - added
Input schema / properties / safe_searchAdded value: +{ + "description": "Enable provider safe-search filtering.", + "type": "boolean" +} - added
Input schema / properties / search_depthAdded value: +{ + "description": "Search depth. Providers may use this to balance speed, relevance, and cost.", + "enum": [ + "basic", + "advanced", + "fast", + "ultra-fast" + ], + "type": "string" +} - added
Input schema / properties / time_rangeAdded value: +{ + "description": "Only return results from this recent time range.", + "enum": [ + "day", + "week", + "month", + "year" + ], + "type": "string" +} - added
Input schema / properties / topicAdded value: +{ + "description": "Search topic category.", + "enum": [ + "general", + "news", + "finance" + ], + "type": "string" +}
3 tool updates
v0.0.29- Changed
ai_search7 fields changed- added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - added
Input schema / properties / limit / maximumAdded value: +50 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - changed
Input schema / properties / query / descriptionPrevious value: -"Question or search query"New value: +"Search query" - added
Input schema / properties / query / minLengthAdded value: +1 - added
Input schema / properties / query / patternAdded value: +"\\S"
- Changed
web_extract3 fields changed- added
Input schema / properties / include_raw_contentsAdded value: +{ + "description": "Whether extraction responses should include per-URL raw_contents alongside combined content (default: true).", + "type": "boolean" +} - added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - changed
Input schema / properties / url / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "items": { - "type": "string" - }, - "type": "array" - } -]New value: +[ + { + "format": "uri", + "pattern": "^https?:\\/\\/", + "type": "string" + }, + { + "items": { + "format": "uri", + "pattern": "^https?:\\/\\/", + "type": "string" + }, + "maxItems": 10, + "minItems": 1, + "type": "array" + } +]
- Changed
web_search10 fields changed- added
Input schema / properties / exclude_domains / items / patternAdded value: +"^(?:\\*\\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\\.)+[a-zA-Z]{2,63}$" - added
Input schema / properties / exclude_domains / maxItemsAdded value: +20 - added
Input schema / properties / include_domains / items / patternAdded value: +"^(?:\\*\\.)?(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?\\.)+[a-zA-Z]{2,63}$" - added
Input schema / properties / include_domains / maxItemsAdded value: +20 - added
Input schema / properties / large_result_modeAdded value: +{ + "description": "How to handle oversized responses for this request. Use inline for remote/container transports; file is local shared-filesystem behavior. Defaults to OMNISEARCH_LARGE_RESULT_MODE or file.", + "enum": [ + "inline", + "file" + ], + "type": "string" +} - added
Input schema / properties / limit / maximumAdded value: +50 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - added
Input schema / properties / query / minLengthAdded value: +1 - added
Input schema / properties / query / patternAdded value: +"\\S"
8 tool updates
v0.0.4- Changed
ai_search6 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Result limit"New value: +"Maximum number of results (default: 10)" - removed
Input schema / properties / provider / anyOfRemoved value: -[ - { - "const": "perplexity" - }, - { - "const": "kagi_fastgpt" - }, - { - "const": "exa_answer" - } -] - changed
Input schema / properties / provider / descriptionPrevious value: -"AI provider"New value: +"AI search provider to use" - added
Input schema / properties / provider / enumAdded value: +[ + "kagi_fastgpt" +] - added
Input schema / properties / provider / typeAdded value: +"string" - changed
Input schema / properties / query / descriptionPrevious value: -"Query"New value: +"Question or search query"
- Removed
firecrawl_process - Removed
jina_grounding_enhance - Removed
kagi_enrichment_enhance - Removed
kagi_summarizer_process - Removed
tavily_extract_process - Added
web_extract - Changed
web_search8 fields changed- changed
Input schema / properties / exclude_domains / descriptionPrevious value: -"Domains to exclude"New value: +"Exclude results from these domains" - changed
Input schema / properties / include_domains / descriptionPrevious value: -"Domains to include"New value: +"Only return results from these domains" - changed
Input schema / properties / limit / descriptionPrevious value: -"Result limit"New value: +"Maximum number of results (default: 10)" - removed
Input schema / properties / provider / anyOfRemoved value: -[ - { - "const": "tavily" - }, - { - "const": "brave" - }, - { - "const": "kagi" - }, - { - "const": "exa" - } -] - changed
Input schema / properties / provider / descriptionPrevious value: -"Search provider"New value: +"Search provider to use" - added
Input schema / properties / provider / enumAdded value: +[ + "tavily", + "brave", + "kagi", + "kagi_enrichment" +] - added
Input schema / properties / provider / typeAdded value: +"string" - changed
Input schema / properties / query / descriptionPrevious value: -"Query"New value: +"Search query"
18 tool updates
v1.0.0- Added
ai_search - Removed
brave_search - Removed
firecrawl_actions_process - Removed
firecrawl_crawl_process - Removed
firecrawl_extract_process - Removed
firecrawl_map_process - Added
firecrawl_process - Removed
firecrawl_scrape_process - Changed
jina_grounding_enhance2 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / content / descriptionPrevious value: -"Content to enhance"New value: +"Content"
- Removed
jina_reader_process - Changed
kagi_enrichment_enhance2 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / content / descriptionPrevious value: -"Content to enhance"New value: +"Content"
- Removed
kagi_fastgpt_search - Removed
kagi_search - Changed
kagi_summarizer_process9 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / extract_depth / anyOfAdded value: +[ + { + "const": "basic" + }, + { + "const": "advanced" + } +] - removed
Input schema / properties / extract_depth / defaultRemoved value: -"basic" - changed
Input schema / properties / extract_depth / descriptionPrevious value: -"The depth of the extraction process. \"advanced\" retrieves more data but costs more credits."New value: +"Extraction depth" - removed
Input schema / properties / extract_depth / enumRemoved value: -[ - "basic", - "advanced" -] - removed
Input schema / properties / extract_depth / typeRemoved value: -"string" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + } +] - added
Input schema / properties / url / descriptionAdded value: +"URL(s)" - removed
Input schema / properties / url / oneOfRemoved value: -[ - { - "description": "Single URL to process", - "type": "string" - }, - { - "description": "Multiple URLs to process", - "items": { - "type": "string" - }, - "type": "array" - } -]
- Removed
perplexity_search - Changed
tavily_extract_process9 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / extract_depth / anyOfAdded value: +[ + { + "const": "basic" + }, + { + "const": "advanced" + } +] - removed
Input schema / properties / extract_depth / defaultRemoved value: -"basic" - changed
Input schema / properties / extract_depth / descriptionPrevious value: -"The depth of the extraction process. \"advanced\" retrieves more data but costs more credits."New value: +"Extraction depth" - removed
Input schema / properties / extract_depth / enumRemoved value: -[ - "basic", - "advanced" -] - removed
Input schema / properties / extract_depth / typeRemoved value: -"string" - added
Input schema / properties / url / anyOfAdded value: +[ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + } +] - added
Input schema / properties / url / descriptionAdded value: +"URL(s)" - removed
Input schema / properties / url / oneOfRemoved value: -[ - { - "description": "Single URL to process", - "type": "string" - }, - { - "description": "Multiple URLs to process", - "items": { - "type": "string" - }, - "type": "array" - } -]
- Removed
tavily_search - Added
web_search
15 tool updates
- First observed
brave_search - First observed
firecrawl_actions_process - First observed
firecrawl_crawl_process - First observed
firecrawl_extract_process - First observed
firecrawl_map_process - First observed
firecrawl_scrape_process - First observed
jina_grounding_enhance - First observed
jina_reader_process - First observed
kagi_enrichment_enhance - First observed
kagi_fastgpt_search - First observed
kagi_search - First observed
kagi_summarizer_process - First observed
perplexity_search - First observed
tavily_extract_process - First observed
tavily_search
TDQS
Scored across 3 tools
Each tool has a clearly distinct job: web_search returns raw search results, ai_search returns synthesized answers with citations, and web_extract processes specific URLs. There is no meaningful overlap or ambiguity between them.
All tool names follow a consistent snake_case pattern combining a domain prefix with an action: web_search, ai_search, web_extract. The naming style is uniform and predictable.
Three tools is well-scoped for an omnisearch server covering the core needs of searching, getting AI answers, and extracting web content. Each tool earns its place without redundancy.
The toolset covers the full search-to-insight workflow: finding sources, getting synthesized answers, and extracting or summarizing content from URLs. There are no obvious dead ends or missing core operations for the stated purpose.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Jina AI Reader/Search MCP — turn any URL into clean LLM-ready markdown, plus web search.
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Related MCP Servers
- AlicenseAqualityCmaintenanceA Model Context Protocol server that enables web search, scraping, crawling, and content extraction through multiple engines including SearXNG, Firecrawl, and Tavily.4338 npm143MIT
- -licenseNot gradedqualityNot gradedmaintenanceA unified Model Context Protocol server that integrates multiple search providers including Brave, Tavily, Exa, Semantic Scholar, and arXiv. It enables users to perform web, news, and image searches alongside academic research and citation analysis through a single interface.-
- AlicenseBqualityAmaintenanceA Model Context Protocol server that gives AI assistants access to 7 search providers with intelligent auto-routing. Analyzes query intent and picks the best provider automatically — no manual switching needed. Install, configure your keys, and go.2417 PyPI5MIT
- AlicenseNot gradedqualityDmaintenanceA unified MCP server aggregating 15 web search and extraction tools across 5 providers (Jina, Tavily, Exa, Firecrawl, Bocha) with automatic API key validation and plugin architecture.MIT