qdrant-mcp-ollama
qdrant-mcp-ollama
Un servidor de Model Context Protocol (MCP) para la base de datos vectorial de Qdrant que utiliza Ollama para embeddings acelerados por GPU.
¿Por qué no el mcp-server-qdrant oficial?
El servidor MCP oficial de Qdrant usa FastEmbed para los embeddings, lo que:
Solo funciona en CPU — es lento en bases de código grandes y no aprovecha las GPU modernas
Utiliza un modelo pequeño (
all-MiniLM-L6-v2, 384 dimensiones) — embeddings de menor calidadProtección de un solo proceso en modo local — solo un cliente MCP puede acceder a la base de datos a la vez
Este servidor resuelve los tres problemas:
|
| |
Motor de embeddings | FastEmbed (CPU) | Ollama (GPU) |
Modelo predeterminado | all-MiniLM-L6-v2 (384-dim, 80 MB) | bge-m3 (1024-dim, 1,2 GB) |
Acceso concurrente | No (modo local) | Sí (servidor Qdrant) |
Flexibilidad de modelos | Solo modelos FastEmbed | Cualquier modelo de embeddings de Ollama |
Related MCP server: Claude Context MCP
Arquitectura
┌──────────────┐ ┌────────────────────┐ ┌─────────────┐
│ MCP Client │────>│ qdrant-mcp-ollama │────>│ Ollama │
│ (Claude Code, │ │ (server.py) │ │ (GPU) │
│ Kilo Code, │<────│ │ └─────────────┘
│ Cursor, etc) │ └────────┬───────────┘
└──────────────┘ │
v
┌────────────────────┐
│ Qdrant Server │
│ (Docker, :6333) │
│ Storage: local │
│ disk / cloud │
└────────────────────┘Requisitos previos
Ollama — instalado y en ejecución, con un modelo de embeddings descargado mediante
ollama pullDocker — para ejecutar el servidor de Qdrant
uv — gestor de paquetes de Python (recomendado) o
pip
Inicio rápido
1. Descarga un modelo de embeddings en Ollama
ollama pull bge-m32. Inicia el servidor de Qdrant
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v qdrant-storage:/qdrant/storage \
--restart unless-stopped \
qdrant/qdrant:latest3. Ejecuta el servidor MCP
# No install needed — uv downloads dependencies on-the-fly:
QDRANT_URL="http://localhost:6333" \
EMBEDDING_MODEL="bge-m3" \
uv run --with fastmcp --with qdrant-client --with httpx python server.py4. Indexa una base de código
uv run --with qdrant-client --with httpx python embed_codebase.py \
/path/to/your/project my-project --preset python5. Busca desde tu cliente MCP
Una vez configurado (consulta las secciones de abajo), pregúntale a tu asistente de IA:
"Busca la lógica de autenticación en la base de código"
Utilizará la herramienta qdrant_find para devolver fragmentos de código semánticamente relevantes.
Configurar el servidor de Qdrant
Opción A: Docker (recomendado)
Almacena los datos en una unidad concreta (p. ej., E: en Windows):
# Create storage directories
mkdir -p E:/qdrant-storage E:/qdrant-snapshots
# Start Qdrant with persistent storage
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v E:/qdrant-storage:/qdrant/storage \
-v E:/qdrant-snapshots:/qdrant/snapshots \
--restart unless-stopped \
qdrant/qdrant:latestEn Linux/macOS:
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v ~/qdrant-storage:/qdrant/storage \
--restart unless-stopped \
qdrant/qdrant:latestLa opción --restart unless-stopped hace que Qdrant se inicie automáticamente con Docker Desktop.
Comprobación de funcionamiento:
docker ps --filter name=qdrant-server
# Or open http://localhost:6333/dashboard in your browserOpción B: Qdrant Cloud
Regístrate en cloud.qdrant.io y obtén tu URL y clave de API. Después define:
QDRANT_URL="https://your-cluster.cloud.qdrant.io:6333"
QDRANT_API_KEY="your-api-key"Nota: la variable de entorno
QDRANT_API_KEYse transfiere automáticamente al cliente de Qdrant.
Indexar una base de código
El script embed_codebase.py escanea un directorio, divide los archivos fuente en fragmentos y los indexa por lotes en Qdrant utilizando Ollama y la GPU.
Uso básico
uv run --with qdrant-client --with httpx python embed_codebase.py <directory> <collection-name>Uso de presets de extensiones
# Python project
python embed_codebase.py ./my-api api-backend --preset python
# Full-stack web project
python embed_codebase.py ./my-app frontend --preset web
# R / bioinformatics project
python embed_codebase.py ./analysis bio-analysis --preset r
# Everything
python embed_codebase.py ./mono-repo all-code --preset allExtensiones personalizadas
python embed_codebase.py ./project my-collection --extensions .py .sql .sh .yamlPresets disponibles
Preset | Extensiones |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| Todas las extensiones de código más comunes |
Si no se proporciona --preset ni --extensions, el script detecta automáticamente los tipos de archivo.
Todas las opciones
usage: embed_codebase.py <directory> <collection> [options]
positional arguments:
directory Path to the codebase directory
collection Qdrant collection name
options:
--extensions EXT [EXT ...] File extensions to include (e.g. .py .ts)
--preset PRESET Use a preset group of extensions
--model MODEL Ollama embedding model (default: bge-m3)
--qdrant-url URL Qdrant server URL (default: http://localhost:6333)
--ollama-url URL Ollama server URL (default: http://localhost:11434)
--chunk-size N Max lines per chunk (default: 80)
--chunk-overlap N Overlap lines between chunks (default: 10)
--batch-size N Upload batch size for Qdrant (default: 500)
--append Append to existing collection instead of replacingModo append
Por defecto, al volver a ejecutar el script se reemplaza la colección. Usa --append para añadir a una colección existente:
# First embed
python embed_codebase.py ./src main-code --preset typescript
# Add more files later
python embed_codebase.py ./docs main-code --extensions .md --appendUso con varias bases de código
Utiliza colecciones separadas para cada base de código y consigue que los resultados de búsqueda estén delimitados y sean relevantes:
# Project A
python embed_codebase.py ~/projects/api-server api-server --preset python
# Project B
python embed_codebase.py ~/projects/web-app web-app --preset web
# Project C
python embed_codebase.py ~/projects/data-pipeline data-pipeline --preset pythonAl configurar el servidor MCP:
Sin
COLLECTION_NAME: de versión por consulta. Es ideal cuando un mismo servidor MCP da servicio a varios proyectos.Con
COLLECTION_NAME: se usa automáticamente una colección por defecto. Configúralo por proyecto si tu cliente MCP admite la configuración por proyecto.
Configurar Claude Code
Añade el servidor MCP
claude mcp add qdrant -s user \
-e QDRANT_URL="http://localhost:6333" \
-e OLLAMA_URL="http://localhost:11434" \
-e EMBEDDING_MODEL="bge-m3" \
-- uv run --with fastmcp --with qdrant-client --with httpx \
python /path/to/qdrant-mcp-ollama/server.pyReemplaza /path/to/qdrant-mcp-ollama/ por la ruta real donde clonaste este repositorio.
Con una colección por defecto
Si trabajas principalmente en un solo proyecto:
claude mcp add qdrant -s user \
-e QDRANT_URL="http://localhost:6333" \
-e OLLAMA_URL="http://localhost:11434" \
-e EMBEDDING_MODEL="bge-m3" \
-e COLLECTION_NAME="my-project" \
-- uv run --with fastmcp --with qdrant-client --with httpx \
python /path/to/qdrant-mcp-ollama/server.pyVerificación
claude mcp list
# Should show: qdrant: ... ✓ Connected
claude mcp get qdrant
# Shows full configuration detailsUso en Claude Code
Cuando esté configurado, Claude Code puede usar estas herramientas:
qdrant_store— Guarda información: "Almacena este patrón de autenticación en Qdrant"qdrant_find— Busca: "Encuentra código relacionado con migraciones de base de datos"
En configuraciones con múltiples colecciones (sin colección por defecto), especifica la colección:
"Busca la lógica de limitación de peticiones en la colección
api-server"
Configurar Kilo Code (extensión de VS Code)
Kilo Code es una extensión de VS Code con soporte MCP integrado.
Opción 1: Configuración manual de MCP
Abre la configuración de Kilo Code en VS Code.
Ve a la configuración de servidores MCP.
Añade un nuevo servidor con:
Campo | Valor |
Nombre |
|
Comando |
|
Argumentos |
|
Define las variables de entorno:
Variable | Valor |
|
|
|
|
|
|
| El nombre de la colección de tu proyecto (p. ej., |
Opción 2: settings.json de VS Code
Añade a tu settings.json de VS Code (Ctrl+Shift+P > Preferencias: Abrir Configuración de Usuario (JSON)):
{
"kilocode.mcpServers": {
"qdrant": {
"command": "uv",
"args": [
"run", "--with", "fastmcp", "--with", "qdrant-client", "--with", "httpx",
"python", "/path/to/qdrant-mcp-ollama/server.py"
],
"env": {
"QDRANT_URL": "http://localhost:6333",
"OLLAMA_URL": "http://localhost:11434",
"EMBEDDING_MODEL": "bge-m3",
"COLLECTION_NAME": "my-project"
}
}
}
}Configuración por proyecto en Kilo Code
Para configuraciones con múltiples bases de código, configura Kilo Code en el ámbito de proyecto (no global) con un COLLECTION_NAME específico de cada proyecto. Así, cada espacio de trabajo buscará únicamente en su propia base de código.
Configurar otros clientes MCP
Cursor / Windsurf
Ejecuta el servidor con transporte SSE para clientes con capacidad de conexión remota:
QDRANT_URL="http://localhost:6333" \
OLLAMA_URL="http://localhost:11434" \
EMBEDDING_MODEL="bge-m3" \
FASTMCP_PORT=8000 \
uv run --with fastmcp --with qdrant-client --with httpx \
python server.py --transport sseLuego, en la configuración de MCP de Cursor/Windsurf, conéctate a: http://localhost:8000/sse
Cliente MCP genérico (stdio)
El transporte predeterminado es stdio. Cualquier cliente MCP que admita stdio puede usar este servidor ejecutando:
uv run --with fastmcp --with qdrant-client --with httpx python server.pyReferencia de configuración
Variables de entorno del servidor MCP
Variable | Descripción | Predeterminado |
| URL del servidor de Qdrant |
|
| Clave de API para Qdrant Cloud | Ninguna |
| URL del servidor de Ollama |
|
| Nombre del modelo de embeddings de Ollama |
|
| Colección por defecto (vacío = debe especificarse por llamada) | *(vacío)* |
Elección del modelo de embeddings
Todos los modelos siguientes se descargan con ollama pull <model>:
Modelo | Dimensiones | Tamaño | Velocidad | Calidad | Mejor para |
| 1024 | 1,2 GB | Moderada | Alta | Uso general, multilingüe |
| 768 | 274 MB | Rápida | Buena | Ligero, para español e inglés |
| 1024 | 670 MB | Moderada | Alta | Inglés, alta calidad |
| 1024 | 1,2 GB | Moderada | Muy alta | Mejor calidad, inglés |
| 384 | 46 MB | Muy rápida | Aceptable | Recursos mínimos |
Recomendación: empieza con bge-m3. Se comporta bien con código, admite contenido multilingüe (comentarios en cualquier idioma) y equilibra calidad y velocidad.
Importante: el modelo de embeddings usado para indexar una colección debe coincidir con el modelo usado en las consultas. Si vuelves a indexar con un modelo diferente, elimina y vuelve a crear la colección.
Uso de GPU
Los modelos más grandes utilizan más GPU. Si tu GPU está infrautilizada:
Cambia de
nomic-embed-text(274 MB) abge-m3(1,2 GB) o a un modelo mayor.El script de indexación envía todos los textos en un solo lote para maximizar la saturación de la GPU.
En las consultas individuales (a través de
qdrant_find), los picos de GPU son breves y normales; indexar una sola consulta toma milisegundos.
Comprueba el uso de GPU: nvidia-smi (NVIDIA) o rocm-smi (AMD)
Herramientas MCP
qdrant_store
Almacena información en la base de datos de Qdrant.
Parámetro | Tipo | Obligatorio | Descripción |
| string | Sí | Texto que se debe almacenar y poder buscar |
| string | Si no hay valor por defecto | Colección de destino |
| dict | No | Metadatos opcionales que adjuntar |
qdrant_find
Busca información relevante mediante similitud semántica.
Parámetro | Tipo | Obligatorio | Descripción |
| string | Yes | Consulta de búsqueda en lenguaje natural |
| string | Si no hay valor por defecto | Colección en la que buscar |
| int | No | Número máximo de resultados por defecto (por defecto: 5) |
Solución de problemas
«Connection closed» / el servidor MCP no se inicia
¿Está Ollama ejecutándose? Compruébalo con
ollama list. Si es necesario, inícialo conollama serve.¿Se ha descargado el modelo de embeddings? Ejecuta
ollama pull bge-m3.¿Está Qdrant ejecutuado? Compruébalo con
docker ps --filter name=qdrant-server.
«La colección no existe»
La colección se crea al ejecutar el script de indexación o en la primera llamada a qdrant_store. Haz una de estas dos cosas:
Ejecuta
embed_codebase.pypara indexar tu base de código en primer lugar.O guarda algo con
qdrant_storepara que la colección se cree automáticamente.
Errores de dimensiones
Esto sucede cuando la colección se creó con un modelo de embeddings pero se consulta con otro. Solución:
Elimina la colección: visita
http://localhost:6333/dashboard.Vuelve a indexar con el modelo correcto.
Asegúrate de que
EMBEDDING_MODELen la configuración del servidor MCP coincide con el que usaste para indexar.
«La carpeta de almacenamiento ya está en uso por otra instancia»
Este mensaje es del oficial mcp-server-qdrant cuando se utiliza en modo local (QDRANT_LOCAL_PATH). Este proyecto evita ese problema conectando a un servidor de Qdrant mediante URL. Asegúrate de no ejecutar ambos servidores apuntando a la misma ruta local.
Indexación lenta / GPU
Usa modelos más grandes y verifica con nvidia-smi (NVIDIA) o rocm-smi (AMD).
Usa un modelo más grande:
bge-m3(1.2 GB) en lugar denomic-embed-text(274 MB)El script de incrustación envía todos los textos en un solo lote — si tienes miles de fragmentos, esto maximiza el uso de la GPU
Para bases de código muy grandes (10,000+ archivos), considera dividir en varias ejecuciones por directorio
Licencia
Apache License 2.0 — ver LICENSE.
Available Tools
2 toolsqdrant_findC
Search for relevant information in the Qdrant database using semantic similarity.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Natural language query to search for. The query is embedded using the same GPU model used for storage, ensuring accurate results. | |
| top_k | No | Maximum number of results to return (default: 5). | |
| collection_name | No | Name of the collection to search in. Required if no default collection is configured. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read-only operation but does not explicitly state that no data is modified, does not mention return behavior, error conditions, or limitations. The single sentence provides minimal behavioral disclosure beyond the literal action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no redundancy or filler. It is front-loaded with the verb and resource. While extremely brief, it is not a tautology and conveys the essential purpose. It avoids unnecessary words while being clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a sibling (qdrant_store) and an output schema (which covers return format), the description is still incomplete. It lacks any usage context, such as when to choose this over storage or how the search integrates with the workflow. The presence of an output schema reduces the need to explain returns, but the description does not cover the selection decision or behavioral expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all three parameters (query, top_k, collection_name) with descriptions, so the baseline is 3. The description adds nothing beyond the schema; it mentions 'semantic similarity' which is already implied by the query parameter's embedding mention. No additional value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search'), the target resource ('Qdrant database'), and the method ('semantic similarity'). This distinguishes it from the sibling qdrant_store, which likely stores information. However, it does not explicitly name the sibling or contrast with it, so it lacks the full differentiation seen in higher-scoring examples.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the alternative qdrant_store, nor any mention of prerequisites or context. The description only states the action without any direction on selection or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
qdrant_storeC
Store information in the Qdrant database with GPU-accelerated embeddings.
| Name | Required | Description | Default |
|---|---|---|---|
| metadata | No | Optional metadata dictionary to attach to the stored point. | |
| information | Yes | The text information to store. This will be embedded and made searchable via semantic similarity. | |
| collection_name | No | Name of the collection to store in. Required if no default collection is configured. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states that information will be stored with embeddings, but does not disclose potential side effects such as whether existing points are overwritten, whether collections are auto-created, or any error behavior. The mutation is implied but not explicitly flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that conveys the core action. 'GPU-accelerated embeddings' adds a performance detail that may be useful context, but it could be considered extraneous. Overall, it is appropriately sized and front-loaded with the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple store operation with only 3 parameters and an output schema present, the description covers the basic action. However, it omits guidance on when a collection_name is required and does not mention any setup steps or constraints. It meets a minimum viable level but leaves gaps that an agent might need to handle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific detail beyond what the schema already provides. The only minor addition is implying that information gets embedded, which is already stated in the schema. This meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Store information in the Qdrant database') with a specific resource and purpose. It implies a write operation distinct from the sibling qdrant_find, though it doesn't explicitly differentiate. The mention of 'GPU-accelerated embeddings' adds implementation detail but doesn't obscure the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the sibling qdrant_find. The description does not say 'use this to add data, use qdrant_find to search' or mention any prerequisites like collection existence. An agent would have to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
qdrant_find - First observed
qdrant_store
TDQS
Scored across 2 tools
The two tools, qdrant_store and qdrant_find, have entirely distinct purposes—one writes data, the other retrieves it. There is zero ambiguity between them.
Both tools follow a consistent 'qdrant_<verb>' pattern, using clear action verbs (store, find). The naming is predictable and uniform.
With only two tools, the server feels thin for what is typically a database domain, but it is not an extreme mismatch. It sits at the borderline of adequacy.
The server only provides store and find, lacking any management operations like delete, update, or list. For a database, this is a significant gap that will limit workflow coverage.
Maintenance
Related MCP Connectors
Connect AI assistants to your GitHub-hosted Obsidian vault to seamlessly access, search, and analy…
Search your knowledge bases from any AI assistant using hybrid RAG.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables semantic code search across codebases using Qdrant vector database and OpenAI embeddings, allowing users to find code by meaning rather than just keywords through natural language queries.2MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to index and search codebases using semantic search powered by multiple embedding providers (OpenAI, VoyageAI, Gemini, Ollama) and vector database storage.-
- FlicenseNot gradedqualityDmaintenanceEnables semantic code search across multi-language codebases using natural language queries, integrated with Qdrant vector database for fast, cached retrieval.1-
- AlicenseNot gradedqualityFmaintenanceIndexes codebases into Qdrant for semantic search, enabling AI assistants to find relevant code by meaning without re-exploring the repo.MIT