qdrant-mcp-ollama
qdrant-mcp-ollama
Ein Model Context Protocol (MCP)-Server für die Qdrant-Vektordatenbank, der Ollama für GPU-beschleunigte Embeddings verwendet.
Warum nicht das offizielle mcp-server-qdrant?
Der offizielle Qdrant-MCP-Server verwendet FastEmbed für Embeddings, was:
Nur auf der CPU läuft – langsam bei großen Codebasen, moderne GPUs werden nicht ausgenutzt
Ein anderes Modell verwendet (
all-MiniLM-L6-v2, 384-dim) – leistungsfähigere Embeddings sind nicht möglichIm lokalen Modus eine Einzelprozess-/Single-Process-Sperre – nur ein MCP-Client kann gleichzeitig auf die Datenbank zugreifen
Dieser Server behebt alle drei Probleme:
Offizielles |
| |
Embedding-Engine | FastMobile (CPU) | Ollama (GPU) |
Standardmodell | all-MiniLM-L6-v2 (384-dim, 80MB) | bge-m3 (1024-dim, 1.2GB) |
Paralleler Zugriff | Nein (lokaler Modus) | Ja (Qdrant-Server) |
Modellflexibilität | Nur Fassung - nur FastEmbed-Modelle | Beliebiges Ollama-Embedding-Modell |
Related MCP server: Claude Context MCP
Architektur
┌──────────────┐ ┌────────────────────┐ ┌─────────────┐
│ MCP Client │────>│ qdrant-mcp-ollama │────>│ Ollama │
│ (Claude Code, │ │ (server.py) │ │ (GPU) │
│ Kilo Code, │<────│ │ └─────────────┘
│ Cursor, etc) │ └────────┬───────────┘
└──────────────┘ │
v
┌────────────────────┐
│ Qdrant Server │
│ (Docker, :6333) │
│ Storage: local │
│ disk / cloud │
└────────────────────┘Voraussetzungen
Ollama – installiert und läuft, ein Embedding-Modell ist heruntergeladen
Docker – muss eingerichtet (läuft), um den Qdrant-Server zu betreiben
uv – Python-Paketmanager (empfohlen) oder
pip
Schnellstart
1. Ein Ollama-Embedding-Modell herunterladen
ollama pull bge-m32. Qdrant-Server starten
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v qdrant-storage:/qdrant/storage \
--restart unless-stopped \
qdrant/qdrant:latest3. MCP-Start up
# No install needed — uv downloads dependencies on-the-fly:
QDRANT_URL="http://localhost:6333" \
EMBEDDING_MODEL="bge-m3" \
uv run --with fastmcp --with qdrant-client --with httpx python server.py4. Codebasis einbetten
uv run --with qdrant-client --with httpx python embed_codebase.py \
/path/to/your/project my-project --preset python5. In deinem MCP-Client suchen
Nach der Konfiguration (siehe Abschnitte unten) frage deinen KI-Assistenten:
„Durchsuche die Codebasis nach Authentifizierungslogik.
Er nutzt das qdrant_' 'формулирует" – actually I'm about to write "Er nutzt das qdrant_find-Tool, um semantisch relevante Codeanteile zurückzugeben."
So:
„Durchsuche die Codebasis nach Authentifizierungslogik."
S**Fnd-Tool liefert dann semantisch relevante Code-Ausschnitte zurück.
Qdrant-Server einrichten
Option A: Docker (empfohlen)
Speichere Daten auf einem bestimmten Laufwerk (z. B. E: unter Windows):
# Create storage directories
mkdir -p E:/qdrant-storage E:/qdrant-snapshots
# Start Qdrant with persistent storage
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v E:/qdrant-storage:/qdrant/storage \
-v E:/qdrant-snapshots:/qdrant/snapshots \
--restart unless-stopped \
qdrant/qdrant:latestUnter Linux/macOS:
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v ~/qdrant-storage:/qdrant/storage \
--restart unless-stopped \
qdrant/qdrant:latestDas restart unless-stopped-Flag sorgt dafür, dass Qdrant automatisch mit Docker Desktop startet.
Überprüfen, dass Qdrant läuft:
docker ps --filter name=qdrant-server
# Or open http://localhost:6333/dashboard in your browserOption B: Qdrant Cloud
Registrieren bei cloud.qdrant.io und URL sowie API-Schlüssel notieren. Dann setzen:
QDRANT_URL="https://your-cluster.cloud.qdrant.io:6333"
QDRANT_API_KEY="your-api-key"Hinweis: Die Umgebungsvariable
QDRANT_API_KEYwird automatisch an den Qdrant-Client durchgereicht.
Codebasis einbetten
Das Skript embed_codebase.py durchsucht ein Verzeichnis, zerlegt Quelldateien in Chunks und bettet sie über Ollama auf der GPU in Qdrant ein.
Grundlegende Verwendung
uv run --with qdrant-client --with httpx python embed_codebase.py <directory> <collection-name>Mit Erweiterungs-Presets
# Python project
python embed_codebase.py ./my-api api-backend --preset python
# Full-stack web project
python embed_codebase.py ./my-app frontend --preset web
# R / bioinformatics project
python embed_codebase.py ./analysis bio-analysis --preset r
# Everything
python embed_codebase.py ./mono-repo all-code --preset allEigene Datei-Erweiterungen
python embed_codebase.py ./project my-collection --extensions .py .sql .sh .yamlVerfügbare Presets
Preset | Datei-Erweiterungen |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| Alle gängigen Quelldatei-Erweiterungen |
Werden weder --preset noch --extensions angegeben, erkennt das Skript die Dateitypen automatisch.
Alle Optionen
usage: embed_codebase.py <directory> <collection> [options]
positional arguments:
directory Path to the codebase directory
collection Qdrant collection name
options:
--extensions EXT [EXT ...] File extensions to include (e.g. .py .ts)
--preset PRESET Use a preset group of extensions
--model MODEL Ollama embedding model (default: bge-m3)
--qdrant-url URL Qdrant server URL (default: http://localhost:6333)
--ollama-url URL Ollama server URL (default: http://localhost:11434)
--chunk-size N Max lines per chunk (default: 80)
--chunk-overlap N Overlap lines between chunks (default: 10)
--batch-size N Upload batch size for Qdrant (default: 500)
--append Append to existing collection instead of replacingAppend-Modus
Standardmäßig ersetzt ein erneuter Skriptlauf die Collection. Mit --append wird an eine bestehende Collection angehängt:
# First embed
python embed_codebase.py ./src main-code --preset typescript
# Add more files later
python embed_codebase.py ./docs main-code --extensions .md --appendMehrere Codebasen verwenden
Verwende getrennte Collections für jede Codebasis, um die Suchergebnisse klar umgrenzt und relevant zu halten:
# Project A
python embed_codebase.py ~/projects/api-server api-server --preset python
# Project B
python embed_codebase.py ~/projects/web-app web-app --preset web
# Project C
python embed_codebase.py ~/projects/data-pipeline data-pipeline --preset pythonBei der Konfiguration des MCP-Servers:
Ohne
COLLECTION_NAME: Du musst die Collection pro Anfrage angeben. Das ist ideal, wenn ein MCP-Server mehrere Projekte bedient.Mit
COLLECTION_NAME: Eine Standard-Collection wird automatisch verwendet. Setze diesen Wert pro Projekt, wenn dein MCP-Client eine Projekt-spezifische Konfiguration unterstützt.
Claude Code konfigurieren
MCP-Server hinzufügen
claude mcp add qdrant -s user \
-e QDRANT_URL="http://localhost:6333" \
-e OLLAMA_URL="http://localhost:11434" \
-e EMBEDDING_MODEL="bge-m3" \
-- uv run --with fastmcp --with qdrant-client --with httpx \
python /path/to/qdrant-mcp-ollama/server.pyErsetze /path/to/qdrant-mcp-ollama/ durch den tatsächlichen Pfad, in den du dieses Repository geklont hast.
Mit einem Standard-Collectionnamen
Wenn du hauptsächlich an einem Projekt arbeitest:
claude mcp add qdrant -s user \
-e QDRANT_URL="http://localhost:6333" \
-e OLLAMA_URL="http://localhost:11434" \
-e EMBEDDING_MODEL="bge-m3" \
-e COLLECTION_NAME="my-project" \
-- uv run --with fastmcp --with qdrant-client --with httpx \
python /path/to/qdrant-mcp-ollama/server.pyVerifizieren
claude mcp list
# Should show: qdrant: ... ✓ Connected
claude mcp get qdrant
# Shows full configuration detailsVerwendung in Claude Code
Nach der Konfiguration kann Claude Code diese Tools nutzen:
qdrant_store– Informationen speichern: „Speichere dieses Authentifizierungsmuster in Qdrant“qdrant_find– Suchen: „Finde Code zu Datenbank-kautionen“
Für Setups mit mehreren Collections (ohne COLLECTION_NAME), definiere die Collection in der Abfrage:
„Durchsuche die
api-server-Collection nach Rate-Limiting-Logic“
Konfiguration von Kilo Code (VS-Code Extension)
Kilo Code ist eine VS-Code-Erweiterung with integrierter MCP-Unterstützung.
Option 1: Manuelle MCP-Konfiguration
Öffne die Kilo-Code-Einstellungen in VS Code
Wechsle zur MCP-Server-Konfiguration
Füge einen neuen Server hinzu:
Feld | Wert |
Name |
|
Command |
|
Argumente |
|
Umgebungsvariablen setzen:
Variable | Wert |
|
|
|
|
|
|
| Name deiner Projekt-Collection (z.B. |
Option 2: VS-Code settings.json
Füge in den VS-Code settings.json (Ctrl+Shift+P > Einstellungen: Benutzer-Einstellungen (JSON) öffnen) hinzu:
{
"kilocode.mcpServers": {
"qdrant": {
"command": "uv",
"args": [
"run", "--with", "fastmcp", "--with", "qdrant-client", "--with", "httpx",
"python", "/path/to/qdrant-mcp-ollama/server.py"
],
"env": {
"QDRANT_URL": "http://localhost:6333",
"OLLAMA_URL": "http://localhost:11434",
"EMBEDDING_MODEL": "bge-m3",
"COLLECTION_NAME": "my-project"
}
}
}
}Projektspezifisches Setup in Kilo Code
Für Multi-Codebase-Setups konfiguriere Kilo Code unter **projekt** (nicht global) mit einer projektbezogenen COLLECTION_NAME`. So durchsucht jedes Workspace nur seine eigene Codebasis.
Andere MCP-Clients Konfigurieren
Cursor / Windsurf
Führe den Server mit SSE-Transport aus, damit Remote-fähige Clients ihn erreichen können:
QDRANT_URL="http://localhost:6333" \
OLLAMA_URL="http://localhost:11434" \
EMBEDDING_MODEL="bge-m3" \
FASTMCP_PORT=8000 \
uv run --with fastmcp --with qdrant-client --with httpx \
python server.py --transport sseStelle dann in den MCP-Einstellungen von Cursor/Windsurf die Verbindung her: http://localhost:8000/sse
Allgemeiner MCP-Client (stdio)
Der Standard-Transport ist stdio. Jeder MCP-Client, der stdio unterstützt, kann den Server wie folgt nutzen:
uv run --with fastmcp --with qdrant-client --with httpx python server.pyKonfigurationsreferenz
Umgebungsvariablen des MCP-Servers
Variable | Beschreibung | Standard |
| Qdrant-Server-URL |
|
| API-Key für Qdrant Cloud | Keine |
| Ollama-Server-URL |
|
| Name des Ollama-Embedding-Modells |
|
| Standard-Collection (leer = pro Aufruf erforderlich) | (leer) |
Auswahl eines Embedding-Modells
Alle Modelle unten sind verfügbar über ollama pull <Model>:
Modell | Dimensionen | Größe | Geschwindigkeit | Qualität | Geeignet für |
| 1024 | 1.2 GB | Mittel | Hoch | Allgemein, mehrsprachig |
| 768 | 274 MB | Schnell | Gut | Ressourcenschonend, englisch |
| 1024 | 670 MB | Mittel | Hoch | Englisch, hohe Qualität |
| 1024 | 1.2 GB | Mittel | Sehr hoch | Beste Qualität, Englisch |
| 384 | 46 MB | Sehr schnell | Akzeptabel | Minimale Ressourcen |
Empfehlung: Beginne mit bge-m3. Es verarbeitet Code gut, unterstützt mehrsprachige Inhalte (Kommentare in jeder Sprache) und balanciert Qualität mit Geschwindigkeit.
Wichtig: Das Embedding-Modell, mit dem sie installiert wurde, muss mit dem Modell für die Suche übereinstimmen. Wenn du ein anderes Modell zum erneuten Einbetten verwendest, wollte ich die Collection wiederherstellen.
GPU-Auslastung
Größere Modelle belegen mehr GPU. Falls deine GPU nicht vollständig ausgelastet wird:
Wechsle von
nomic-embed-text(274 MB) zubge-m3(1.2 GB) oder einem größeren ModellDas Einbettungsskript sendet alle Texte in einem einzelnen Batch, um die GPU-Sättigung (Auslastung g) zu maximieren
Bei einzelnen Abfragen (z. B.
qdrant_find) sind kurze GPU-Ausschläge normal – das Einbetten einer einzelnen Abfrage SomeTime ist Millisekunden nötig
GPU-Auslastung prüfen: nvidia-smi (NVIDIA) oder rocm-smi (AMD)
MCP-Tools
qdrant_store
Speichert Informationen in der Qdrant-Datenbank.
Parameter | Typ | Erforderlich | Beschreibung |
| string | Ja | Text, der gespeichert und durchsucht werden |
| string | Wenn kein Standardwert wäre | Ziel-Collection |
| dict | Nein | Optionale Metadaten zum Anfügen |
qdrant_find
Sucht mit semantischer Ähnlichkeit nach relevanten Informationen.
Parameter | Typ | Erforderlich | Beschreibung |
| string | Ja | Suchanfrage in natürlichsprachiger Form |
| string | Wenn kein Standardwert wäre | Zu durchsuchende Sammlung |
| int | Nein | Maximale Trefferzahl, Norm:5 (default) |
Fehlerbehebung – troubleshooting
„Connection closed“ / MCP-Connection server won't start
Läuft Ollama? Überprüfe mit
ollama list. Falls erforderlich starten mitollama serve.Ist das Embedding-Modell geladen?
Einfollama pull bge-m3`.Läuft Qdrant? Prüfe mit
docker ps --filter name=qdrant-server.
„Collection existiert nicht“
Die COLLECTION wird beim Einbetten oder beim ersten qdrant_store-Aufruf erzeugt. Du hast zwei Optionen:
embed_code.pyausführen, um die Sammlung zu indizierenOder mit
qdrant_storeeinen Eintrag speichern, um die Collection automatisch und zu herzustellen
Dimension-mismatch ``-Fehler
Dies tritt auf, wenn die Collection mit einem Modell erstellt wurde, aber die Abfrage mit einem anderen Embedding-Modell durchgeführt wird.
Lösungsziele:
Collection löschen: Besuche
http://localhost:6333/dashboardMit dem richtigen Modell neu einbetten
Stelle sicher, dass
EMBEDDING_MODELin der MCP-Serverkonfiguration mit dem verwendeten Einbetmodel übereinstimmt.
„Storage folder is already accessed by another instance“
Das ist der Fehler vom offiziellen mcp-server-qdrant Modus – verwendet wird ein lokaler Speicherpfad (QDRANT_LOCAL_PATH). Dieses Projekt umgeht situ용 durch die Verbindung über eine URL zur Qdrant-Server-instanz. Stelle sicher, dass nicht beide Server auf denselben lokalen Pfad gelenkt werden.
Langsame Einbettungen / Geringe GPU-Auslastung
Höheres GPU-ModellGenisch:
Siehe Wahl des Embedding-Modell-Abschnittss oben.
Verwende ein größeres Modell:
bge-m3(1,2 GB) stattnomic-embed-text(274 MB)Das Embedding-Skript sendet alle Texte in einem Batch — wenn du Tausende von Chunks hast, maximiert das die GPU-Auslastung
Bei sehr großen Codebasen (10.000+ Dateien) solltest du erwägen, die Verarbeitung in mehrere Durchläufe pro Verzeichnis aufzuteilen
Lizenz
Apache License 2.0 — siehe LICENSE.
Available Tools
2 toolsqdrant_findC
Search for relevant information in the Qdrant database using semantic similarity.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Natural language query to search for. The query is embedded using the same GPU model used for storage, ensuring accurate results. | |
| top_k | No | Maximum number of results to return (default: 5). | |
| collection_name | No | Name of the collection to search in. Required if no default collection is configured. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read-only operation but does not explicitly state that no data is modified, does not mention return behavior, error conditions, or limitations. The single sentence provides minimal behavioral disclosure beyond the literal action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no redundancy or filler. It is front-loaded with the verb and resource. While extremely brief, it is not a tautology and conveys the essential purpose. It avoids unnecessary words while being clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a sibling (qdrant_store) and an output schema (which covers return format), the description is still incomplete. It lacks any usage context, such as when to choose this over storage or how the search integrates with the workflow. The presence of an output schema reduces the need to explain returns, but the description does not cover the selection decision or behavioral expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all three parameters (query, top_k, collection_name) with descriptions, so the baseline is 3. The description adds nothing beyond the schema; it mentions 'semantic similarity' which is already implied by the query parameter's embedding mention. No additional value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search'), the target resource ('Qdrant database'), and the method ('semantic similarity'). This distinguishes it from the sibling qdrant_store, which likely stores information. However, it does not explicitly name the sibling or contrast with it, so it lacks the full differentiation seen in higher-scoring examples.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the alternative qdrant_store, nor any mention of prerequisites or context. The description only states the action without any direction on selection or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
qdrant_storeC
Store information in the Qdrant database with GPU-accelerated embeddings.
| Name | Required | Description | Default |
|---|---|---|---|
| metadata | No | Optional metadata dictionary to attach to the stored point. | |
| information | Yes | The text information to store. This will be embedded and made searchable via semantic similarity. | |
| collection_name | No | Name of the collection to store in. Required if no default collection is configured. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states that information will be stored with embeddings, but does not disclose potential side effects such as whether existing points are overwritten, whether collections are auto-created, or any error behavior. The mutation is implied but not explicitly flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that conveys the core action. 'GPU-accelerated embeddings' adds a performance detail that may be useful context, but it could be considered extraneous. Overall, it is appropriately sized and front-loaded with the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple store operation with only 3 parameters and an output schema present, the description covers the basic action. However, it omits guidance on when a collection_name is required and does not mention any setup steps or constraints. It meets a minimum viable level but leaves gaps that an agent might need to handle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific detail beyond what the schema already provides. The only minor addition is implying that information gets embedded, which is already stated in the schema. This meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Store information in the Qdrant database') with a specific resource and purpose. It implies a write operation distinct from the sibling qdrant_find, though it doesn't explicitly differentiate. The mention of 'GPU-accelerated embeddings' adds implementation detail but doesn't obscure the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the sibling qdrant_find. The description does not say 'use this to add data, use qdrant_find to search' or mention any prerequisites like collection existence. An agent would have to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
qdrant_find - First observed
qdrant_store
TDQS
Scored across 2 tools
The two tools, qdrant_store and qdrant_find, have entirely distinct purposes—one writes data, the other retrieves it. There is zero ambiguity between them.
Both tools follow a consistent 'qdrant_<verb>' pattern, using clear action verbs (store, find). The naming is predictable and uniform.
With only two tools, the server feels thin for what is typically a database domain, but it is not an extreme mismatch. It sits at the borderline of adequacy.
The server only provides store and find, lacking any management operations like delete, update, or list. For a database, this is a significant gap that will limit workflow coverage.
Maintenance
Related MCP Connectors
Connect AI assistants to your GitHub-hosted Obsidian vault to seamlessly access, search, and analy…
Search your knowledge bases from any AI assistant using hybrid RAG.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables semantic code search across codebases using Qdrant vector database and OpenAI embeddings, allowing users to find code by meaning rather than just keywords through natural language queries.2MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to index and search codebases using semantic search powered by multiple embedding providers (OpenAI, VoyageAI, Gemini, Ollama) and vector database storage.-
- FlicenseNot gradedqualityDmaintenanceEnables semantic code search across multi-language codebases using natural language queries, integrated with Qdrant vector database for fast, cached retrieval.1-
- AlicenseNot gradedqualityFmaintenanceIndexes codebases into Qdrant for semantic search, enabling AI assistants to find relevant code by meaning without re-exploring the repo.MIT