qdrant-mcp-ollama
qdrant-mcp-ollama
Сервер Model Context Protocol (MCP) для векторной базы данных Qdrant, использующий Ollama для GPU-ускоренных эмбеддингов.
Почему не официальный mcp-server-qdrant?
Официальный MCP-сервер Qdrant использует FastEmbed для эмбеддингов, которые:
Работают только на CPU — медленно на больших кодовых базах, не задействуют современные GPU
Используют маленькую модель (
all-MiniLM-L6-v2, 384-dim) — менее качественные эмбеддингиОднопроцессная блокировка в локальном режиме — только один MCP-клиент может одновременно обращаться к базе данных
Этот сервер решает все три проблемы:
Официальный |
| |
Движок эмбеддинга | FastEmbed (CPU) | Ollama (GPU) |
Модель по умолчанию | all-MiniLM-L6-v2 (384-dim, 80MB) | bge-m3 (1024-dim, 1.2GB) |
Одновременный доступ | Нет (локальный режим) | Да (сервер Qdrant) |
Гибкость моделей | Только модели FastEmbed | Любая эмбеддинг-модель Ollama |
Related MCP server: Claude Context MCP
Архитектура
┌──────────────┐ ┌────────────────────┐ ┌─────────────┐
│ MCP Client │────>│ qdrant-mcp-ollama │────>│ Ollama │
│ (Claude Code, │ │ (server.py) │ │ (GPU) │
│ Kilo Code, │<────│ │ └─────────────┘
│ Cursor, etc) │ └────────┬───────────┘
└──────────────┘ │
v
┌────────────────────┐
│ Qdrant Server │
│ (Docker, :6333) │
│ Storage: local │
│ disk / cloud │
└────────────────────┘Предварительные требования
Ollama — установлена и запущена, эмбеддинг-модель скачана
Docker — для запуска сервера Qdrant
uv — менеджер пакетов Python (рекомендуется) или
pip
Быстрый старт
1. Скачайте эмбеддинг-модель в Ollama
ollama pull bge-m32. Запустите сервер Qdrant
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v qdrant-storage:/qdrant/storage \
--restart unless-stopped \
qdrant/qdrant:latest3. Запустите MCP-сервер
# No install needed — uv downloads dependencies on-the-fly:
QDRANT_URL="http://localhost:6333" \
EMBEDDING_MODEL="bge-m3" \
uv run --with fastmcp --with qdrant-client --with httpx python server.py4. Индексируйте кодовую базу
uv run --with qdrant-client --with httpx python embed_codebase.py \
/path/to/your/project my-project --preset python5. Поиск через вашего MCP-клиента
После настройки (см. разделы ниже) спросите вашего ИИ-ассистента:
"Найди в кодовой базе логику аутентификации"
Он использует инструмент qdrant_find, чтобы вернуть семантически релевантные фрагменты кода.
Настройка сервера Qdrant
Вариант A: Docker (рекомендуется)
Храните данные на отдельном диске (например, E: в Windows):
# Create storage directories
mkdir -p E:/qdrant-storage E:/qdrant-snapshots
# Start Qdrant with persistent storage
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v E:/qdrant-storage:/qdrant/storage \
-v E:/qdrant-snapshots:/qdrant/snapshots \
--restart unless-stopped \
qdrant/qdrant:latestНа Linux/macOS:
docker run -d --name qdrant-server \
-p 6333:6333 -p 6334:6334 \
-v ~/qdrant-storage:/qdrant/storage \
--restart unless-stopped \
qdrant/qdrant:latestФлаг --restart unless-stopped гарантирует, что Qdrant запускается автоматически вместе с Docker Desktop.
Проверьте, что он запущен:
docker ps --filter name=qdrant-server
# Or open http://localhost:6333/dashboard in your browserВариант B: Qdrant Cloud
Зарегистрируйтесь на cloud.qdrant.io и получите URL и API-ключ. Затем укажите их:
QDRANT_URL="https://your-cluster.cloud.qdrant.io:6333"
QDRANT_API_KEY="your-api-key"Примечание: Переменная окружения
QDRANT_API_KEYавтоматически передаётся клиенту Qdrant.
Эмбеддинг кодовой базы
Скрипт embed_codebase.py сканирует каталог, разбивает исходные файлы на фрагменты и массово индексирует их в Qdrant, используя Ollama на GPU.
Базовое использование
uv run --with qdrant-client --with httpx python embed_codebase.py <directory> <collection-name>Использование пресетов расширений
# Python project
python embed_codebase.py ./my-api api-backend --preset python
# Full-stack web project
python embed_codebase.py ./my-app frontend --preset web
# R / bioinformatics project
python embed_codebase.py ./analysis bio-analysis --preset r
# Everything
python embed_codebase.py ./mono-repo all-code --preset allПользовательские расширения
python embed_codebase.py ./project my-collection --extensions .py .sql .sh .yamlДоступные пресеты
Пресет | Расширения |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| Все распространённые расширения исходного кода |
Если не указаны --preset или --extensions, скрипт определяет типы файлов автоматически.
Все параметры
usage: embed_codebase.py <directory> <collection> [options]
positional arguments:
directory Path to the codebase directory
collection Qdrant collection name
options:
--extensions EXT [EXT ...] File extensions to include (e.g. .py .ts)
--preset PRESET Use a preset group of extensions
--model MODEL Ollama embedding model (default: bge-m3)
--qdrant-url URL Qdrant server URL (default: http://localhost:6333)
--ollama-url URL Ollama server URL (default: http://localhost:11434)
--chunk-size N Max lines per chunk (default: 80)
--chunk-overlap N Overlap lines between chunks (default: 10)
--batch-size N Upload batch size for Qdrant (default: 500)
--append Append to existing collection instead of replacingРежим дополнения
По умолчанию повторный запуск скрипта заменяет коллекцию. Используйте --append, чтобы добавить данные в существующую коллекцию:
# First embed
python embed_codebase.py ./src main-code --preset typescript
# Add more files later
python embed_codebase.py ./docs main-code --extensions .md --appendИспользование нескольких кодовых баз
Используйте отдельные коллекции для каждой кодовой базы, чтобы результаты поиска были релевантными и не смешивались:
# Project A
python embed_codebase.py ~/projects/api-server api-server --preset python
# Project B
python embed_codebase.py ~/projects/web-app web-app --preset web
# Project C
python embed_codebase.py ~/projects/data-pipeline data-pipeline --preset pythonПри настройке MCP-сервера:
Без
COLLECTION_NAME: коллекцию необходимо указывать при каждом запросе. Идеально, когда один MCP-сервер обслуживает несколько проектов.С
COLLECTION_NAME: коллекция по умолчанию используется автоматически. Задавайте её отдельно для каждого проекта, если ваш MCP-клиент поддерживает конфигурацию в рамках проекта.
Настройка Claude Code
Добавление MCP-сервера
claude mcp add qdrant -s user \
-e QDRANT_URL="http://localhost:6333" \
-e OLLAMA_URL="http://localhost:11434" \
-e EMBEDDING_MODEL="bge-m3" \
-- uv run --with fastmcp --with qdrant-client --with httpx \
python /path/to/qdrant-mcp-ollama/server.pyЗамените /path/to/qdrant-mcp-ollama/ на фактический путь, куда вы клонировали этот репозиторий.
С коллекцией по умолчанию
Если вы работаете преимущественно с одним проектом:
claude mcp add qdrant -s user \
-e QDRANT_URL="http://localhost:6333" \
-e OLLAMA_URL="http://localhost:11434" \
-e EMBEDDING_MODEL="bge-m3" \
-e COLLECTION_NAME="my-project" \
-- uv run --with fastmcp --with qdrant-client --with httpx \
python /path/to/qdrant-mcp-ollama/server.pyПроверка
claude mcp list
# Should show: qdrant: ... ✓ Connected
claude mcp get qdrant
# Shows full configuration detailsИспользование в Claude Code
После настройки Claude Code сможет использовать такие инструменты:
qdrant_store— сохранить информацию: "Сохрани этот паттерн аутентификации в Qdrant"qdrant_find— поиск: "Найди код, связанный с миграциями базы данных"
Для конфигураций с несколькими коллекциями (без коллекции по умолчанию) указывайте коллекцию явно:
"Найди в коллекции
api-serverлогику ограничения частоты запросов"
Настройка Kilo Code (расширение VS Code)
Kilo Code — расширение VS Code со встроенной поддержкой MCP.
Вариант 1: Настройка MCP вручную
Откройте настройки Kilo Code в VS Code.
Перейдите к конфигурации MCP-серверов.
Добавьте новый сервер со следующими параметрами:
Поле | Значения |
Имя |
|
Команда |
|
Аргументы |
|
Задайте переменные окружения:
Переменная | Значение |
|
|
|
|
|
|
| Название коллекции вашего проекта (например, |
Вариант 2: settings.json в VS Code
Добавьте в ваш файл settings.json в VS Code (Ctrl+Shift+P > Preferences: Open User Settings (JSON)):
{
"kilocode.mcpServers": {
"qdrant": {
"command": "uv",
"args": [
"run", "--with", "fastmcp", "--with", "qdrant-client", "--with", "httpx",
"python", "/path/to/qdrant-mcp-ollama/server.py"
],
"env": {
"QDRANT_URL": "http://localhost:6333",
"OLLAMA_URL": "http://localhost:11434",
"EMBEDDING_MODEL": "bge-m3",
"COLLECTION_NAME": "my-project"
}
}
}
}Настройка на уровне проекта в Kilo Code
Для конфигураций с несколькими кодовыми базами настройте Kilo Code на уровне проекта (не глобально), указав COLLECTION_NAME для конкретного проекта. Так каждая рабочая область будет искать только в своей кодовой базе.
Настройка других MCP-клиентов
Cursor / Windsurf
Запустите сервер с транспортом SSE для клиентов с удалённым доступом:
QDRANT_URL="http://localhost:6333" \
OLLAMA_URL="http://localhost:11434" \
EMBEDDING_MODEL="bge-m3" \
FASTMCP_PORT=8000 \
uv run --with fastmcp --with qdrant-client --with httpx \
python server.py --transport sseЗатем в настройках MCP в Cursor/Windsurf подключитесь к адресу: http://localhost:8000/sse
Стандартный MCP-клиент (stdio)
Транспорт по умолчанию — stdio. Любой MCP-клиент с поддержкой stdio может использовать этот сервер, запустив команду:
uv run --with fastmcp --with qdrant-client --with httpx python server.pyСправочник по конфигурации
Переменные окружения MCP-сервера
Переменные | Описание | По умолчанию |
| URL сервера Qdrant |
|
| API-ключ для Qdrant Cloud | Нет |
| URL сервера Ollama |
|
| Название эмбеддинг-модели Ollama |
|
| Коллекция по умолчанию (пусто = указывать в каждом вызове) | (пусто) |
Выбор эмбеддинг-модели
Все модели ниже доступны через ollama pull <model>:
Модель | Размерность | Размер | Скорость | Качество | Для чего подходит |
| 1024 | 1.2 GB | Средняя | Высокое | Общее назначение, многоязычность |
| 768 | 274 MB | Быстрая | Хорошее | Лёгкая, ориентирована на английский |
| 1024 | 670 MB | Средняя | Высокое | Английский, высокое качество |
| 1024 | 1.2 GB | Средняя | Очень высокое | Лучшее качество, английский |
| 384 | 46 MB | Очень быстрая | Среднее | Минимальные ресурсы |
Рекомендация: Начните с bge-m3. Она хорошо работает с кодом, поддерживает многоязычный контент (комментарии на любом языке) и балансирует между качеством и скоростью.
Важно: Эмбеддинг-модель, использованная для индексации коллекции, должна совпадать с моделью, используемой для запросов. Если вы переиндексируете с другой моделью, удалите и создайте коллекцию заново.
Использование GPU
Более крупные модели занимают больше GPU. Если GPU недогружен:
Перейдите с
nomic-embed-text(274 MB) наbge-m3(1.2 GB) или более крупную модель.Скрипт эмбеддинга отправляет все тексты одним пакетом, чтобы максимально нагрузить GPU.
При отдельных запросах (через
qdrant_find) нагрузка на GPU кратковременна — это нормально, эмбеддинг одиночного запроса занимает миллисекунды.
Проверьте использование GPU: nvidia-smi (NVIDIA) или rocm-smi (AMD).
Инструменты MCP
qdrant_store
Сохраняет информацию в базе данных Qdrant.
Параметры | Тип | Обязателен | Описание |
| string | Да | Текст для сохранения и последующего поиска |
| string | Если не задана коллекция по умолчанию | Целевая коллекция |
| dict | Нет | Необязательные метаданные |
qdrant_find
Выполняет поиск релевантной информации с использованием семантической близости.
Параметры | Тип | Обязателен | Описание |
| string | Да | Поисковый запрос на естественном языке |
| string | Если не задана коллекция по умолчанию | Коллекция для поиска |
| int | Нет | Максимальное количество результатов (по умолчанию: 5) |
Устранение неполадок
«Connection closed» / MCP-сервер не запускается
Запущен ли Ollama? Проверьте командой
ollama list. При необходимости запустите с помощьюollama serve.Загружена ли эмбеддинг-модель? Выполните
ollama pull bge-m3.Запущен ли Qdrant? Проверьте командой
docker ps --filter name=qdrant-server.
Ошибка «Collection does not exist»
Коллекция создаётся скриптом эмбеддинга или при первом вызове qdrant_store. Выполните одно из действий:
Запустите
embed_codebase.py, чтобы сначала проиндексировать кодовую базу.Или сохраните что-нибудь через
qdrant_store, чтобы коллекция создалась автоматически.
Ошибки несовпадения размерности
Это происходит, когда коллекция была создана с одной эмбеддинг-моделью, а запрос выполняется другой. Исправление:
Удалите коллекцию: откройте
http://localhost:6333/dashboard.Заново выполните эмбеддинг с правильной моделью.
Убедитесь, что
EMBEDDING_MODELв конфигурации MCP-сервера совпадает с той, что использовалась для эмбеддинга.
Ошибка «Storage folder is already accessed by another instance»
Эта ошибка возникает в официальном mcp-server-qdrant при использовании локального режима (QDRANT_LOCAL_PATH). Данный проект избегает эту ошибку, подключаясь к серверу Qdrant по URL. Убедитесь, что оба сервера не используют один локальный путь.
Медленный эмбеддинг / низкая загрузка GPU
Используйте более крупную модель:
bge-m3(1.2 ГБ) вместоnomic-embed-text(274 МБ)Скрипт эмбеддинга отправляет все тексты одним пакетом — если у вас тысячи фрагментов, это максимизирует загрузку GPU
Для очень больших кодовых баз (10 000+ файлов) рассмотрите возможность разбиения на несколько запусков по каталогам
Лицензия
Apache License 2.0 — см. LICENSE.
Available Tools
2 toolsqdrant_findC
Search for relevant information in the Qdrant database using semantic similarity.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Natural language query to search for. The query is embedded using the same GPU model used for storage, ensuring accurate results. | |
| top_k | No | Maximum number of results to return (default: 5). | |
| collection_name | No | Name of the collection to search in. Required if no default collection is configured. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read-only operation but does not explicitly state that no data is modified, does not mention return behavior, error conditions, or limitations. The single sentence provides minimal behavioral disclosure beyond the literal action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no redundancy or filler. It is front-loaded with the verb and resource. While extremely brief, it is not a tautology and conveys the essential purpose. It avoids unnecessary words while being clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a sibling (qdrant_store) and an output schema (which covers return format), the description is still incomplete. It lacks any usage context, such as when to choose this over storage or how the search integrates with the workflow. The presence of an output schema reduces the need to explain returns, but the description does not cover the selection decision or behavioral expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all three parameters (query, top_k, collection_name) with descriptions, so the baseline is 3. The description adds nothing beyond the schema; it mentions 'semantic similarity' which is already implied by the query parameter's embedding mention. No additional value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search'), the target resource ('Qdrant database'), and the method ('semantic similarity'). This distinguishes it from the sibling qdrant_store, which likely stores information. However, it does not explicitly name the sibling or contrast with it, so it lacks the full differentiation seen in higher-scoring examples.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the alternative qdrant_store, nor any mention of prerequisites or context. The description only states the action without any direction on selection or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
qdrant_storeC
Store information in the Qdrant database with GPU-accelerated embeddings.
| Name | Required | Description | Default |
|---|---|---|---|
| metadata | No | Optional metadata dictionary to attach to the stored point. | |
| information | Yes | The text information to store. This will be embedded and made searchable via semantic similarity. | |
| collection_name | No | Name of the collection to store in. Required if no default collection is configured. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states that information will be stored with embeddings, but does not disclose potential side effects such as whether existing points are overwritten, whether collections are auto-created, or any error behavior. The mutation is implied but not explicitly flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that conveys the core action. 'GPU-accelerated embeddings' adds a performance detail that may be useful context, but it could be considered extraneous. Overall, it is appropriately sized and front-loaded with the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple store operation with only 3 parameters and an output schema present, the description covers the basic action. However, it omits guidance on when a collection_name is required and does not mention any setup steps or constraints. It meets a minimum viable level but leaves gaps that an agent might need to handle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific detail beyond what the schema already provides. The only minor addition is implying that information gets embedded, which is already stated in the schema. This meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Store information in the Qdrant database') with a specific resource and purpose. It implies a write operation distinct from the sibling qdrant_find, though it doesn't explicitly differentiate. The mention of 'GPU-accelerated embeddings' adds implementation detail but doesn't obscure the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the sibling qdrant_find. The description does not say 'use this to add data, use qdrant_find to search' or mention any prerequisites like collection existence. An agent would have to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
qdrant_find - First observed
qdrant_store
TDQS
Scored across 2 tools
The two tools, qdrant_store and qdrant_find, have entirely distinct purposes—one writes data, the other retrieves it. There is zero ambiguity between them.
Both tools follow a consistent 'qdrant_<verb>' pattern, using clear action verbs (store, find). The naming is predictable and uniform.
With only two tools, the server feels thin for what is typically a database domain, but it is not an extreme mismatch. It sits at the borderline of adequacy.
The server only provides store and find, lacking any management operations like delete, update, or list. For a database, this is a significant gap that will limit workflow coverage.
Maintenance
Related MCP Connectors
Connect AI assistants to your GitHub-hosted Obsidian vault to seamlessly access, search, and analy…
Search your knowledge bases from any AI assistant using hybrid RAG.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables semantic code search across codebases using Qdrant vector database and OpenAI embeddings, allowing users to find code by meaning rather than just keywords through natural language queries.2MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to index and search codebases using semantic search powered by multiple embedding providers (OpenAI, VoyageAI, Gemini, Ollama) and vector database storage.-
- FlicenseNot gradedqualityDmaintenanceEnables semantic code search across multi-language codebases using natural language queries, integrated with Qdrant vector database for fast, cached retrieval.1-
- AlicenseNot gradedqualityFmaintenanceIndexes codebases into Qdrant for semantic search, enabling AI assistants to find relevant code by meaning without re-exploring the repo.MIT