Skip to main content
Glama
michal7kw
by michal7kw

qdrant-mcp-ollama

Qdrant ベクトルデータベース用の Model Context Protocol (MCP) サーバーです。Ollama を使用して GPU アクセラレーションによる埋め込みを実現します。

なぜ公式の mcp-server-qdrant ではないのか?

公式 Qdrant MCP サーバー は埋め込みに FastEmbed を使用しており、以下の問題があります。

  • CPU のみで動作 — 大規模なコードベースでは遅く、最新の GPU を十分に活用できない

  • 小さなモデルを使用 (all-MiniLM-L6-v2, 384-dim) — 埋め込みの品質が低い

  • ローカルモードではシングルプロセスロック — 同時にアクセスできる MCP クライアントは 1 つだけ

このサーバーはこれら 3 つの問題をすべて解決します。

公式 mcp-server-qdrant

qdrant-mcp-ollama

埋め込みエンジン

FastEmbed (CPU)

Ollama (GPU)

デフォルトモデル

% — all-MiniLM-L6-v2 (384-dim, 80MB)

bge-m3 (1024-dim, 1.2GB)

同時アクセス

なし (ローカルモード)

あり (Qdrant サーバー)

モデルの柔軟性

FastEmbed モデルのみ

任意の Ollama 埋め込みモデル

Related MCP server: Claude Context MCP

Architecture

┌──────────────┐     ┌────────────────────┐     ┌─────────────┐
│  MCP Client   │────>│  qdrant-mcp-ollama │────>│   Ollama    │
│ (Claude Code, │     │    (server.py)      │     │  (GPU)      │
│  Kilo Code,   │<────│                    │     └─────────────┘
│  Cursor, etc) │     └────────┬───────────┘
└──────────────┘              │
                              v
                   ┌────────────────────┐
                   │   Qdrant Server    │
                   │  (Docker, :6333)   │
                   │  Storage: local    │
                   │  disk / cloud      │
                   └────────────────────┘

前提条件

  • Ollama — 埋め込みモデルをプルした状態でインストール済み・実行中であること

  • Docker — Qdrant サーバーを実行するため

  • uv — Python パッケージマネージャー(推奨)または pip

クイックスタート

1. Ollama で埋め込みモデルをプルする

ollama pull bge-m3

2. Qdrant サーバーを起動する

docker run -d --name qdrant-server \
  -p 6333:6333 -p 6334:6334 \
  -v qdrant-storage:/qdrant/storage \
  --restart unless-stopped \
  qdrant/qdrant:latest

3. MCP サーバーを実行する

# No install needed — uv downloads dependencies on-the-fly:
QDRANT_URL="http://localhost:6333" \
EMBEDDING_MODEL="bge-m3" \
uv run --with fastmcp --with qdrant-client --with httpx python server.py

4. コードベースを埋め込む

uv run --with qdrant-client --with httpx python embed_codebase.py \
  /path/to/your/project my-project --preset python

5. MCP クライアントから検索する

設定が完了したら(下記のセクション参照)、AI アシスタントに次のように依頼します:

"コードベースの認証ロジックを検索して"

すると qdrant_find ツールを使い、意味的に関連するコードチャンクを返します。


Qdrant サーバーのセットアップ

オプション A: Docker(推奨)

特定のドライブ(例: Windows の E:)にデータを保存する場合:

# Create storage directories
mkdir -p E:/qdrant-storage E:/qdrant-snapshots

# Start Qdrant with persistent storage
docker run -d --name qdrant-server \
  -p 6333:6333 -p 6334:6334 \
  -v E:/qdrant-storage:/qdrant/storage \
  -v E:/qdrant-snapshots:/qdrant/snapshots \
  --restart unless-stopped \
  qdrant/qdrant:latest

Linux/macOS の場合:

docker run -d --name qdrant-server \
  -p 6333:6333 -p 6334:6334 \
  -v ~/qdrant-storage:/qdrant/storage \
  --restart unless-stopped \
  qdrant/qdrant:latest

--restart unless-stopped フラグにより、Qdrant は Docker Desktop と一緒に自動起動します。

起動を確認する:

docker ps --filter name=qdrant-server
# Or open http://localhost:6333/dashboard in your browser

オプション B: Qdrant Cloud

cloud.qdrant.io にサインアップして、URL と API キーを取得します。次に設定します:

QDRANT_URL="https://your-cluster.cloud.qdrant.io:6333"
QDRANT_API_KEY="your-api-key"

注意: QDRANT_API_KEY 環境変数は Qdrant クライアントに自動的に渡されます。


コードベースの埋め込み

embed_codebase.py スクリプトはディレクトリをスキャンし、ソースファイルをまとまりとしてチャンク化して、GPU 上の Ollama を使用して Qdrant に一括埋め込みします。

基本的な使い方

uv run --with qdrant-client --with httpx python embed_codebase.py <directory> <collection-name>

拡張子のプリセットを使用する

# Python project
python embed_codebase.py ./my-api api-backend --preset python

# Full-stack web project
python embed_codebase.py ./my-app frontend --preset web

# R / bioinformatics project
python embed_codebase.py ./analysis bio-analysis --preset r

# Everything
python embed_codebase.py ./mono-repo all-code --preset all

カスタム拡張子

python embed_codebase.py ./project my-collection --extensions .py .sql .sh .yaml

利用可能なプリセット

プリセット

拡張子

python

.py .pyi

javascript

.js .jsx .mjs .cjs

typescript

.ts .tsx

web

.js .jsx .ts .tsx .vue .svelte .html .css .scss

r

.R.r .Rmd .rmd

java

.java

csharp

.cs

go

.go

rust

.rs

cpp

.cpp .hpp .cc .hh .c .h

all

一般的なソース拡張子すべて

--preset または --extensions を指定しない場合、スクリプトがファイルタイプを自動検出します。

すべてのオプション

usage: embed_codebase.py <directory> <collection> [options]

positional arguments:
  directory              Path to the codebase directory
  collection             Qdrant collection name

options:
  --extensions EXT [EXT ...]  File extensions to include (e.g. .py .ts)
  --preset PRESET             Use a preset group of extensions
  --model MODEL               Ollama embedding model (default: bge-m3)
  --qdrant-url URL            Qdrant server URL (default: http://localhost:6333)
  --ollama-url URL            Ollama server URL (default: http://localhost:11434)
  --chunk-size N              Max lines per chunk (default: 80)
  --chunk-overlap N           Overlap lines between chunks (default: 10)
  --batch-size N              Upload batch size for Qdrant (default: 500)
  --append                    Append to existing collection instead of replacing

アペンドモード

デフォルトではスクリプトを再実行すると、コレクションが置き換えられます。既存のコレクションに追加するには --append を使用します:

# First embed
python embed_codebase.py ./src main-code --preset typescript

# Add more files later
python embed_codebase.py ./docs main-code --extensions .md --append

複数コードベースの使用

コードベースごとに別々のコレクションを使用すると、検索結果の範囲が適切に限定され、関連性が保たれます:

# Project A
python embed_codebase.py ~/projects/api-server api-server --preset python

# Project B
python embed_codebase.py ~/projects/web-app web-app --preset web

# Project C
python embed_codebase.py ~/projects/data-pipeline data-pipeline --preset python

MCP サーバーを設定するときは、次の点に注意してください。

  • COLLECTION_NAME を設定しない場合 — クエリごとにコレクションを指定する必要があります。これは 1 つの MCP サーバーで複数プロジェクトを扱う場合に最適です。

  • COLLECTION_NAME を設定する場合 — デフォルトのコレクションが自動的に使用されます。MCP クライアントがプロジェクト単位の設定をサポートしている場合は、プロジェクトごとに設定してください。


Claude Code の設定

MCP サーバーを追加する

claude mcp add qdrant -s user \
  -e QDRANT_URL="http://localhost:6333" \
  -e OLLAMA_URL="http://localhost:11434" \
  -e EMBEDDING_MODEL="bge-m3" \
  -- uv run --with fastmcp --with qdrant-client --with httpx \
     python /path/to/qdrant-mcp-ollama/server.py

/path/to/qdrant-mcp-ollama/ は、このリポジトリをクローンしたときの実際のパスに置き換えてください。

デフォルトコレクションを設定する場合

主に 1 つのプロジェクトで作業している場合は、次のように設定します:

claude mcp add qdrant -s user \
  -e QDRANT_URL="http://localhost:6333" \
  -e OLLAMA_URL="http://localhost:11434" \
  -e EMBEDDING_MODEL="bge-m3" \
  -e COLLECTION_NAME="my-project" \
  -- uv run --with fastmcp --with qdrant-client --with httpx \
     python /path/to/qdrant-mcp-ollama/server.py

動作確認

claude mcp list
# Should show: qdrant: ... ✓ Connected

claude mcp get qdrant
# Shows full configuration details

Claude Code での使用方法

設定すると、Claude Code で次のツールを使用できるようになります。

  • qdrant_store — 情報の保存:"この認証パターンを Qdrant に保存"

  • qdrant_find — 検索:"データベースマイグレーション関連のコードを検索"

複数コレクションの設定(デフォルトなし)では、コレクションを指定します。

"api-server コレクションからレート制限ロジックを検索"


Kilo Code(VS Code 拡張機能)の設定

Kilo Code は、MCP 対応の組み込み VS Code 拡張機能です。

オプション 1: 手動の MCP 設定

  1. VS Code で Kilo Code の設定を開きます。

  2. MCP サーバーの設定に移動します。

  3. 次の内容で新しいサーバーを追加します。

フィールド

値

名前

qdrant

コマンド

uv

引数

run --with fastmcp --with qdrant-client --with httpx python /path/to/server.py

  1. 環境変数を設定します。

環境変数

値

QDRANT_URL

http://localhost:6333

OLLAMA_URL

http://localhost:11434

EMBEDDING_MODEL

bge-m3

COLLECTION_NAME

使用するプロジェクトのコレクション名(例: my-project)

オプション 2: VS Code の settings.json

settings.json に次を追加します。(Ctrl+Shift+P > Preferences: Open User Settings (JSON))

{
  "kilocode.mcpServers": {
    "qdrant": {
      "command": "uv",
      "args": [
        "run", "--with", "fastmcp", "--with", "qdrant-client", "--with", "httpx",
        "python", "/path/to/qdrant-mcp-ollama/server.py"
      ],
      "env": {
        "QDRANT_URL": "http://localhost:6333",
        "OLLAMA_URL": "http://localhost:11434",
        "EMBEDDING_MODEL": "bge-m3",
        "COLLECTION_NAME": "my-project"
      }
    }
  }
}

Kilo Code のプロジェクト別設定

複数コードベースを扱う場合は、プロジェクト固有の COLLECTION_NAME を指定して、Kilo Code を プロジェクトスコープ(グローバルではなく)で設定します。これにより、各ワークスペースは自分のコードベースだけを検索するようになります。


その他の MCP クライアントの設定

Cursor / Windsurf

リモート対応クライアントでは、SSE トランスポートでサーバーを実行します。

QDRANT_URL="http://localhost:6333" \
OLLAMA_URL="http://localhost:11434" \
EMBEDDING_MODEL="bge-m3" \
FASTMCP_PORT=8000 \
uv run --with fastmcp --with qdrant-client --with httpx \
  python server.py --transport sse

Cursor / Windsurf の MCP 設定で、http://localhost:8000/sse に接続します。

汎用 MCP クライアント (stdio)

デフォルトのトランスポートは stdio です。stdio をサポートする MCP クライアントは、このコマンドでこのサーバーを使用できます。

uv run --with fastmcp --with qdrant-client --with httpx python server.py

設定リファレンス

MCP サーバー環境変数

環境変数

説明

デフォルト

QDRANT_URL

Qdrant サーバーの URL(例 : http://localhost:6333)

http://localhost:6333

QDRANT_API_KEY

Qdrant Cloud 用 API キー

None

OLLAMA_URL

Ollama サーバーの URL(例 : http://localhost:11434)

http://localhost:11434

EMBEDDING_MODEL

Ollama の埋め込みモデル名

bge-m3

COLLECTION_NAME

デフォルトのコレクション(空 = 呼び出しごとに指定が必要)

(空)

埋め込みモデルの選選択

以下のモデルはすべて ollama pull <model> で利用でできます。

モデル

次元数数

サイズズ

速度伏

品質

最適する用用途途

bge-m3

1024

1.2GB

中程度度

高

汎用・多言言語

nomic-embnej

768

274MB

高速速

良

軽量、英語中心心

mxbai-emb-shell

mxbai-emb-large → 1024

670MB

中程度度

高

英語・高品質

snowmsflake-artgeneric-emodd2

snowflake-arctic-embed2 → 1024

1.2GB

中程度

非と高

最高品質・英語英

all-minilm

386

46MB

非常高速速

普通

最小リソース源

推奨: まずは bge-m3 から始めてください。コードの扱いに優れ、多言語コンテンツ(あらゆる言語のコメント)に対応し、品質と速度のバランスが良好です。

重要: コレクションのインデックスに使用した埋め込みモデルは、クエリに使用するモデルと一致している必要があります。別のモデルで再埋め込みする場合は、コレクションを削除して再度作成してください。

GPU 使用率

大きなモデルほど GPU を使用します。GPU がフル活用されていない場合は:

  • nomic-embed-text(274 MB) から bge-m3(1.2 GB) 以上のモデルに切代え替える

  • 埋に込みスクリプトはすべての文テキストを一度に送信し、GPU 利利率率を最最化

  • 単一クエリリ(qdrant_find)でる GPる使用率率のスパイククは短瞬間で正常で(単一問一併わせの埋め込込みは数ミリリ秒で完了)

GPU 使使用率を確認確認する: nvidia-smi (NVIDIA) または rocm-smi (AMD)


MCPツールズ

qrant_store

Qドランt デデータベースに情報を保存します。

パラメーター

型

必必須

説明

information

string

必須

保存して検索可能対象となるテキスト

collection_name

string

デフォルトが未設定の場合

対象コレクション

metadata

dict

任意

添付する任意のメタデータ

qrant_fund

セマンティックス度検索を検索します。

パラメメーータター

型

必要須

説明語

query

string

必省

自然言語の検索クエリ

collection_name

string

デフォルト未定設定の場場合

対对象 collectionターゲット

top_k

int

任意

れターンする最最大件数 (デフォールルト: 5)


トブルシューティング

「接続が断されまし」 / M CPサーババーが起動んでない

  • は Ollama が動作動していますか? ollama list で確認確認してください。必要なければ ollama serve で起動動してください。

  • は埋め込こみコルデルがルル済みみか? llama pull bge-m3 を実実行してください。

  • は Odrant が件動してますか? docker ls --filter name=qdrant-serve で確認してください。

「コレクションが存在しません」

コードベースの作コレクションは、コードベースの埋め込みスクリプト作るか、初回の qdrant_store 呼び出しで自動作成されます。次のいずれかを実行してください。

  • まずコードベースのインデックス化を行うため embed_codebase.py を実行する

  • または qdrant_store で情報を保存してコレクションを自動生成する

次元誤差

埋め込み時に使用するモデルと、クエリ時に使用するぶモデルが異なる場合に発生します。対処法は以下の通りです。

  1. コレクションを削除する: http://localhost:6333/ダッシュボード を訪ねてください

  2. 正しいモデルで再埋め込みする

  3. MCP サーバー設定の EMBEDDING_MODEL が、埋め込みに使用したモデルと一致しているか確認する

。「ストレージフォルダががすに他のインスタンスへアクセスされていきます」

このこの問題は、公式 mcp-server-qdrant がを使用する場合に、ローカルモード (QDRANT_LOCAL_PATH) で発生します。このリポジがは URL 経由で Qdrant サーバに接続して避けています。両方のサーバーが同じ蹊ローカルパスを指し示すことらないようご注意ください。

埋め込みが遅い / GPU の活用率が低い

  • より大きなモデルを使用する: nomic-embed-text (274 MB) の代わりに bge-m3 (1.2 GB)

  • 埋め込みスクリプトはすべてのテキストを1つのバッチで送信します — 数千のチャンクがある場合、これによりGPU使用率が最大化されます

  • 非常に大きなコードベース(10,000+ファイル)の場合は、ディレクトリごとに複数回に分けて実行することを検討してください


License

Apache License 2.0 — LICENSE を参照してください。

Available Tools

2 tools
qdrant_findC

Search for relevant information in the Qdrant database using semantic similarity.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural language query to search for. The query is embedded using the same GPU model used for storage, ensuring accurate results.
top_kNoMaximum number of results to return (default: 5).
collection_nameNoName of the collection to search in. Required if no default collection is configured.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read-only operation but does not explicitly state that no data is modified, does not mention return behavior, error conditions, or limitations. The single sentence provides minimal behavioral disclosure beyond the literal action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no redundancy or filler. It is front-loaded with the verb and resource. While extremely brief, it is not a tautology and conveys the essential purpose. It avoids unnecessary words while being clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a sibling (qdrant_store) and an output schema (which covers return format), the description is still incomplete. It lacks any usage context, such as when to choose this over storage or how the search integrates with the workflow. The presence of an output schema reduces the need to explain returns, but the description does not cover the selection decision or behavioral expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all three parameters (query, top_k, collection_name) with descriptions, so the baseline is 3. The description adds nothing beyond the schema; it mentions 'semantic similarity' which is already implied by the query parameter's embedding mention. No additional value is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Search'), the target resource ('Qdrant database'), and the method ('semantic similarity'). This distinguishes it from the sibling qdrant_store, which likely stores information. However, it does not explicitly name the sibling or contrast with it, so it lacks the full differentiation seen in higher-scoring examples.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the alternative qdrant_store, nor any mention of prerequisites or context. The description only states the action without any direction on selection or exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

qdrant_storeC

Store information in the Qdrant database with GPU-accelerated embeddings.

ParametersJSON Schema
NameRequiredDescriptionDefault
metadataNoOptional metadata dictionary to attach to the stored point.
informationYesThe text information to store. This will be embedded and made searchable via semantic similarity.
collection_nameNoName of the collection to store in. Required if no default collection is configured.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states that information will be stored with embeddings, but does not disclose potential side effects such as whether existing points are overwritten, whether collections are auto-created, or any error behavior. The mutation is implied but not explicitly flagged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that conveys the core action. 'GPU-accelerated embeddings' adds a performance detail that may be useful context, but it could be considered extraneous. Overall, it is appropriately sized and front-loaded with the action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple store operation with only 3 parameters and an output schema present, the description covers the basic action. However, it omits guidance on when a collection_name is required and does not mention any setup steps or constraints. It meets a minimum viable level but leaves gaps that an agent might need to handle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific detail beyond what the schema already provides. The only minor addition is implying that information gets embedded, which is already stated in the schema. This meets the baseline but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Store information in the Qdrant database') with a specific resource and purpose. It implies a write operation distinct from the sibling qdrant_find, though it doesn't explicitly differentiate. The mention of 'GPU-accelerated embeddings' adds implementation detail but doesn't obscure the core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the sibling qdrant_find. The description does not say 'use this to add data, use qdrant_find to search' or mention any prerequisites like collection existence. An agent would have to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedqdrant_find
    • First observedqdrant_store

TDQS

B3.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools, qdrant_store and qdrant_find, have entirely distinct purposes—one writes data, the other retrieves it. There is zero ambiguity between them.

Naming Consistency5/5

Both tools follow a consistent 'qdrant_<verb>' pattern, using clear action verbs (store, find). The naming is predictable and uniform.

Tool Count3/5

With only two tools, the server feels thin for what is typically a database domain, but it is not an extreme mismatch. It sits at the borderline of adequacy.

Completeness2/5

The server only provides store and find, lacking any management operations like delete, update, or list. For a database, this is a significant gap that will limit workflow coverage.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers