Qdrant MCP Server
Used as the default embedding provider for generating vector embeddings locally, requiring no API keys. The server works out of the box with Ollama, pulling models such as nomic-embed-text to embed documents and code for semantic search.
Can be selected as the embedding provider (EMBEDDING_PROVIDER=openai) with an OpenAI API key to generate vector embeddings for documents and code search instead of using local models.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Qdrant MCP Serversearch my codebase for authentication middleware"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ce dépôt est un composant MCP du Hephaistos-Kit — utilisable seul, mais conçu pour être cloné en sous-module et installé via
mcp/install.ps1.🙏 Adapté de mhalder/qdrant-mcp-server — merci aux développeurs originaux pour la base qui a permis cette intégration au kit.
Qdrant MCP Server
A Model Context Protocol (MCP) server providing semantic search capabilities using Qdrant vector database with multiple embedding providers.
Features
Zero Setup: Works out of the box with Ollama - no API keys required
Privacy-First: Local embeddings and vector storage - data never leaves your machine
Code Vectorization: Intelligent codebase indexing with AST-aware chunking and semantic code search
Multiple Providers: Ollama (default), OpenAI, Cohere, and Voyage AI
Hybrid Search: Combine semantic and keyword search for better results
Semantic Search: Natural language search with metadata filtering
Incremental Indexing: Efficient updates - only re-index changed files
Configurable Prompts: Create custom prompts for guided workflows without code changes
Rate Limiting: Intelligent throttling with exponential backoff
Full CRUD: Create, search, and manage collections and documents
Flexible Deployment: Run locally (stdio) or as a remote HTTP server
Related MCP server: MCP Codebase Index
Quick Start
Prerequisites
Node.js 20+
Docker and Docker Compose
Installation
# Clone and install
git clone https://github.com/mhalder/qdrant-mcp-server.git
cd qdrant-mcp-server
npm install
# Start services and pull model
docker compose up -d
docker exec ollama ollama pull nomic-embed-text
# Build
npm run buildConfiguration
Local Setup (stdio transport)
Add to ~/.claude/claude_code_config.json:
{
"mcpServers": {
"qdrant": {
"command": "node",
"args": ["/path/to/qdrant-mcp-server/build/index.js"],
"env": {
"QDRANT_URL": "http://localhost:6333",
"EMBEDDING_BASE_URL": "http://localhost:11434"
}
}
}
}Remote Setup (HTTP transport)
⚠️ Security Warning: When deploying the HTTP transport in production:
Always run behind a reverse proxy (nginx, Caddy) with HTTPS
Implement authentication/authorization at the proxy level
Use firewalls to restrict access to trusted networks
Never expose directly to the public internet without protection
Consider implementing rate limiting at the proxy level
Monitor server logs for suspicious activity
Start the server:
TRANSPORT_MODE=http HTTP_PORT=3000 node build/index.jsConfigure client:
{
"mcpServers": {
"qdrant": {
"url": "http://your-server:3000/mcp"
}
}
}Using a different provider:
"env": {
"EMBEDDING_PROVIDER": "nvidia", // or "openai", "cohere", "voyage"
"NVIDIA_API_KEY": "nvapi-...", // provider-specific API key
"QDRANT_URL": "http://localhost:6333"
}Restart after making changes.
See Advanced Configuration section below for all options.
Tools
Collection Management
Tool | Description |
| Create collection with specified distance metric (Cosine/Euclid/Dot) |
| List all collections |
| Get collection details and statistics |
| Delete collection and all documents |
Document Operations
Tool | Description |
| Add documents with automatic embedding (supports string/number IDs, metadata) |
| Natural language search with optional metadata filtering |
| Hybrid search combining semantic and keyword (BM25) search with RRF |
| Delete specific documents by ID |
Code Vectorization
Tool | Description |
| Index a codebase for semantic code search with AST-aware chunking |
| Search indexed codebase using natural language queries |
| Incrementally re-index only changed files (detects added/modified/deleted) |
| Get indexing status and statistics for a codebase |
| Delete all indexed data for a codebase |
Resources
qdrant://collections- List all collectionsqdrant://collection/{name}- Collection details
Configurable Prompts
Create custom prompts tailored to your specific use cases without modifying code. Prompts provide guided workflows for common tasks.
Note: By default, the server looks for prompts.json in the project root directory. If the file exists, prompts are automatically loaded. You can specify a custom path using the PROMPTS_CONFIG_FILE environment variable.
Setup
Create a prompts configuration file (e.g.,
prompts.jsonin the project root):See
prompts.example.jsonfor example configurations you can copy and customize.Configure the server (optional - only needed for custom path):
If you place prompts.json in the project root, no additional configuration is needed. To use a custom path:
{
"mcpServers": {
"qdrant": {
"command": "node",
"args": ["/path/to/qdrant-mcp-server/build/index.js"],
"env": {
"QDRANT_URL": "http://localhost:6333",
"PROMPTS_CONFIG_FILE": "/custom/path/to/prompts.json"
}
}
}
}Use prompts in your AI assistant:
Claude Code:
/mcp__qdrant__find_similar_docs papers "neural networks" 10VSCode:
/mcp.qdrant.find_similar_docs papers "neural networks" 10Example Prompts
See prompts.example.json for ready-to-use prompts including:
find_similar_docs- Semantic search with result explanationsetup_rag_collection- Create RAG-optimized collectionsanalyze_collection- Collection insights and recommendationsbulk_add_documents- Guided bulk document insertionsearch_with_filter- Metadata filtering assistancecompare_search_methods- Semantic vs hybrid search comparisoncollection_maintenance- Maintenance and cleanup workflowsmigrate_to_hybrid- Collection migration guide
Template Syntax
Templates use {{variable}} placeholders:
Required arguments must be provided
Optional arguments use defaults if not specified
Unknown variables are left as-is in the output
Code Vectorization
Intelligently index and search your codebase using semantic code search. Perfect for AI-assisted development, code exploration, and understanding large codebases.
Features
AST-Aware Chunking: Intelligent code splitting at function/class boundaries using tree-sitter
Multi-Language Support: 35+ file types including TypeScript, Python, Java, Go, Rust, C++, and more
Incremental Updates: Only re-index changed files for fast updates
Smart Ignore Patterns: Respects .gitignore, .dockerignore, and custom .contextignore files
Semantic Search: Natural language queries to find relevant code
Metadata Filtering: Filter by file type, path patterns, or language
Local-First: All processing happens locally - your code never leaves your machine
Quick Start
1. Index your codebase:
# Via Claude Code MCP tool
/mcp__qdrant__index_codebase /path/to/your/project2. Search your code:
# Natural language search
/mcp__qdrant__search_code /path/to/your/project "authentication middleware"
# Filter by file type
/mcp__qdrant__search_code /path/to/your/project "database schema" --fileTypes .ts,.js
# Filter by path pattern
/mcp__qdrant__search_code /path/to/your/project "API endpoints" --pathPattern src/api/**3. Update after changes:
# Incrementally re-index only changed files
/mcp__qdrant__reindex_changes /path/to/your/projectUsage Examples
Index a TypeScript Project
// The MCP tool automatically:
// 1. Scans all .ts, .tsx, .js, .jsx files
// 2. Respects .gitignore patterns (skips node_modules, dist, etc.)
// 3. Chunks code at function/class boundaries
// 4. Generates embeddings using your configured provider
// 5. Stores in Qdrant with metadata (file path, line numbers, language)
index_codebase({
path: "/workspace/my-app",
forceReindex: false // Set to true to re-index from scratch
})
// Output:
// ✓ Indexed 247 files (1,823 chunks) in 45.2sSearch for Authentication Code
search_code({
path: "/workspace/my-app",
query: "how does user authentication work?",
limit: 5
})
// Results include file path, line numbers, and code snippets:
// [
// {
// filePath: "src/auth/middleware.ts",
// startLine: 15,
// endLine: 42,
// content: "export async function authenticateUser(req: Request) { ... }",
// score: 0.89,
// language: "typescript"
// },
// ...
// ]Search with Filters
// Only search TypeScript files
search_code({
path: "/workspace/my-app",
query: "error handling patterns",
fileTypes: [".ts", ".tsx"],
limit: 10
})
// Only search in specific directories
search_code({
path: "/workspace/my-app",
query: "API route handlers",
pathPattern: "src/api/**",
limit: 10
})Incremental Re-indexing
// After making changes to your codebase
reindex_changes({
path: "/workspace/my-app"
})
// Output:
// ✓ Updated: +3 files added, ~5 files modified, -1 files deleted
// ✓ Chunks: +47 added, -23 deleted in 8.3sCheck Indexing Status
get_index_status({
path: "/workspace/my-app"
})
// Output:
// {
// isIndexed: true,
// collectionName: "code_a3f8d2e1",
// chunksCount: 1823,
// filesCount: 247,
// lastUpdated: "2025-01-30T10:15:00Z",
// languages: ["typescript", "javascript", "json"]
// }Supported Languages
Programming Languages (35+ file types):
Web: TypeScript, JavaScript, Vue, Svelte
Backend: Python, Java, Go, Rust, Ruby, PHP
Systems: C, C++, C#
Mobile: Swift, Kotlin, Dart
Functional: Scala, Clojure, Haskell, OCaml
Scripting: Bash, Shell, Fish
Data: SQL, GraphQL, Protocol Buffers
Config: JSON, YAML, TOML, XML, Markdown
See configuration for full list and customization options.
Custom Ignore Patterns
Create a .contextignore file in your project root to specify additional patterns to ignore:
# .contextignore
**/test/**
**/*.test.ts
**/*.spec.ts
**/fixtures/**
**/mocks/**
**/__tests__/**Best Practices
Index Once, Update Incrementally: Use
index_codebasefor initial indexing, thenreindex_changesfor updatesUse Filters: Narrow search scope with
fileTypesandpathPatternfor better resultsMeaningful Queries: Use natural language that describes what you're looking for (e.g., "database connection pooling" instead of "db")
Check Status First: Use
get_index_statusto verify a codebase is indexed before searchingLocal Embedding: Use Ollama (default) to keep everything local and private
Performance
Typical performance on a modern laptop (Apple M1/M2 or similar):
Codebase Size | Files | Indexing Time | Search Latency |
Small (10k LOC) | 50 | ~10s | <100ms |
Medium (100k LOC) | 500 | ~2min | <200ms |
Large (500k LOC) | 2,500 | ~10min | <500ms |
Note: Indexing time varies based on embedding provider. Ollama (local) is fastest for initial indexing.
Examples
See examples/ directory for detailed guides:
Basic Usage - Create collections, add documents, search
Knowledge Base - Structured documentation with metadata
Advanced Filtering - Complex boolean filters
Rate Limiting - Batch processing with cloud providers
Code Search - Index codebases and semantic code search
Advanced Configuration
Environment Variables
Core Configuration
Variable | Description | Default |
| "stdio" or "http" | stdio |
| Port for HTTP transport | 3000 |
| "ollama", "nvidia", "openai", "cohere", "voyage" | ollama |
| Qdrant server URL | |
| Path to prompts configuration JSON | prompts.json |
Embedding Configuration
Variable | Description | Default |
| Model name | Provider-specific |
| Custom API URL | Provider-specific |
| Rate limit | Provider-specific |
| Retry count | 3 |
| Initial retry delay (ms) | 1000 |
| OpenAI API key | - |
| Cohere API key | - |
| Voyage AI API key | - |
Code Vectorization Configuration
Variable | Description | Default |
| Maximum chunk size in characters | 2500 |
| Overlap between chunks in characters | 300 |
| Enable AST-aware chunking (tree-sitter) | true |
| Number of chunks to embed in one batch | 100 |
| Additional file extensions (comma-separated) | - |
| Additional ignore patterns (comma-separated) | - |
| Default search result limit | 5 |
Provider Comparison
Provider | Models | Dimensions | Rate Limit | Notes |
Ollama |
| 768, 1024, 384 | None | Local, no API key |
OpenAI |
| 1536, 3072 | 3500/min | Cloud API |
Cohere |
| 1024 | 100/min | Multilingual support |
Voyage |
| 1024, 1536 | 300/min | Code-specialized |
Note: Ollama models require docker exec ollama ollama pull <model-name> before use.
Troubleshooting
Issue | Solution |
Qdrant not running |
|
Collection missing | Create collection first before adding documents |
Ollama not running | Verify with |
Model missing |
|
Rate limit errors | Adjust |
API key errors | Verify correct API key in environment configuration |
Filter errors | Ensure Qdrant filter format, check field names match metadata |
Codebase not indexed | Run |
Slow indexing | Use Ollama (local) for faster indexing, or increase |
Files not found | Check |
Search returns no results | Try broader queries, check if codebase is indexed with |
Out of memory during index | Reduce |
Development
npm run dev # Development with auto-reload
npm run build # Production build
npm run type-check # TypeScript validation
npm test # Run test suite
npm run test:coverage # Coverage reportTesting
422 tests (376 unit + 46 functional) with 98%+ coverage:
Unit Tests: QdrantManager (21), Ollama (31), OpenAI (25), Cohere (29), Voyage (31), Factory (32), MCP Server (19)
Functional Tests: Live API integration, end-to-end workflows (46)
CI/CD: GitHub Actions runs build, type-check, and tests on Node.js 20 & 22 for every push/PR.
Contributing
Contributions welcome! See CONTRIBUTING.md for:
Development workflow
Conventional commit format (
feat:,fix:,BREAKING CHANGE:)Testing requirements (run
npm test,npm run type-check,npm run build)
Automated releases: Semantic versioning via conventional commits - feat: → minor, fix: → patch, BREAKING CHANGE: → major.
Acknowledgments
The code vectorization feature is inspired by and builds upon concepts from the excellent claude-context project (MIT License, Copyright 2025 Zilliz).
License
MIT - see LICENSE file.
This server cannot be deployed
Maintenance
Related MCP Connectors
Ingest, manage, and retrieve documents for RAG-powered AI applications
Versioned documentation registry and semantic search for AI tools and coding assistants.
Search your knowledge bases from any AI assistant using hybrid RAG.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables semantic search and document management using a local Qdrant vector database with OpenAI embeddings. Supports natural language queries, metadata filtering, and collection management for AI-powered document retrieval.101 npm36MIT
- AlicenseNot gradedqualityDmaintenanceEnables semantic search across your codebase using Google's Gemini embeddings and Qdrant Cloud vector storage. Supports 15+ programming languages with smart code chunking and real-time file change monitoring.23 npm20MIT
- FlicenseNot gradedqualityDmaintenanceEnables semantic search and retrieval-augmented generation (RAG) using Qdrant vector database. Supports indexing documents from URLs and local directories, with flexible embedding options using Ollama or OpenAI.2-
- AlicenseNot gradedqualityDmaintenanceEnables semantic search and management of documentation through vector similarity using Qdrant and Ollama/OpenAI embeddings.16 npmApache 2.0