mcp-vector-memory
Provides local embedding generation via Ollama (default qwen3-embedding) for indexing and searching code, docs, and memories, keeping data local without third-party APIs.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-vector-memorysearch all projects for our API error handling conventions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AI coding agents forget everything between sessions, and two agents working on the same project never share what they learned. mcp-vector-memory is a self-hosted Model Context Protocol server that fixes both: it indexes your code and documentation into a vector database, and gives every agent tools to search it and to store decisions, conventions and lessons learned.
Claude Code, Codex or any MCP client connects to the same server and works on the same knowledge base. Embeddings are computed locally with Ollama: no code leaves your infrastructure.
How it works
flowchart TB
subgraph Clients[" "]
direction LR
A1[Claude Code] ~~~ A2[Codex] ~~~ A3[Any MCP client] ~~~ G[Git push webhook]
end
Clients -->|MCP over HTTP · bearer token| S[Auth · rate limit]
S --> C[Chunking<br/>tree-sitter AST · text] --> E[Local embeddings<br/>Ollama]
S --> R[Hybrid search<br/>dense + sparse]
E --> Q[(Qdrant<br/>code · docs · memory)]
R <--> QRelated MCP server: Turbo Quant Memory MCP Server
Features
Code-aware indexing: tree-sitter splits source files along functions and classes, so a result is a whole unit, not half a function
Hybrid search: dense embeddings plus a sparse keyword encoder, with separate relevance thresholds for code, docs and memories
Agent memory: agents store what they learn; duplicates are detected, updates use optimistic locking, deletions are soft with 30-day retention
Local embeddings:
qwen3-embeddingthrough Ollama by default, with distinct instructions for code and textMulti-project, multi-token: each bearer token only sees the projects it is scoped to
Auto re-indexing: one webhook per repository, triggered on push
Hardened: per-token rate limiting, request size limits, security headers
MCP tools
Tool | What it does |
| Semantic and keyword search across code, docs and memories |
| Index or re-index a file |
| Store a memory (decision, convention, lesson learned), with duplicate detection |
| Update a memory, with optimistic locking |
| Soft-delete a memory (30-day retention) |
| Health and metrics of the vector store |
Quick start
git clone https://github.com/BBerthod/mcp-vector-memory.git
cd mcp-vector-memory
cp .env.example .env # set QDRANT_API_KEY (openssl rand -hex 32)
# declare your projects and a random token in config/projects.yml
docker network create dokploy-network # the compose file expects this external network
docker compose -f docker-compose.yml -f docker-compose.local.yml up -dThe first start pulls the embedding model (about 5 GB). The server then listens on http://localhost:3100.
Connect a client, for example Claude Code:
claude mcp add --transport http vector-memory http://localhost:3100/mcp \
--header "Authorization: Bearer <your-token>"Configuration
# config/projects.yml
projects:
my-app:
repo: "github.com/your-user/my-app"
server_repo_path: "/var/repos/my-app"
languages: [php, javascript, vue]
exclude: [vendor, node_modules, storage]
webhook_secret: "${MY_APP_WEBHOOK_SECRET}"
auth:
tokens:
<random-token>:
name: "me"
projects: ["my-app"]Variable | Default | Purpose |
|
| Embedding models |
|
| Embedding dimensions |
|
| Requests per token per hour |
|
| Burst allowance |
Stack
TypeScript · Node.js 22 · MCP SDK · Express · Qdrant · Ollama · tree-sitter · Docker Compose
About this repository
This server runs in production behind my own agents. The repository is a weekly snapshot of the private one it is developed in. Issues and ideas are welcome.
License
MIT. Built by Billy Berthod.
This server cannot be deployed
Maintenance
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Universal memory for AI agents and tools. Save, organize and search context anywhere.
Persistent memory for AI agents. Search and store durable facts, preferences and decisions.
Persistent memory for AI agents. Semantic search, memory graph, W3C DID identity.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides AI coding assistants with persistent project memory to retain architectural decisions, code patterns, and domain knowledge across sessions. It stores data locally in a SQLite database, allowing agents to remember, recall, and manage project-specific context using full-text search.8 npmApache 2.0
- AlicenseNot gradedqualityAmaintenanceProvides persistent, local-first memory with knowledge graph and hybrid search for AI coding agents, reducing token usage by storing decisions, patterns, and codebase context.9MIT
- AlicenseNot gradedqualityDmaintenanceProvides long-term memory for AI coding agents, enabling them to remember, search, and organize information across sessions and platforms like Claude Code, ChatGPT, and Cursor.8 npm8MIT
- AlicenseAqualityBmaintenanceProvides persistent, searchable memory across AI coding agent and chat history (Claude Code, Codex, Gemini CLI, ChatGPT, and more) via retrieval-augmented generation, enabling semantic and hybrid search to retain context across sessions.5MIT