Skip to main content
Glama
README.md
# qdrant-mcp-server (maintained fork)

Fork of the abandoned PyPI package `qdrant-mcp-server` 0.1.0 (author fiyen;
upstream github.com/fiyen/qdrant-mcp-server is 404 — no repo exists to PR against).

## Why this fork exists

The original eagerly loads a `sentence-transformers/all-MiniLM-L6-v2` model in
`__init__`. As a **stdio MCP server**, every agent runtime/session spawns its own
copy, so N idle sessions pin N × ~1.35 GB of duplicated model memory (measured:
5 copies = 5.7 GB). See the generalized "stdio trap" write-up.

## Changes vs PyPI 0.1.0

- **Lazy-load patch**: the embedding model loads on the first
  `generate_embedding` call instead of at startup. Idle copies cost ~30 MB.
  (`lazy-load.patch` in the repo root is the diff for reference.)

## Install (replaces the PyPI install)

    uv tool install --force git+https://github.com/Ripnrip/qdrant-mcp-server

Then the existing MCP configs (`command = /Users/admin/.local/bin/qdrant-mcp-server`)
keep working, and future `uv tool upgrade` pulls from this fork, not dead PyPI 0.1.0.

After any reinstall, restart running qdrant-mcp-server instances (the Swift reaper
LaunchAgent handles orphan cleanup: com.local.qdrant-mcp-reaper).

Context: HAB-354 · ~/Documents/Developer/qdrant-mcp-ram-fix-2026-08-24/