Skip to main content
Glama

Ollqd — MCP Client-Server RAG System

Local-first RAG system that indexes codebases and documents into Qdrant using Ollama embeddings. Exposes everything through MCP (Model Context Protocol) so AI assistants can search your code via tool-calling.

Architecture

┌──────────────────────────────────────────────────────────────────┐
│  User Interface                                                  │
│  ┌─────────────┐  ┌─────────────────────────────────────────┐   │
│  │ ollqd-chat   │  │ Claude Desktop / any MCP host           │   │
│  │ (CLI / REPL) │  │ (connects to ollqd-server directly)     │   │
│  └──────┬───────┘  └────────────────┬────────────────────────┘   │
└─────────┼──────────────────────────┼────────────────────────────┘
          │ stdio JSON-RPC           │ stdio JSON-RPC
┌─────────▼──────────────────────────▼────────────────────────────┐
│  Ollqd MCP Server (FastMCP)                                     │
│  ┌───────────────┐ ┌─────────────────┐ ┌─────────────────────┐  │
│  │index_codebase │ │index_documents  │ │semantic_search      │  │
│  │index docs     │ │markdown/text/rst│ │embed query → Qdrant │  │
│  └───────┬───────┘ └────────┬────────┘ └──────────┬──────────┘  │
│  ┌───────┴──────┐  ┌───────┴────────┐             │             │
│  │list_collections│ │delete_collection│             │             │
│  └──────────────┘  └────────────────┘             │             │
└─────────┬──────────────────────────────────────────┼────────────┘
          │ /api/embed                               │
┌─────────▼──────────┐                    ┌──────────▼─────────┐
│  Ollama             │                    │  Qdrant             │
│  nomic-embed-text   │                    │  cosine similarity  │
│  + chat models      │                    │  payload indexes    │
└────────────────────┘                    └────────────────────┘

How it works

  1. Discovery — Walks the codebase, filters by language (40+ extensions), skips lock files / build artifacts / vendor dirs.

  2. Code-aware chunking — Splits files at natural code boundaries (function defs, class declarations, impl blocks) rather than blindly cutting at token limits. Overlapping windows preserve context.

  3. Embedding — Sends chunks to Ollama's /api/embed in batches. Each chunk is prefixed with file path + language + line range for better semantic grounding.

  4. Storage — Upserts into Qdrant with full metadata payload. Payload indexes on file_path, language, and content_hash enable filtered search and incremental re-indexing.

  5. RAG loop — The client sends user queries to Ollama with MCP tools attached. Ollama decides when to call semantic_search, gets results from the server, and synthesizes a final answer with code citations.

Related MCP server: RagDocs MCP Server

Setup

Prerequisites

  • Ollama running locally with an embedding model pulled

  • Qdrant running (Docker recommended)

  • Python 3.10+

# Pull the embedding model
ollama pull nomic-embed-text

# Pull a chat model (any that supports tool-calling)
ollama pull qwen2.5:14b

# Start Qdrant (and optionally Ollama via Docker)
docker compose up -d

Install

# With uv (recommended)
uv venv && source .venv/bin/activate
uv pip install -e ".[client,dev]"

# Or with pip
pip install -e ".[client,dev]"

Usage

Start the MCP server (standalone)

ollqd-server

The server communicates over stdio using JSON-RPC (MCP protocol). It's meant to be launched by MCP clients, not used directly.

Interactive RAG chat

# Interactive REPL — ask questions about your codebase
ollqd-chat --interactive

# Single query
ollqd-chat "how does the auth middleware work?"

# Use a different chat model
ollqd-chat --interactive --model llama3.1

# Debug mode
ollqd-chat -v "find the database connection setup"

REPL commands:

  • :quit / :q — exit

  • :model <name> — switch chat model on the fly

Use with Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "ollqd": {
      "command": "ollqd-server",
      "args": []
    }
  }
}

Then in Claude Desktop, ask things like:

  • "Index my project at /path/to/codebase"

  • "Search for how authentication is implemented"

  • "What error handling patterns are used?"

  • "List all indexed collections"

MCP Tools

Tool

Description

index_codebase

Walk + chunk + embed + upsert code files from a directory

index_documents

Chunk + embed + upsert document files (markdown, text, rst)

semantic_search

Embed a natural language query and search Qdrant

list_collections

List all Qdrant collections with point counts

delete_collection

Drop a collection (requires confirm=true)

Configuration

Environment variables

Variable

Default

Description

OLLAMA_URL

http://localhost:11434

Ollama base URL

QDRANT_URL

http://localhost:6333

Qdrant REST URL

OLLAMA_CHAT_MODEL

qwen2.5:14b

Chat model for RAG

OLLAMA_EMBED_MODEL

nomic-embed-text

Embedding model

OLLAMA_TIMEOUT_S

120

Request timeout (seconds)

CHUNK_SIZE

512

Approximate tokens per chunk

CHUNK_OVERLAP

64

Overlap tokens between chunks

MAX_TOOL_ROUNDS

6

Max tool-calling rounds per query

ollqd.toml

[ollama]
host = "http://localhost:11434"
chat_model = "qwen2.5:14b"
embed_model = "nomic-embed-text"
timeout = 120

[qdrant]
host = "http://localhost:6333"
default_collection = "codebase"

[indexing]
chunk_size = 512
chunk_overlap = 64
max_file_size_kb = 512

[server]
name = "ollqd-rag-server"
transport = "stdio"

[client]
max_tool_rounds = 6

Project structure

src/ollqd/
├── __init__.py
├── config.py          # AppConfig dataclass + env var overrides
├── errors.py          # Exception hierarchy
├── models.py          # FileInfo, Chunk, SearchResult, IndexingStats
├── chunking.py        # Code-aware + document chunking
├── discovery.py       # File discovery (40+ languages)
├── embedder.py        # OllamaEmbedder wrapping /api/embed
├── vectorstore.py     # QdrantManager (upsert, search, incremental)
├── server/
│   └── main.py        # FastMCP server with 5 tools
└── client/
    ├── mcp_bridge.py  # MCP session over stdio
    ├── ollama_agent.py # Ollama chat with tool-calling
    ├── rag_loop.py    # RAG loop runner
    └── main.py        # CLI entry point

Supported languages

Python, Go, JavaScript, TypeScript, Rust, Java, Kotlin, Scala, C, C++, C#, Ruby, PHP, Swift, Lua, Shell, SQL, R, HTML, CSS, SCSS, YAML, TOML, JSON, Markdown, reStructuredText, Terraform, HCL, Dockerfile, Protobuf, GraphQL.

Embedding models

Any Ollama model that supports /api/embed works. Recommended:

Model

Dimensions

Notes

nomic-embed-text

768

Good balance of quality and speed (default)

mxbai-embed-large

1024

Higher quality, slower

all-minilm

384

Fast, smaller footprint

snowflake-arctic-embed

1024

Strong code understanding

Design decisions

Why MCP? — The Model Context Protocol lets any compatible AI assistant (Claude Desktop, custom clients, IDE extensions) use ollqd's indexing and search tools without custom integration code.

Why not tree-sitter for chunking? — Tree-sitter gives perfect AST-based splits but adds a heavy dependency per language. The heuristic boundary detection covers ~90% of cases with zero extra setup.

Why deterministic point IDs?md5(file_path::chunk_N) means re-indexing the same file overwrites existing points instead of creating duplicates. This makes incremental mode reliable.

Why prefix chunks with metadata? — Embedding models produce better vectors when given context. "File: auth/middleware.go | Language: go | Lines 45-82" followed by the code produces more semantically meaningful vectors.

Legacy scripts

The standalone scripts from v0.1 are still available:

# Bulk index (standalone, no MCP)
python codebase_indexer.py /path/to/project --collection myproject

# Search (standalone, no MCP)
python codebase_search.py "auth middleware" --interactive

See DESIGN.md for the full architecture document with diagrams, security analysis (STRIDE), and detailed API reference.

Available Tools

5 tools
delete_collectionC

Delete a Qdrant collection. Set confirm=true to proceed.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo
collectionYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so description carries full burden. It only states deletion and the confirm requirement, but misses irreversible nature, permissions, side effects, and error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise - two short sentences. No unnecessary information, front-loaded with the action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with no annotations or output schema, the description is severely incomplete. Lacks details on irreversibility, data loss, prerequisites, and what happens after deletion.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description must explain parameters. It only mentions 'confirm' meaning to proceed, but does not explain the required 'collection' parameter at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Delete a Qdrant collection' - a specific verb+resource. It distinguishes from siblings like index_codebase, list_collections, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions setting confirm=true to proceed, but gives no guidance on when to use this tool vs alternatives, or any prerequisites/when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_codebaseB

Index a codebase directory into Qdrant. Walks files, chunks at code boundaries, embeds via Ollama, upserts to Qdrant.

ParametersJSON Schema
NameRequiredDescriptionDefault
root_pathYes
chunk_sizeNo
collectionNocodebase
incrementalNo
chunk_overlapNo
extra_skip_dirsNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description provides some behavioral context (walks files, chunks, embeds, upserts), but misses details on side effects (e.g., overwriting existing data), authorization needs, or idempotency. The term 'upserts' implies potential overwrite but not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences that pack the core process. No superfluous words, front-loading the key action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is incomplete given no output schema and 6 unannotated parameters. It lacks return value description, error handling, parameter defaults, and any mention of the incremental flag or extra_skip_dirs. Users would need external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any parameter meanings. It mentions 'codebase directory' (likely root_path) and 'chunks' (chunk_size), but fails to map clearly to all six parameters, leaving significant gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Index a codebase directory into Qdrant' and outlines the pipeline. It distinguishes from sibling tools like index_documents and semantic_search by specifying 'codebase directory' and 'chunks at code boundaries'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lacks explicit guidance on when to use this tool versus alternatives. It does not mention prerequisites, when not to use it, or refer to sibling tools for other purposes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_documentsC

Index document files (markdown, text, etc.) into Qdrant.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
chunk_sizeNo
collectionNodocuments
source_tagNodocs
chunk_overlapNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must reveal behavioral traits. It does not mention if indexing overwrites existing data, requires authentication, or handles errors. The chunking behavior is implied but not explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but it is too brief given the tool's complexity. Conciseness sacrifices necessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, no output schema, and no annotations, the description lacks critical information about return values, error handling, and behavior beyond the literal indexing action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for parameters, and the description adds no meaning beyond names and defaults. It fails to explain what 'chunk_size' or 'collection' control.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Index' and resource 'document files (markdown, text, etc.)' into Qdrant. It effectively distinguishes from sibling tools like delete_collection or semantic_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool vs alternatives, such as index_codebase for code files. No prerequisites or conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_collectionsA

List all Qdrant collections with point counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates a read-only operation with no side effects, which is accurate. With no annotations, it fully conveys the tool's behavior for a simple list action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, complete sentence with no unnecessary words. It efficiently communicates the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, no output schema, and no annotations, the description fully covers the tool's capability. There is no missing information for its intended use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, so schema coverage is 100%. The description adds no parameter info, but baseline for zero parameters is 4, as no additional meaning is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists all Qdrant collections with point counts. It uses a specific verb ('list') and resource ('collections'), and distinguishes from sibling tools like delete_collection or semantic_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use versus alternatives. For a straightforward listing tool, the purpose implicitly suggests use for overview, but explicit context would improve clarity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.3.0
    • First observeddelete_collection
    • First observedindex_codebase
    • First observedindex_documents
    • First observedlist_collections
    • First observedsemantic_search

TDQS

A3.5/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct operation: deleting collections, indexing codebases, indexing documents, listing collections, and searching. No overlap in functionality; descriptions clearly differentiate them.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (e.g., delete_collection, index_codebase, list_collections). Even 'semantic_search' fits as an adjective-noun pair, maintaining uniformity.

Tool Count5/5

With 5 tools, the server is well-scoped for its purpose of managing Qdrant collections and indexing/searching content. Each tool serves a clear, necessary role without bloat.

Completeness4/5

The tool surface covers key operations: listing, deleting, indexing (two types), and searching. However, it lacks a dedicated create_collection tool (indexing might auto-create, but not explicit) and no tool to remove individual indexed points, which is a minor gap.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and retrieval-augmented generation (RAG) using Qdrant vector database. Supports indexing documents from URLs and local directories, with flexible embedding options using Ollama or OpenAI.
    2
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    A semantic codebase indexer MCP server that chunks source code, generates embeddings via Ollama, and stores them in Qdrant for natural-language code search.
    26 npm
    1
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic code search over a local codebase using Qdrant vector embeddings and OpenAI embeddings, allowing natural language queries from MCP-compatible clients like Claude Desktop.
    -