Skip to main content
Glama
README.md
# DocuMind MCP — Internal Docs RAG Server

**DocuMind MCP** is a citation-grounded RAG (Retrieval-Augmented Generation) system built natively on the Model Context Protocol (MCP). It exposes your organization's internal knowledge base (engineering runbooks, HR policies, onboarding documentation) directly to any MCP client (Claude Desktop, Claude Code, custom Slack agents) with zero custom integration code.

---

## 1. Problem Definition (The Value Proposition)

Before MCP, connecting AI tools to internal data sources suffered from the **N×M integration problem**:
If you have **N** AI clients (Claude Desktop, internal CLI, Slack bots) and **M** data sources (Confluence, local markdown, policy wikis), you must write and maintain **N×M** bespoke integrations. Every client ends up implementing its own authentication, document parsing, retrieval logic, and citation rendering.

**With DocuMind MCP:**
- The knowledge base is exposed once as a unified MCP Server.
- **Every MCP-speaking client gets access for free**, with zero additional integration code.
- The server owns retrieval quality, grounding thresholds, and citation metadata exactly once.

### Protocol Philosophy: Retrieval as a Primitive, Not a Black Box
Unlike typical "MCP wrappers" that expose a single opaque `ask_question` tool containing a hidden LLM call, DocuMind MCP exposes **retrieval and structured context as primitives**:
- **Tools**: `search_docs` performs semantic vector searches and returns structured, similarity-thresholded context snippets with source details.
- **Resources**: `docs://catalog` and `docs://document/{doc_id}/{section}` expose browsable, addressable sections of documents.
This separation leaves the final reasoning and citation synthesis to the calling client's LLM—preserving the reasoning transparency that MCP is built for.

---

## 2. Complete System Architecture

```text
┌───────────────────┐  ┌───────────────────┐  ┌───────────────────┐
│   Claude Desktop  │  │    Claude Code    │  │   Custom Agent    │
└─────────┬─────────┘  └─────────┬─────────┘  └─────────┬─────────┘
          │                      │                      │
          │   MCP over HTTP (+ SSE handshake), all clients hit same protocol
          └──────────────────────┼──────────────────────┘
                                 │
                     ┌───────────▼───────────┐
                     │    Auth Middleware    │
                     │ - API key validation  │
                     │ - Rate limiting       │
                     └───────────┬───────────┘
                                 │
                     ┌───────────▼───────────┐
                     │    MCP Server Core    │
                     │  (Tools & Resources)  │
                     └─────────┬───────┬─────┘
                               │       │
             ┌─────────────────▼─┐   ┌─▼─────────────────┐
             │ Tools:            │   │ Resources:        │
             │ - search_docs     │   │ - docs://catalog  │
             │ - get_doc         │   │ - docs://document │
             └─────────────────┬─┘   └─┬─────────────────┘
                               │       │
                  ┌────────────▼───────▼─────┐
                  │    RAG Retrieval Core    │
                  │   - Local Embeddings     │
                  │   - Similarity filter    │
                  └────────────┬─────────────b
                               │
                      ┌────────▼────────┐
                      │   Vector DB     │
                      │   (Qdrant)      │
                      └─────────────────┘
```

---

## 3. Data Schema & Models

### SQLite (Metadata, Logs & Auth)
1. **`api_clients`**: Stores client credentials (never raw API keys, only SHA-256 hashes), rate limits, and revocation states.
2. **`documents`**: Tracks document metadata, connector sources (`filesystem`, `confluence`, `git`), and `last_synced_at` timestamps.
3. **`mcp_request_log`**: Structured audit log detailing client requests, latency, endpoints called, and summary stats.

### Qdrant (`document_chunks` collection)
- **Vector**: 384 dimensions (`BAAI/bge-small-en-v1.5` dense model).
- **Payload**: `chunk_id` (deterministic UUIDv5), `doc_id`, `doc_title`, `section`, `text`, `source_connector`.

---

## 4. Getting Started

### Local Setup
1. **Initialize virtual environment & install requirements**:
   ```bash
   python -m venv .venv
   .venv\Scripts\activate   # Windows
   source .venv/bin/activate # Unix
   pip install -r requirements.txt
   ```

2. **Run Server**:
   ```bash
   python -m uvicorn app.main:app --port 8000
   ```
   *On first run, the server automatically initializes SQLite and seeds the in-memory Qdrant instance with sample Confluence documentation. It will print a newly generated default API Key to the console.*

### Docker Deployment
```bash
docker compose up --build
```
This spins up both the FastAPI MCP Server and a persistent Qdrant instance running on local port `6333`.

---

## 5. Integrating with Claude Desktop / Claude Code

To add this server to your Claude client, configure the SSE transport by pointing to the server endpoint with the generated API key:

### `claude_desktop_config.json`
```json
{
  "mcpServers": {
    "documind-rag": {
      "command": "npx",
      "args": [
        "-y",
        "@modelcontextprotocol/inspector",
        "http://localhost:8000/mcp/sse?api_key=YOUR_GENERATED_KEY"
      ]
    }
  }
}
```

---

## 6. Verification and Testing

### Unit Tests
To run all tests (authentication, rate-limiting, and retrieval filtering):
```bash
python -m pytest
```

### Self-Contained Handshake Test
Run the client script to simulate a complete protocol-level SSE handshake and JSON-RPC query cycle:
```bash
python tests/test_client.py
```