Skip to main content
Glama
ansh-7666

RAG Chatbot MCP Server

by ansh-7666
README.md
# RAG Chatbot

A small numpy-based vector index (`vector_store.py`) over your own documents — no compiled/native
dependencies beyond numpy, so it needs no admin rights to install. The LLM/embeddings backend is
pluggable via `config.LLM_PROVIDER`: `"ollama"` (100% local, no API key) or `"gemini"` (Google's
cloud API, needs a free key).

## Setup

**Provider: Gemini** (`config.py` → `LLM_PROVIDER = "gemini"`, the current default)

1. Get a free API key from [Google AI Studio](https://aistudio.google.com/apikey).
2. Set it as an environment variable (don't paste it into files or commit it):

```bash
# macOS/Linux
export GOOGLE_API_KEY="your-key-here"
```

```powershell
# Windows PowerShell
$env:GOOGLE_API_KEY = "your-key-here"
```

**Provider: Ollama** (set `LLM_PROVIDER = "ollama"` in `config.py`)

1. Install [Ollama](https://ollama.com/download) and pull the models used:

```bash
ollama pull nomic-embed-text
ollama pull llama3.1
```

Then, for either provider, install Python dependencies:

```bash
pip install -r requirements.txt
```

## Usage — CLI

1. Drop your files (`.pdf`, `.txt`, `.md`, `.docx`, `.csv`, `.xlsx`) into the `docs/` folder.
2. Build the index:

```bash
python ingest.py
```

3. Chat:

```bash
python chat.py
```

Type `exit` to quit.

## Usage — Web app (frontend + backend)

Runs a FastAPI backend that serves a JSON API and a static chat UI, all on one port.

```bash
python -m uvicorn server:app --reload --port 8000
```

Open http://localhost:8000 in a browser. From there you can:
- Upload files (drag/select, click **Upload**)
- Click **Rebuild Index** to (re)embed everything currently in `docs/`
- Chat in the main panel — answers include source file names

API endpoints, if you want to script against it directly:
- `GET  /api/status` — index/model info
- `POST /api/upload` — multipart file upload, saved into `docs/`
- `POST /api/ingest` — rebuilds the index from `docs/`
- `POST /api/chat` — `{"question": "..."}` → `{"answer": "...", "sources": [...]}`

## Usage — MCP server

Exposes the document index as MCP tools (`ask`, `search`, `rebuild_index`, `status`) so any
MCP client (Claude Desktop, Claude Code, etc.) can query your docs. Runs over stdio - the
client launches it as a subprocess, no port involved.

Add it to your MCP client config, e.g. Claude Desktop's `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "rag-chatbot": {
      "command": "python",
      "args": ["C:/Users/Nikita_Admin/Desktop/mcm/rag-chatbot/mcp_server.py"]
    }
  }
}
```

For Claude Code, run:

```bash
claude mcp add rag-chatbot -- python C:/Users/Nikita_Admin/Desktop/mcm/rag-chatbot/mcp_server.py
```

Restart the client afterward. The index must already exist (`python ingest.py`), or call the
`rebuild_index` tool from within the chat once files are in `docs/`.

## Notes

- Re-run `python ingest.py` after adding/changing files in `docs/`. It rebuilds `index.npz` from scratch each time.
- Change models or chunking behavior in `config.py`.
- Larger/more capable local models (e.g. `llama3.1:70b`, `mixtral`) give better answers but need more RAM/VRAM — swap `LLM_MODEL` in `config.py`.
- The vector index is a single `index.npz` file (numpy arrays + JSON), fine for personal/small document sets. For large corpora, swap `vector_store.py` for a proper vector DB.