RAG Chatbot MCP Server
by ansh-7666
README.md
# RAG Chatbot
A small numpy-based vector index (`vector_store.py`) over your own documents — no compiled/native
dependencies beyond numpy, so it needs no admin rights to install. The LLM/embeddings backend is
pluggable via `config.LLM_PROVIDER`: `"ollama"` (100% local, no API key) or `"gemini"` (Google's
cloud API, needs a free key).
## Setup
**Provider: Gemini** (`config.py` → `LLM_PROVIDER = "gemini"`, the current default)
1. Get a free API key from [Google AI Studio](https://aistudio.google.com/apikey).
2. Set it as an environment variable (don't paste it into files or commit it):
```bash
# macOS/Linux
export GOOGLE_API_KEY="your-key-here"
```
```powershell
# Windows PowerShell
$env:GOOGLE_API_KEY = "your-key-here"
```
**Provider: Ollama** (set `LLM_PROVIDER = "ollama"` in `config.py`)
1. Install [Ollama](https://ollama.com/download) and pull the models used:
```bash
ollama pull nomic-embed-text
ollama pull llama3.1
```
Then, for either provider, install Python dependencies:
```bash
pip install -r requirements.txt
```
## Usage — CLI
1. Drop your files (`.pdf`, `.txt`, `.md`, `.docx`, `.csv`, `.xlsx`) into the `docs/` folder.
2. Build the index:
```bash
python ingest.py
```
3. Chat:
```bash
python chat.py
```
Type `exit` to quit.
## Usage — Web app (frontend + backend)
Runs a FastAPI backend that serves a JSON API and a static chat UI, all on one port.
```bash
python -m uvicorn server:app --reload --port 8000
```
Open http://localhost:8000 in a browser. From there you can:
- Upload files (drag/select, click **Upload**)
- Click **Rebuild Index** to (re)embed everything currently in `docs/`
- Chat in the main panel — answers include source file names
API endpoints, if you want to script against it directly:
- `GET /api/status` — index/model info
- `POST /api/upload` — multipart file upload, saved into `docs/`
- `POST /api/ingest` — rebuilds the index from `docs/`
- `POST /api/chat` — `{"question": "..."}` → `{"answer": "...", "sources": [...]}`
## Usage — MCP server
Exposes the document index as MCP tools (`ask`, `search`, `rebuild_index`, `status`) so any
MCP client (Claude Desktop, Claude Code, etc.) can query your docs. Runs over stdio - the
client launches it as a subprocess, no port involved.
Add it to your MCP client config, e.g. Claude Desktop's `claude_desktop_config.json`:
```json
{
"mcpServers": {
"rag-chatbot": {
"command": "python",
"args": ["C:/Users/Nikita_Admin/Desktop/mcm/rag-chatbot/mcp_server.py"]
}
}
}
```
For Claude Code, run:
```bash
claude mcp add rag-chatbot -- python C:/Users/Nikita_Admin/Desktop/mcm/rag-chatbot/mcp_server.py
```
Restart the client afterward. The index must already exist (`python ingest.py`), or call the
`rebuild_index` tool from within the chat once files are in `docs/`.
## Notes
- Re-run `python ingest.py` after adding/changing files in `docs/`. It rebuilds `index.npz` from scratch each time.
- Change models or chunking behavior in `config.py`.
- Larger/more capable local models (e.g. `llama3.1:70b`, `mixtral`) give better answers but need more RAM/VRAM — swap `LLM_MODEL` in `config.py`.
- The vector index is a single `index.npz` file (numpy arrays + JSON), fine for personal/small document sets. For large corpora, swap `vector_store.py` for a proper vector DB.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues