genai-mcp-assistant
by mykabowd
README.md
π«π· [Version franΓ§aise](README.fr.md)
# genai-mcp-assistant
A RAG + agent + **MCP server** assistant that answers technical questions
about **DevOps / Kubernetes / Docker** using a local knowledge base:
core concepts, networking pitfalls, and a personal troubleshooting journal.
Built to concretely apply, on a real use case, a modern GenAI stack:
- **Python** (backend, data pipeline, MCP server)
- **RAG** (Retrieval-Augmented Generation): document chunking,
**vector embeddings**, a vector database (**ChromaDB**), semantic search
- **Local LLM** via **Ollama** (no external API key required)
- **LangChain / LangGraph** for the RAG chain and an **agent** that decides
on its own when to query the knowledge base (agentic AI)
- **MCP server** (official `mcp` SDK) exposing the assistant as
**MCP tools** usable by any MCP client (Cursor, Claude Desktop, etc.)
- **Docker** for containerization
- **Kubernetes / Helm** for deployment (full chart included)
- **Automated tests** (pytest), CI-ready
## Why this project
This project grew out of wanting to get hands-on with a modern
GenAI/agentic stack β RAG, embeddings, a local LLM, LangChain/LangGraph,
and an MCP server β by building something that actually works end to end,
rather than just reading about it. Kubernetes/Docker troubleshooting
turned out to be a good playground for this: there's a mix of stable
reference concepts and narrower, concrete troubleshooting knowledge that
genuinely benefits from being looked up rather than guessed by the LLM
alone.
It complements the two other projects in this portfolio by exploring a
different technical angle β retrieval-augmented generation, agents, and
generative AI β where [`ml-critical-systems-lab`](../ml-critical-systems-lab)
focuses on classic data science and
[`engine-health-platform`](../engine-health-platform) on full-stack
development and real deployment.
The knowledge base is intentionally generic β it is not tied to any
specific company, job application, or individual. It's meant to be shown
as a standalone technical project.
## Architecture
```
ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
β data/*.txt β ---> β ingest.py β ---> β ChromaDB β
β (K8s/Docker docs,β β chunking + β β (vectorstore/) β
β troubleshooting)β β embeddings β ββββββββββ¬ββββββββββ
ββββββββββββββββββββ ββββββββββββββββββββ β
β semantic
βΌ search
ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
β rag.py β <--> β agent.py β β mcp_server.py β
β retrieval + LLM β β LangGraph ReAct β β FastMCP tools β
β (Ollama) β β (agentic loop) β β (stdio / HTTP) β
ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
```
## Quickstart (local)
Prerequisites: [Ollama](https://ollama.com) installed and running
(`brew install ollama && brew services start ollama`), with the following
models pulled:
```bash
ollama pull llama3.2:3b
ollama pull nomic-embed-text
```
Then:
```bash
make venv install # create the venv and install dependencies
make ingest # chunk the knowledge base, compute embeddings, persist to ChromaDB
make test # run the automated test suite (pytest)
```
### Query the assistant via the agent (CLI)
```bash
make agent
# or with a custom question:
.venv/bin/python -m genai_mcp_assistant.agent "How do I debug a Kubernetes pod stuck in CrashLoopBackOff?"
```
### Run the MCP server
```bash
make mcp # stdio transport, to plug into an MCP client (Cursor, Claude Desktop, ...)
```
Tools exposed (`mcp_server.py`):
| Tool | Description |
|---|---|
| `list_knowledge_topics` | Lists the ingested knowledge base files |
| `search_docs(query, k)` | Raw semantic search (no generation) |
| `ask_devops_question(question)` | Full RAG pipeline (retrieval + generated answer) |
| `search_troubleshooting_journal(symptom)` | Targeted search in the personal troubleshooting journal, from a symptom description |
## Containerized deployment
```bash
make docker-up # runs Ollama + the MCP server (HTTP transport) via docker-compose
```
## Kubernetes deployment (Helm)
A complete Helm chart is provided in `helm/genai-mcp-assistant/`:
```bash
helm install genai-mcp-assistant ./helm/genai-mcp-assistant \
--set image.repository=<your-registry>/genai-mcp-assistant \
--set image.tag=0.1.0
```
The chart deploys: a `Deployment`, a `Service`, a `ConfigMap` (env vars),
and a `PersistentVolumeClaim` for the vector store. Ollama is assumed to
run as a separate service in the cluster (`OLLAMA_BASE_URL` configurable
via `values.yaml`).
## Project structure
```
genai-mcp-assistant/
βββ data/
β βββ knowledge_base/ # Kubernetes/Docker docs + troubleshooting journal (RAG knowledge base)
βββ src/genai_mcp_assistant/
β βββ config.py # centralized configuration (env vars)
β βββ ingest.py # loading + chunking + embeddings -> ChromaDB
β βββ rag.py # retrieval + generation (LangChain + Ollama)
β βββ agent.py # LangGraph agent (agentic AI, tool-use)
β βββ mcp_server.py # MCP server (FastMCP) exposing the tools
βββ tests/ # automated tests (pytest)
βββ helm/genai-mcp-assistant/ # Helm chart (Deployment, Service, ConfigMap, PVC)
βββ Dockerfile
βββ docker-compose.yml
βββ Makefile
```
## What this project demonstrates (and what it doesn't)
**Demonstrates**: end-to-end Python development, RAG, vector embeddings,
semantic search, LLM orchestration (LangChain/LangGraph), designing an
agent that autonomously decides when to use its tools, building an MCP
server exposing MCP tools, Docker containerization, writing a Helm chart
for Kubernetes, automated testing.
**Does NOT demonstrate**: training or fine-tuning ML/deep-learning models
(the LLM and the embedding model are used purely at inference time,
pre-trained and served via Ollama), classical statistical analysis, data
visualization, or scientific publication work. It's a focused,
infrastructure-and-orchestration-layer project, not a machine-learning-
research project.
## Known limitations / possible improvements
- `llama3.2:3b` is deliberately lightweight to run without a GPU; answer
quality (and citation accuracy) improves noticeably with a larger model.
- No automated answer-quality re-evaluation yet (e.g. RAGAS or similar).
- The LangGraph agent only has one tool for now; adding a second tool
(e.g. a structured lookup against a small severity/runbook table) would
better illustrate genuine multi-step agentic reasoning.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues