MCP RAG Agent
by elandy
README.md
# MCP RAG Agent
A local-first AI backend that combines **Model Context Protocol**, **Retrieval-Augmented Generation**, and an agent runtime to answer questions from a private knowledge base.
The project runs an MCP server that exposes retrieval tools. A client-side agent connects to that server, lets the LLM decide when to call tools, retrieves relevant knowledge chunks, and returns grounded answers with source citations.
## Features
* MCP server exposing retrieval tools
* CLI agent client
* Interactive chat mode
* RAG pipeline over a local knowledge base
* ChromaDB vector store
* Pluggable LLM providers (Ollama, AWS Bedrock)
* Pluggable embedding providers (Ollama, AWS Bedrock)
* Optional reranking pipeline
* Source-grounded answers with citations
* Modular service layout for retrieval, embeddings, chunking, vector stores, and rerankers
## Example Usage
Ask a single question:
```bash
uv run python -m client.src.main "Explain how MCP tools work with RAG"
```
Start interactive mode:
```bash
uv run python -m client.src.main --interactive
```
Example interactive session:
```text
MCP RAG Agent
Type a question and press Enter.
Type 'exit' or 'quit' to stop.
You: Explain how MCP tools work with RAG
Assistant:
MCP tools allow the agent to access external capabilities, such as document retrieval, through a standard protocol...
Sources:
- knowledge/mcp.md#3
- knowledge/rag.md#1
```
## Architecture
```mermaid
flowchart TD
User[User] --> Agent[LangGraph Agent]
Agent --> Runtime[Agent Runtime]
Runtime --> LLM["LLM Provider<br/>(Ollama / Bedrock)"]
Runtime --> MCPClient[MCP Client]
MCPClient --> MCPServer[MCP Server]
MCPServer --> RetrievalTool["retrieve_documents()"]
MCPServer --> OtherTools["Other MCP Tools"]
RetrievalTool --> Embed["Embedding Provider<br/>(Ollama / Bedrock)"]
Embed --> VectorStore[ChromaDB Vector Store]
VectorStore --> Reranker[Reranker]
Reranker --> Context[Context Builder]
Context --> RetrievalTool
RetrievalTool --> MCPServer
MCPServer --> MCPClient
Runtime --> Answer[Grounded Answer with Citations]
```
## How It Works
1. The user asks a question through the CLI.
2. The client starts and connects to the MCP server.
3. The agent discovers available MCP tools.
4. The LLM decides whether it needs to call a retrieval tool.
5. The MCP server runs the RAG pipeline:
* embeds the query
* searches ChromaDB
* optionally reranks retrieved chunks
* builds a context payload
* returns source metadata
6. The agent generates a final answer using the retrieved context.
7. The final response includes citations such as:
```text
[knowledge/mcp.md#2]
```
## Project Structure
```text
client/
└── src/
├── main.py # CLI entry point
├── agent_runtime.py # Agent loop and tool-calling runtime
├── mcp_client.py # MCP client wrapper
├── tool_executor.py # Executes MCP tools
├── tool_mapper.py # Converts MCP tools to Ollama tool schema
└── llm/ # LLM provider implementations
├── base.py
├── ollama.py
└── bedrock.py
server/
└── src/
├── server.py # MCP server entry point
└── tools/ # MCP tool registrations
services/
├── rag/
├── chunking/ # Text chunking strategies
├── dto/ # API response DTOs
├── embedding/ # Embedding providers
| ├── base.py
| ├── ollama_provider.py
| └── bedrock_provider.py
├── loaders/ # Document loaders
├── mappers/ # Domain-to-response mappers
├── models/ # RAG domain models
├── pipelines/ # Indexing and retrieval pipelines
├── rerankers/ # Reranking implementations
└── vectorstores/ # Vector store adapters
knowledge/ # Local knowledge base documents
scripts/ # Utility scripts
chroma_db/ # Local ChromaDB data
```
## Requirements
Required
* Python 3.13+
* uv
Optional
* Ollama (for local inference)
* AWS credentials with Bedrock access (for Bedrock providers)
Example:
```bash
ollama pull llama3.1
```
If your embedding provider uses a separate embedding model, pull that model too.
## Setup
Install dependencies:
```bash
uv sync
```
Start Ollama:
```bash
ollama serve
```
Run a question:
```bash
uv run python -m client.src.main "Explain how MCP tools work with RAG"
```
Run interactive chat:
```bash
uv run python -m client.src.main --interactive
```
## CLI Options
```bash
uv run python -m client.src.main --help
```
Available options:
```text
question Optional question to ask the agent
-i, --interactive Start an interactive chat session
--model Ollama chat model to use
--max-steps Maximum number of tool-calling steps
```
Examples:
```bash
uv run python -m client.src.main "What is AWS Bedrock?" --model llama3.1
```
```bash
uv run python -m client.src.main --interactive --max-steps 8
```
## Source Citations
The retrieval tool returns source metadata for every retrieved chunk:
```json
{
"context": "...",
"sources": [
{
"document_id": "knowledge/mcp.md",
"chunk_index": 2
}
]
}
```
The agent is instructed to cite retrieved sources in the final answer using this format:
```text
[document_id#chunk_index]
```
Example:
```text
MCP lets an agent call external tools through a standard protocol, which makes the RAG system easier to separate from the LLM runtime [knowledge/mcp.md#2].
Sources:
- knowledge/mcp.md#2
- knowledge/rag.md#4
```
## Why MCP + RAG?
Traditional RAG systems often tightly couple the agent, retriever, vector store, and application logic. MCP creates a cleaner boundary:
* The agent does not need to know how retrieval is implemented.
* Retrieval can be exposed as a reusable tool.
* Other MCP-compatible clients can use the same backend.
* The server can add more tools without rewriting the agent runtime.
## Current Retrieval Tool
The MCP server exposes:
```python
retrieve_documents(query: str, top_k: int = 20, top_n: int = 5)
```
It returns:
* retrieved context
* document IDs
* chunk indexes
* vector similarity scores
* rerank scores when available
## Retrieval Evaluation
Run retrieval evaluation:
```bash
uv run python -m scripts.evaluate_retrieval
```
Run with custom retrieval settings:
```bash
uv run python -m scripts.evaluate_retrieval --top-k 8 --top-n 4
```
Write detailed results to JSON:
```bash
uv run python -m scripts.evaluate_retrieval --json-output scripts/eval/results/latest.json
```
## Web API and Browser Demo
Install API dependencies:
```bash
uv add fastapi uvicorn
```
Run the FastAPI server:
```bash
uv run uvicorn api.main:app --reload
```
Open the browser UI:
```text
[http://localhost:8000](http://localhost:8000)
```
Open the interactive API docs:
```text
[http://localhost:8000/docs](http://localhost:8000/docs)
```
### API Endpoints
Health check:
```bash
curl [http://localhost:8000/health](http://localhost:8000/health)
```
Non-streaming chat:
```bash
curl -X POST [http://localhost:8000/chat](http://localhost:8000/chat)
-H "Content-Type: application/json"
-d '{"message": "Explain how MCP tools work with RAG"}'
```
Streaming chat:
```bash
curl -N -X POST [http://localhost:8000/chat/stream](http://localhost:8000/chat/stream)
-H "Content-Type: application/json"
-d '{"message": "Explain how MCP tools work with RAG"}'
```
Direct retrieval:
```bash
curl -X POST [http://localhost:8000/retrieve](http://localhost:8000/retrieve)
-H "Content-Type: application/json"
-d '{"query": "What is Model Context Protocol?", "top_k": 8, "top_n": 4}'
```
List available MCP tools:
```bash
curl [http://localhost:8000/tools](http://localhost:8000/tools)
```
The browser UI uses the streaming endpoint so the answer appears token-by-token while the final response is generated.
## Configuration
Chat generation and embeddings are configured independently, allowing combinations such as Bedrock + Ollama or Ollama + Bedrock.
Create a local `.env` file from the example:
```bash
cp .env.example .env
```
### Use Ollama for chat and embeddings
```bash
CLIENT_LLM_PROVIDER=ollama OLLAMA_CHAT_MODEL=llama3.1 RAG_EMBEDDING_PROVIDER=ollama OLLAMA_EMBEDDING_MODEL=nomic-embed-text
```
### Use Bedrock for chat and Ollama for embeddings
```bash
CLIENT_LLM_PROVIDER=bedrock BEDROCK_CHAT_MODEL=anthropic.claude-3-5-sonnet-20240620-v1:0 AWS_REGION=us-east-1
RAG_EMBEDDING_PROVIDER=ollama OLLAMA_EMBEDDING_MODEL=nomic-embed-text
```
### Use Bedrock for chat and embeddings
```bash
CLIENT_LLM_PROVIDER=bedrock BEDROCK_CHAT_MODEL=anthropic.claude-3-5-sonnet-20240620-v1:0
RAG_EMBEDDING_PROVIDER=bedrock BEDROCK_EMBEDDING_MODEL=amazon.titan-embed-text-v2:0
AWS_REGION=us-east-1
```
If you switch embedding providers, rebuild or re-index the vector database because vectors from different embedding models are not compatible.
## Architecture Decisions
- **MCP** defines the boundary between the agent runtime and backend capabilities.
- **LangGraph** orchestrates workflows but does not implement retrieval logic.
- **Retrieval is exposed as an MCP tool**, allowing any MCP-compatible client to reuse the backend.
- **LLM providers** are swappable through a common interface (Ollama or AWS Bedrock).
- **Embedding providers** are independent of the chat model.
- **Vector stores** are abstracted behind a common interface (ChromaDB by default).
- **Retrieval results** are exposed as DTOs to avoid leaking internal domain models through the MCP API.
## Portfolio Highlights
This project demonstrates:
* MCP tool design
* RAG architecture
* Provider abstraction (Ollama / Bedrock)
* vector search
* source-grounded generation
* agent tool-calling loops
* modular backend design
* Python async programming
* clean CLI UX
## Future Improvements
Planned improvements:
* Docker Compose setup
* streaming responses
* structured logging
* query rewriting
* hybrid keyword/vector retrieval
* automated tests
* CI workflow
TDQS
C2.9/5.0
Scored across 1 tool
Disambiguation5/5
With only one tool, there is no possibility of confusion between tools.
Naming Consistency4/5
The single tool name 'retrieve_documents' follows a clear verb_noun pattern and is descriptive, but there is no set of tools to evaluate consistency.
Tool Count2/5
A RAG agent typically requires multiple tools (e.g., retrieval and generation). A single tool is too few for the stated purpose.
Completeness2/5
The tool set covers only retrieval, missing essential capabilities like generation or query processing for a RAG agent.
Maintenance
ActivitySlowing
ResponsivenessNo issues