RAG Factory
by Zyxac24
README.md
# RAG Factory — AutoML for RAG
**RAG Factory** turns RAG configuration from guesswork into measurement. From a short
conversation plus a sample of your own documents, it recommends candidate RAG
configurations, **builds** an isolated sub-environment for each, **evaluates** it on your
own data, and **optimizes** across the candidate grid — then registers the winning
environments as a persistent fleet you can query directly or over MCP.
> **Status:** v1.0 snapshot. This repository is a curated portfolio snapshot of a project
> that continues to evolve in a private development repository.
> 📸 _Screenshots / demo of the web UI will be added here._
---
## The problem
A good RAG setup **cannot be reliably predicted up front**. MTEB scores do not predict
performance in your domain, and the chunking strategy can matter as much as the embedding
model. The only trustworthy answer is to **measure candidates on your own corpus** — but
doing that by hand (generating a QA set, building indexes, running ranking metrics,
comparing) is slow and expensive.
## The solution
RAG Factory closes the loop:
```
[Recommender] -> [Factory] -> [Evaluator]
input build measure
^ |
+-------- optimization loop ------+
```
- **Recommender (input).** A conversational interview plus a data-profile turns your intent
and a document sample into a *search space* of candidate configurations — not a single
guessed "truth".
- **Factory (build).** For a given configuration it builds a sub-environment: parsing →
chunking → embedding → index.
- **Evaluator (output).** It measures each candidate on your own data (MRR, NDCG,
faithfulness / completeness), and the result feeds back to select the winner.
- **Registry (fleet).** Winning environments are persisted and can be served — locally, via
a Gradio web UI, or through an **MCP server**.
## What is built here vs. adopted
Honesty about the boundary matters, so it is explicit:
| Layer | Decision | Why |
|---|---|---|
| Build + eval + optimize core, ranking metrics (MRR/NDCG) | **Adopted — [AutoRAG](https://github.com/Marker-Inc-Korea/AutoRAG)** | The optimizer and validated metrics are a solved problem; reimplementing them would be worse and slower |
| QA ground-truth generation | **Adopted** (engine-provided) | Hand-curating a test set dominates development time |
| **Conversational recommender** (input layer) | **Built here** | No existing tool offers this; the unique, agentic part |
| **Fleet registry / environment manager** | **Built here** | Tools optimize "one best pipeline for one dataset"; they do not manage N persistent environments |
| **MCP server + Gradio web UI + optimization orchestration** | **Built here** | The serving and human-in-the-loop layer wrapping the engine |
Principle: **reinvent zero metrics and zero optimizer; build exactly what is new.**
## Tech stack
Python · [AutoRAG](https://github.com/Marker-Inc-Korea/AutoRAG) · LlamaIndex (OpenAI-compatible LLMs) ·
[MCP](https://modelcontextprotocol.io/) (FastMCP) · Gradio · pandas / numpy · langdetect · PyYAML ·
LLM access via [OpenRouter](https://openrouter.ai/).
## Package layout
| Module | Responsibility |
|---|---|
| `ragfactory/interview.py`, `profiler.py` | Conversational recommender: intent interview + data profiling |
| `ragfactory/recommendation.py`, `grid.py` | Candidate search space & finalist recommendation |
| `ragfactory/corpus.py`, `solution.py`, `yaml_builder.py` | Corpus prep, QA generation, environment build (AutoRAG YAML) |
| `ragfactory/runner.py`, `optimize.py`, `_optimize_worker.py` | Build/eval execution and the optimization loop |
| `ragfactory/generation.py`, `reports.py`, `emission.py` | Answer generation, reporting, artifact emission |
| `ragfactory/registry.py`, `serving.py` | Fleet registry & query serving |
| `ragfactory/mcp_server.py` | MCP server (FastMCP, stdio) exposing the fleet |
| `ragfactory/app.py`, `web_ui.py` | Gradio web application |
## Getting started
```bash
# 1. Environment
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# 2. Configure access (copy the template and fill in your key)
cp .env.example .env
# then edit .env: OPENROUTER_API_KEY, OPENROUTER_BASE_URL, OPENROUTER_MODEL
```
Run the web application:
```bash
python -m ragfactory.app
```
Run the MCP server (stdio transport):
```bash
python -m ragfactory.mcp_server
# optional: RAGFACTORY_REGISTRY_DIR=/path/to/registry python -m ragfactory.mcp_server
```
> **Note:** some build/evaluation paths run AutoRAG and may require an embedding model and a
> GPU. A corpus is **not** bundled — you point RAG Factory at your own documents.
## Testing
```bash
pytest # unit / fast tests (default; costly tests deselected)
pytest -m integration # slower tests that run AutoRAG / GPU
pytest -m costly # tests that spend money on live LLM calls (OpenRouter)
```
The suite is substantial (24 test modules) and mirrors the production package closely.
Markers are defined in `pytest.ini`.
## Documentation
- `docs/ARCHITECTURE.md` — design decisions, the build-vs-adopt rationale, and evaluation
methodology.
## License
Released under the [MIT License](LICENSE) © 2026 Dariusz Poślad.
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues