aichallenge-rag
Integrates with OpenAI-compatible embedding APIs to generate vector embeddings for document indexing and search queries.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@aichallenge-ragsearch my indexed docs for chunk strategies using full reranking"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
aichallenge-rag
Standalone RAG index / search service for AIChallenge and any client that can call HTTP or Streamable HTTP MCP.
Vector backend: Qdrant — embedded on disk by default (RAG_DATA_DIR/qdrant), or a remote server via QDRANT_URL.
Index a folder of docs (or upload files), retrieve chunks with cosine similarity + metadata filters, optionally filter + heuristic-rerank, then feed the context into your LLM. This process never calls a chat model for answers — only embeddings (API, local, or fake).
You run | AIChallenge / client sees |
Qdrant index + search on your machine (or stand) | HTTP |
Embedded store under | Bearer token only |
Origin in the monorepo: apps/rag. This repo is the public, self-contained copy.
Features
Qdrant vector store — COSINE dense vectors, payload filters (
scope,owner_id, …)Corpus index — walk
RAG_CORPUS_DIR(md/txt/pdf with text layer, common source suffixes)Two chunk strategies —
fixedwindow vsstructural(headings / files)User uploads —
POST /v1/documents/uploadwith optionalowner_idscopeRetrieval modes —
raw·filtered(min score) ·full(rewrite + filter + heuristic rerank)/v1/ask— returns a readycontext/prompt_suffixblock (no LLM inside)MCP tools —
rag_stats,rag_search,rag_indexon/mcpEmbeddings — OpenAI-compatible API, optional
sentence-transformers, or deterministicfake
Related MCP server: RAG MCP Server
5-minute local run
git clone https://github.com/ArtemKyslicyn/aichallenge-rag.git
cd aichallenge-rag
cp .env.example .env
# optional: set RAG_SHARED_TOKEN to a long random string
uv sync
uv run python -m aichallenge_ragHealth: http://127.0.0.1:18766/health
OpenAPI: http://127.0.0.1:18766/docs
Stats (
backend=qdrant):GET /v1/stats
With the bundled corpus/sample.md and EMBEDDING_PROVIDER=fake:
curl -sS -X POST http://127.0.0.1:18766/v1/search \
-H 'Content-Type: application/json' \
-d '{"query":"chunk strategies","mode":"full"}'Tests:
uv run pytestQdrant modes
Mode | Config | When |
Embedded (default) |
| Local demos, Compose stand, CI |
Server |
| Shared / larger corpora |
# optional remote
docker run -d --name qdrant -p 6333:6333 qdrant/qdrant
export QDRANT_URL=http://127.0.0.1:6333Collection name: QDRANT_COLLECTION (default aichallenge_rag).
Connect to AIChallenge (Guest MCP)
Run the service on loopback (
RAG_HOST=127.0.0.1).Set the same long token in
.envasRAG_SHARED_TOKEN.Tunnel HTTPS to the process, e.g.
cloudflared tunnel --url http://127.0.0.1:18766.On the site: Profile → Подключения (Guest MCP)
URL:
https://<tunnel-host>/mcpToken: same as
RAG_SHARED_TOKEN
Full walkthrough: docs/connect-aichallenge.md.
HTTP API (summary)
Method | Path | Purpose |
|
| Liveness (no auth) |
|
| Chunk / vector / Qdrant backend |
|
| Rebuild corpus ( |
|
| Index if empty / rebuild if inconsistent |
|
| Add plain-text document (JSON) |
|
| Upload file (multipart) |
|
| List docs ( |
|
| Retrieve hits + scores |
|
| Search + |
|
| Flip embed provider / chunk strategy |
Details and examples: docs/api.md.
When RAG_SHARED_TOKEN is set, send Authorization: Bearer <token> on all routes except /health, /, /docs, /openapi.json.
Docker
docker compose up --buildBinds loopback only 127.0.0.1:18766. Never publish on public :443 / :8443 without your own edge plan.
Optional local embeddings: uv sync --extra local, then PATCH /v1/settings {"local_embeddings": true}.
Docs
Doc | Topic |
Qdrant pipeline, modes | |
HTTP + MCP contracts | |
Tunnel + Guest MCP | |
Tokens, bind, risks |
Env (names only)
See .env.example. Important names:
RAG_SHARED_TOKEN— Bearer for HTTP + MCP (empty = open local/dev)RAG_CORPUS_DIR/RAG_DATA_DIR— corpus + Qdrant path / metaQDRANT_URL/QDRANT_COLLECTION— remote vs embeddedEMBEDDING_PROVIDER—api|local|fakeLLM_API_KEY/ROUTERAI_KEY— for API embeddingsRAG_MODE/RAG_MIN_SCORE/RAG_TOP_K_*— retrieval defaults
Requirements
Python 3.12+
uv recommended
qdrant-client(pulled byuv sync)Optional: embedding API key, or
uv sync --extra localOptional: cloudflared / ngrok for Guest MCP
Security
Whoever has URL + Bearer can reindex, search, and upload into your index. Keep bind on loopback; only expose via a tunnel you control. See docs/security.md.
License
MIT — see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
Serve a folder of Markdown notes as an MCP server: hybrid search, reading, and sourced answers.
Ingest, manage, and retrieve documents for RAG-powered AI applications
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
Related MCP Servers
- AlicenseAqualityDmaintenanceLocal-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.313 npmMIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol (MCP) server for Retrieval-Augmented Generation (RAG) operations. It provides tools for building and querying vector-based knowledge bases from document collections, enabling semantic search and document retrieval capabilities.3MIT
- FlicenseNot gradedqualityDmaintenanceIndexes PDF documents into Qdrant and exposes semantic search as MCP tools, enabling RAG-based interactions with your documents.-
- AlicenseNot gradedqualityBmaintenanceLocal RAG over a directory, served as an MCP tool plus CLI, with incremental indexing, zero-config defaults, and offline CPU-only operation.MIT