Skip to main content
Glama

aichallenge-rag

Standalone RAG index / search service for AIChallenge and any client that can call HTTP or Streamable HTTP MCP.

Vector backend: Qdrant — embedded on disk by default (RAG_DATA_DIR/qdrant), or a remote server via QDRANT_URL.

Index a folder of docs (or upload files), retrieve chunks with cosine similarity + metadata filters, optionally filter + heuristic-rerank, then feed the context into your LLM. This process never calls a chat model for answers — only embeddings (API, local, or fake).

You run

AIChallenge / client sees

Qdrant index + search on your machine (or stand)

HTTP /v1/* or Guest MCP /mcp

Embedded store under RAG_DATA_DIR (or remote Qdrant)

Bearer token only

Origin in the monorepo: apps/rag. This repo is the public, self-contained copy.

Features

  • Qdrant vector store — COSINE dense vectors, payload filters (scope, owner_id, …)

  • Corpus index — walk RAG_CORPUS_DIR (md/txt/pdf with text layer, common source suffixes)

  • Two chunk strategies — fixed window vs structural (headings / files)

  • User uploads — POST /v1/documents/upload with optional owner_id scope

  • Retrieval modes — raw · filtered (min score) · full (rewrite + filter + heuristic rerank)

  • /v1/ask — returns a ready context / prompt_suffix block (no LLM inside)

  • MCP tools — rag_stats, rag_search, rag_index on /mcp

  • Embeddings — OpenAI-compatible API, optional sentence-transformers, or deterministic fake

Related MCP server: RAG MCP Server

5-minute local run

git clone https://github.com/ArtemKyslicyn/aichallenge-rag.git
cd aichallenge-rag
cp .env.example .env
# optional: set RAG_SHARED_TOKEN to a long random string
uv sync
uv run python -m aichallenge_rag

With the bundled corpus/sample.md and EMBEDDING_PROVIDER=fake:

curl -sS -X POST http://127.0.0.1:18766/v1/search \
  -H 'Content-Type: application/json' \
  -d '{"query":"chunk strategies","mode":"full"}'

Tests:

uv run pytest

Qdrant modes

Mode

Config

When

Embedded (default)

QDRANT_URL= empty

Local demos, Compose stand, CI

Server

QDRANT_URL=http://host:6333

Shared / larger corpora

# optional remote
docker run -d --name qdrant -p 6333:6333 qdrant/qdrant
export QDRANT_URL=http://127.0.0.1:6333

Collection name: QDRANT_COLLECTION (default aichallenge_rag).

Connect to AIChallenge (Guest MCP)

  1. Run the service on loopback (RAG_HOST=127.0.0.1).

  2. Set the same long token in .env as RAG_SHARED_TOKEN.

  3. Tunnel HTTPS to the process, e.g. cloudflared tunnel --url http://127.0.0.1:18766.

  4. On the site: Profile → Подключения (Guest MCP)

    • URL: https://<tunnel-host>/mcp

    • Token: same as RAG_SHARED_TOKEN

Full walkthrough: docs/connect-aichallenge.md.

HTTP API (summary)

Method

Path

Purpose

GET

/health

Liveness (no auth)

GET

/v1/stats

Chunk / vector / Qdrant backend

POST

/v1/index

Rebuild corpus (fixed | structural)

POST

/v1/heal

Index if empty / rebuild if inconsistent

POST

/v1/documents

Add plain-text document (JSON)

POST

/v1/documents/upload

Upload file (multipart)

GET

/v1/documents

List docs (owner_id, include_stand)

POST

/v1/search

Retrieve hits + scores

POST

/v1/ask

Search + context / prompt_suffix

PATCH

/v1/settings

Flip embed provider / chunk strategy

Details and examples: docs/api.md.

When RAG_SHARED_TOKEN is set, send Authorization: Bearer <token> on all routes except /health, /, /docs, /openapi.json.

Docker

docker compose up --build

Binds loopback only 127.0.0.1:18766. Never publish on public :443 / :8443 without your own edge plan.

Optional local embeddings: uv sync --extra local, then PATCH /v1/settings {"local_embeddings": true}.

Docs

Doc

Topic

docs/architecture.md

Qdrant pipeline, modes

docs/api.md

HTTP + MCP contracts

docs/connect-aichallenge.md

Tunnel + Guest MCP

docs/security.md

Tokens, bind, risks

Env (names only)

See .env.example. Important names:

  • RAG_SHARED_TOKEN — Bearer for HTTP + MCP (empty = open local/dev)

  • RAG_CORPUS_DIR / RAG_DATA_DIR — corpus + Qdrant path / meta

  • QDRANT_URL / QDRANT_COLLECTION — remote vs embedded

  • EMBEDDING_PROVIDER — api | local | fake

  • LLM_API_KEY / ROUTERAI_KEY — for API embeddings

  • RAG_MODE / RAG_MIN_SCORE / RAG_TOP_K_* — retrieval defaults

Requirements

  • Python 3.12+

  • uv recommended

  • qdrant-client (pulled by uv sync)

  • Optional: embedding API key, or uv sync --extra local

  • Optional: cloudflared / ngrok for Guest MCP

Security

Whoever has URL + Bearer can reindex, search, and upload into your index. Keep bind on loopback; only expose via a tunnel you control. See docs/security.md.

License

MIT — see LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Local-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.
    3
    13 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A Model Context Protocol (MCP) server for Retrieval-Augmented Generation (RAG) operations. It provides tools for building and querying vector-based knowledge bases from document collections, enabling semantic search and document retrieval capabilities.
    3
    MIT