Skip to main content
Glama
cwarre33
by cwarre33

wiki-rag-mcp

Retrieval-augmented Q&A over a markdown knowledge base. It has three parts:

  • a LangGraph agent that grades its own retrieval and refuses when it has no grounded answer

  • an MCP server that exposes the wiki to any MCP client, such as Claude Code, Cursor, or the included client

  • an evaluation harness with retrieval metrics and Ragas LLM-as-a-judge scoring

It indexes the public pages of my engineering wiki (an Obsidian-style vault of production-system notes, ADRs, and Kaggle writeups). It works on any folder of markdown with YAML frontmatter.

            ┌──────────────┐   search_wiki / get_page / ask   ┌──────────────────┐
 MCP client │ Claude Code, │ ───────────────────────────────▶ │ wikirag.server   │
 (stdio)    │ Cursor, CLI  │ ◀─────────────────────────────── │ (MCP, stdio)     │
            └──────────────┘                                  └────────┬─────────┘
                                                                       │ ask
                        retrieve ─▶ grade ─(relevant)─▶ generate ─▶ cited answer
                           ▲          │
                           └─ rewrite ◀┘ (none relevant, 1 retry) ─▶ refuse
                                                                       │
                             FAISS inner-product index ◀───────────────┘
                             all-MiniLM-L6-v2 embeddings, 384 chunks / 50 pages

What's in it

Module

What it does

wikirag/corpus.py

Loads pages and applies the publication filter: visibility: public only, minus redacted sections and security tags. Parses Obsidian frontmatter that isn't valid YAML (related: [[a]], [[b]]). Renders wikilinks as words. Chunks by heading with overlapping windows.

wikirag/index.py

sentence-transformers embeddings in a FAISS IndexFlatIP (cosine on normalized vectors). Saves and loads to disk, and refuses to load an index built with a different embedding model.

wikirag/graph.py

LangGraph state machine: retrieve → LLM relevance grading → generate with numbered citations. On no relevant chunks it does one query rewrite, then refuses. Citations to sources the model wasn't given are stripped.

wikirag/server.py

MCP server (official Python SDK, MCPServer). Tools: search_wiki, get_page, ask. Resource: wiki://pages. The LLM is built lazily, so search-only clients never need one running.

wikirag/client.py

MCP client. Launches the server over stdio, lists its tools, and runs a search.

eval/run_eval.py

Scores the pipeline against eval/golden.jsonl. Hand-written questions with reference answers, pages and passages, plus one question the wiki can't answer.

Related MCP server: data-olympus MCP server

Quickstart

uv venv && uv pip install -e ".[eval,dev]"
wikirag-index ../cameron-wiki/wiki --out .index          # build the index
python -m wikirag.client "why run CLIP as a subprocess"   # MCP client → server → FAISS

ollama pull llama3.2:3b                                  # local LLM for the agent and judge
wikirag-ask "Why does SofaScope use metadata scoring instead of embeddings?"

Use it from Claude Code:

claude mcp add wiki -- /path/to/.venv/bin/python -m wikirag.server --index /path/to/.index

Evaluation

python eval/run_eval.py --index .index          # retrieval metrics (no LLM)
python eval/run_eval.py --index .index --llm    # + Ragas faithfulness, answer relevancy, context recall

Retrieval baseline, 15 answerable questions, k = 5 (eval/results/2026-09-30.json):

Metric

Score

hit@5 (a reference page in the top 5)

0.93

MRR

0.86

Snippet recall (reference passages present verbatim in retrieved text)

0.57

The one miss asks about the trade-offs of the SofaScope metadata-scoring ADR. It retrieves the system overview and the hybrid-routing page instead, which discuss the same decision. Snippet recall is lower than hit rate because heading-based windows split some reference passages across chunks. That is the next thing to tune.

The --llm run evaluates generation with Ragas and uses the local Ollama model as both agent and judge:

  • Faithfulness

  • ResponseRelevancy

  • LLMContextRecall

  • refusal accuracy on the unanswerable question

Generation, 2026-09-30, 15 answerable questions plus 1 unanswerable. qwen2.5:7b is both the agent and the Ragas judge, running locally on Ollama:

Metric

qwen2.5:7b

llama3.2:3b (first run)

Faithfulness

0.73

0.50

LLM context recall

0.81

0.80

Answer relevancy

0.52 ⚠️

~1.00 ⚠️

False refusals (answerable questions refused)

0

—

Judge exceptions

0

25

Why the 3B run isn't trustworthy. It threw 25 judge exceptions, and Ragas records a failed judgment as NaN. Pandas' mean() silently skips NaN, so each 3B average may cover only a few of the 15 questions. That is the most likely reason answer relevancy came out at a suspiciously perfect ~1.00.

The eval now:

  • records how many samples each metric actually scored (<metric>_scored)

  • saves per-question scores to the results file

  • writes one results file per judge model instead of overwriting

Answer relevancy is under investigation. Ragas ResponseRelevancy asks the judge to generate questions from the answer. It scores the cosine similarity between those questions and the original, using the MiniLM embedder, and scores 0 when the judge flags the answer as noncommittal. Short, cited answers and a small embedder both pull the score down. The per-question scores from the next run will show which is driving it.

Tests

pytest -q        # 25 tests; deterministic hash embedder + fake chat model, no downloads
ruff check .

The tests cover the publication filter (private and security-tagged pages never reach the index or the MCP surface), chunking, index persistence, and every graph path: answer, rewrite then recover, rewrite then refuse. They also cover citation stripping, and the MCP tools and resource over an in-process client.

Design notes

  • Grading before generating. Small local models are easily distracted by off-topic context, so only chunks the grader keeps reach the answer prompt. (Measuring this with --llm runs, with and without grading, is on the to-do list.) Grading also gives the refusal path a clean trigger.

  • The publication filter lives at the index boundary. The server only ever sees what was indexed, so there is no per-request access check to get wrong.

  • One rewrite, not a loop. It keeps latency bounded and makes refusal deterministic on questions the wiki can't answer.

License

MIT

Related MCP Connectors

Related MCP Servers