evidence-rag
Provides local embedding and chat model inference for the Evidence RAG server, enabling dense retrieval, hybrid search via reciprocal-rank fusion, and cited answer generation through Ollama.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@evidence-ragwhat evidence supports the merge gesture hint?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Evidence RAG
Local search and MCP with evidence packets you can recheck after sources change. Index selected files, preserve the passages used in a document, and verify their current revision and approval.
Retrieved passages include source hashes, line ranges and approval status. Answers cite the passages with exact supporting quotes.
Evidence packets preserve selected source text outside the index. Re-indexing a document does not rewrite an exported packet; verification reports its integrity, current approval and source drift separately.
Explicit .md / .txt files
→ bounded chunks + SHA-256 + source status
→ SQLite FTS5 + optional Ollama embeddings
→ reciprocal-rank fusion
→ bounded source context → Ollama → citation / quote validationRun
Requires Python 3.10+ with SQLite FTS5. Lexical search works without a model server.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
evidence-rag index --root examples/corpus --approved \
onboarding-study.md economy-review.md career-approved.md \
launch-plan.md release-review.md
evidence-rag search "What evidence supports the merge gesture hint?"
evidence-rag evaluate --cases examples/eval_cases.json
evidence-rag packet create "What evidence supports the merge gesture hint?" \
--output .data/tutorial-evidence.json
evidence-rag packet verify .data/tutorial-evidence.jsonPacket creation and verification work without Ollama or an API key. On Windows, create the virtual environment with python -m venv .venv and activate it with .venv\Scripts\Activate.ps1 before running the same commands. See the packet workflow for a standalone example and the meaning of each verification result.
All examples are synthetic. Alex Rivera is a fictional candidate; the career claims do not describe the repository's author. untrusted-note.md tests hostile source instructions and is excluded from the example index command.
The operator approves documents at ingestion with --approved; default searches exclude unapproved documents. Use --status historical or --status superseded to retain older records outside the default current-evidence view. Re-index a file to update its status or contents. evidence-rag remove study.md deletes the source and all its spans, even if the original file is gone.
Related MCP server: memory-bank-mcp
Add local inference
Install Ollama and pull the two models:
ollama pull qwen3.5:4b
ollama pull qwen3-embedding:0.6b
evidence-rag index --root examples/corpus --approved --dense \
onboarding-study.md economy-review.md career-approved.md \
launch-plan.md release-review.md
evidence-rag search --mode hybrid "What is known about the tutorial?"
evidence-rag ask --mode hybrid "What are the conflicting Orbit launch dates?"Hybrid retrieval combines FTS5/BM25 and cosine ranking through reciprocal-rank fusion (k=60). Dense vectors come from Ollama's /api/embed. Provider failures return errors without substituting synthetic vectors. All eligible chunks require the same embedding model and dimensions; create a fresh database when changing models.
The default chat model is qwen3.5:4b. The tested 24 GB M5 Pro works better with --chat-model qwen3.5:9b; the 27B model caused substantial memory pressure and was removed. The Windows RTX 3090 profile uses qwen3.8:27b-q4_K_M, a 32,768-token chat context and a 4,096-token embedding context. Both models were observed on GPU together. See PC measurements and embedding context for the measured scope and limits.
Global flags precede the command:
evidence-rag --db .data/product.sqlite3 \
--ollama-url http://127.0.0.1:11434 \
--chat-model qwen3.8:27b-q4_K_M \
--embedding-model qwen3-embedding:0.6b serveEnvironment equivalents: EVIDENCE_RAG_DB, OLLAMA_BASE_URL, OLLAMA_CHAT_MODEL, OLLAMA_EMBEDDING_MODEL, OLLAMA_CONTEXT_TOKENS, and OLLAMA_MAX_OUTPUT_TOKENS. --context-tokens defaults to 8192; --max-output-tokens defaults to 1024. --provider-timeout / OLLAMA_TIMEOUT sets the network inactivity timeout (CLI default 300 seconds). No environment files are read automatically. For a Mac talking to a PC, an SSH tunnel keeps the PC's Ollama listener on loopback:
ssh -N -L 11435:127.0.0.1:11434 user@your-workstation
evidence-rag --ollama-url http://127.0.0.1:11435 serveAPI and application boundary
evidence-rag serve listens at 127.0.0.1:8091. The HTTP API provides health, retrieval, individual approved spans and answers. Ingestion is CLI-only; the API cannot crawl directories or write to a career or product application.
curl http://127.0.0.1:8091/v1/search \
-H 'Content-Type: application/json' \
-d '{"query":"Orbit launch date","corpus_id":"default","limit":5,"mode":"lexical","approved_only":true,"statuses":["current"]}'Spans contain chunk_id, document_id, relative source, source_sha256, start_line, end_line, start_char, end_char, text, approved, status, and complete:false. Character offsets are zero-based, end-exclusive Unicode code-point positions in decoded UTF-8 text. Lines are one-based and inclusive. Spans cover excerpts, not complete documents.
POST /v1/answer accepts the same request plus optional max_context_chars (default 10,000). It returns status, answer, full citation spans, exact supporting_quotes, and an abstention reason when applicable. GET /v1/spans/{chunk_id} resolves an approved span; GET /health reports index counts.
The model receives short source aliases and returns structured statements with references. Code validates each alias and exact quote, then renders citations with persisted chunk IDs. Invalid output gets at most one correction attempt using the same sources. A conservative UTF-8 byte bound reserves space for completion and the chat template before selecting whole spans. If none fit, the service returns context_budget_exceeded; it never relies on Ollama silently truncating evidence.
In a career workbench, retrieve approved evidence and inspect the spans before freezing the reviewed input packet. In a product workbench, retrieve study excerpts and decision notes before drafting a proposal. Keep retrieval outside already approved packets and application writes. The companion Local Workbench provides a separate local task API.
Set EVIDENCE_RAG_API_TOKEN to require a bearer token. A token is required for binding beyond loopback; use an authenticated private network or TLS proxy for remote access. There is no multi-user authorization, and approved_only:false permits reading unapproved indexed data. A shared instance should contain only material the operator has approved for indexing.
MCP
Install the adapter with python -m pip install -e '.[mcp]', then run evidence-rag mcp. It exposes read-only search_evidence, get_evidence_span, answer_from_evidence, create_evidence_packet, and verify_evidence_packet tools over stdio. Packet creation returns JSON; the client decides whether to save it. Searches default to approved, current sources. The tools cannot ingest files, approve sources or edit an application. The official MCP Python SDK is pinned to the API version used by the adapter.
Evaluation and limits
See local measurements for the 20-document retrieval comparison and real-model answer checks. requirements.lock records the Python dependency versions used for these runs.
python -m pip install -e '.[dev]'
ruff check .
pytest -q
evidence-rag evaluate --cases examples/eval_cases.json
python scripts/packet_demo.pyThe packet lifecycle report exercises source edits, re-indexing, approval revocation, deletion, index removal and packet tampering with synthetic material and zero model calls. It checks deterministic state transitions, not the truth of the source's claims.
The retrieval report measures document recall at k and reciprocal rank on a small synthetic corpus. Unanswerable questions appear without scores because dense retrieval returns nearest candidates even when none support an answer. That decision needs evaluation at answer time; the retrieval report sets generation_evaluated:false. Offline tests cover source boundaries, stale-chunk replacement, approval/status filters, embedding mismatches, citation rejection, exact quotes, abstention and conflicting evidence. Fake providers exist only in tests.
Dense retrieval scans stored vectors in Python and is intended for personal corpora, not large collections. Import explicitly named UTF-8 .md or .txt files of up to 2 MB each. Hidden or credential-like paths and symlinks are rejected. Re-indexing is atomic per document; files deleted from disk remain indexed until removed through the CLI. The index stores model tags and preprocessing versions but not model artifact digests. Rebuild it if a model changes behind an existing tag.
Citation validation checks that cited spans were supplied and quoted accurately. Whether an answer follows from those sources still needs review: the model can misread evidence or miss a conflict. Answers are drafts, and a few synthetic fixtures cannot establish production reliability. Source instructions remain untrusted data. The answer model has no tools, filesystem access or application write capability.
MIT licensed. No private application code, real applicant history, credentials, or proprietary research is included.
This server cannot be deployed
Maintenance
Related MCP Connectors
Citable retrieval across papers, books, patents, Wikipedia, and live social sources.
Machine-readable entity discovery with provenance, trust and verified source evidence.
Authenticated public evidence search, verification, research jobs, exports, and webhooks.
341Search source collections your team vetted. Every passage cites its URL and fetch date.
Related MCP Servers
- AlicenseBqualityBmaintenanceProvides read-only hybrid RAG search and discovery over a local-first AI knowledge corpus, enabling semantic and keyword search, browse, digest, and status tools.4PolyForm Noncommercial 1.0.0
- AlicenseBqualityBmaintenanceProvides a read-only MCP interface to query and retrieve verifiable evidence from a local memory bank, supporting search, dossier, chronology, source, and evidence tools.6BSD Zero Clause
- FlicenseNot gradedqualityBmaintenanceProvides read-only access to a personal RAG knowledge base, enabling hybrid search, evidence-grounded retrieval with citations, and knowledge gap tracking for LLM agents.-
- FlicenseAqualityCmaintenanceProvides read-only MCP tools to search and retrieve evidence-grounded knowledge compiled from video content, including hybrid semantic and lexical search with citations.5-