io.github.Brightdotdev/darwin-rag
OfficialDarwin RAG
A local-first RAG engine that ingests documents, indexes them with BM25 + dense embeddings, and exposes search via an MCP server for AI agent integration.
Ingestion — PDF, Markdown, HTML, images (OCR), CSV, Excel, ODS, URLs
Indexing — BM25 keyword + dense embedding hybrid index with configurable chunking strategies
Search — Hybrid, semantic, or keyword retrieval with reranking, diversity rerank, and structural penalties
Generation — LLM-backed answer synthesis via LiteLLM (OpenAI, Anthropic, Gemini, etc.)
Observability — Structured logging with per-session history, queryable via MCP
Isolation — Multiple independent stores for tenant/project separation
Deployment — stdio (AI agent subprocess), SSE, or Streamable HTTP; Docker-ready
Local-first — Everything runs locally, fully offline-capable after setup
Built by BrightDotDev.
Quick Start
1. Install
pip install darwin-rag2. Set up models
# Interactive — detects hardware, pick your models
darwin-admin setup interactive
# Or one-shot (embedding-only, no prompts)
darwin-admin setup --preset required3. Start the MCP server & connect
# stdio mode — for AI agent subprocess (Claude Desktop, Cursor, etc.)
darwin mcp
# Or HTTP mode — for remote clients
darwin mcp --http --port 8765Configure your MCP client:
{
"mcpServers": {
"darwin": {
"command": "darwin",
"args": ["mcp"],
"env": {
"OPENAI_API_KEY": "sk-..." // At least one LLM provider key
}
}
}
}Or generate config automatically:
darwin config claude # Claude Desktop config
darwin config cursor # Cursor config
darwin config all --copy # All clients + copy to clipboardMCP Tools
Tool | Description |
| Query the knowledge base with hybrid/semantic/keyword search |
| Search structured data (CSVs, JSON arrays) by field values |
| List saved search results |
| Load a saved search result by filename |
| Inspect schemas for structured files (keys, types, record counts) |
| Ingest + index documents from a path or URL |
| Delete pipeline artifacts for specific files |
| Create a new isolated data store |
| List all tracked files with pipeline status |
| Detailed status for a single file across all stages |
| Query session logs (oldest first, INFO excluded) |
Full documentation: docs/mcp.md
Remote / HTTP Mode
Start the server on a network-accessible endpoint:
# SSE transport (legacy)
darwin mcp --sse --host 0.0.0.0 --port 8765
# Streamable HTTP transport (recommended for production)
darwin mcp --http --host 0.0.0.0 --port 8765Configure your MCP client with the URL:
{
"mcpServers": {
"darwin": {
"url": "http://your-host:8765/mcp" // or /sse for SSE mode
}
}
}Environment Variables
Variable | Required | Description |
| No* | OpenAI provider key |
| No* | Anthropic provider key |
| No* | Google Gemini provider key |
| No* | Mistral AI provider key |
| No* | Groq provider key |
| No* | Cohere provider key |
| No* | Together AI provider key |
| No* | OpenRouter provider key |
| No* | DeepSeek provider key |
| No | Override the base data directory |
| No | Set to any value to disable ANSI color output |
* At least one LLM provider key is required for answer generation. Search/indexing works without any.
Setup Details
Command | What it does |
| Guided setup — detect hardware, choose models |
| Download embedding model only (fastest) |
| Embedding + reranker + OCR models |
| Reconfigure logging only |
| Validate current setup |
See docs/setup.md for the full walkthrough including Docker, from-source install, and API key configuration.
CLI Reference
darwin — User CLI
Command | Description |
| Start MCP server (stdio, |
| Generate MCP client config |
darwin-admin — Power-user CLI
Command | Description |
| Setup models, logging, and configuration |
| System status overview |
| Model registry: list, install, switch, keys |
| Data store: status, files, audit, health, repair |
| Ingestion pipeline: run, ingest, index, purge |
| Interactive search |
| Structured log viewer and management |
| System information |
| Remove Darwin data and configuration |
See docs/admin.md for the full command reference.
Python API
For embedding darwin-rag as a library in your own app:
from core import Darwin
d = Darwin()
d.ingest("./papers", recursive=True)
results = d.search("what is this paper about")
records = d.search_records(filters={"status": "active"})Full reference: docs/api.md
Documentation
Doc | What |
Full setup walkthrough | |
MCP server, tools, resources, transports | |
Admin CLI reference | |
For developers and contributors | |
Ingestion & indexing | |
Search engine | |
DarwinStore | |
Model registry & inference | |
Structured logging | |
High-level business logic | |
Python API (Darwin class) |
Contributing
Found a bug? Want to add something? You're welcome here.
Issues — open one at github.com/BrightDotDev/DARWIN/issues
Code — fork, branch, PR. Keep it minimal.
AI-generated code is fine — but you own what you ship. Test it before submitting.
Read CONTRIBUTING.md for the full guidelines.
License
MIT with Attribution — see LICENSE.
Core architecture and implementation by BrightDotDev.