Tea Rags MCP
Enriches code chunks with authorship, timestamps, churn metrics, and task IDs extracted from commit history and git blame data.
Supports extracting GitHub task IDs from commit messages to provide context and linking between code and project issues.
Enables extraction of JIRA task IDs from commit messages to associate indexed code chunks with specific project tickets.
Integrates with Ollama for local, privacy-first embedding generation and semantic codebase search.
Supports OpenAI embedding models for semantic vectorization and high-performance code search.
Provides specialized Ruby AST-aware chunking to improve the accuracy and relevance of semantic search in Ruby codebases.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Tea Rags MCPsearch for where user authentication is implemented"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Your coding agent copies the first code it finds β not the right one.
TeaRAGs is a Codebase Intelligence layer your agent queries over MCP. It indexes the repository on your machine into five layers β three that read the code and two that judge its interfaces:
π What it does β semantic and hybrid search over AST-aware chunks
πΈοΈ How it is connected β callers, callees, fan-in, transitive impact
𧬠How it has lived β churn, bug-fix rate, ownership, age
ποΈ Whether it is laid out right β dependency direction, leaking facades, files that change together with no edge between them
π€ What the project calls things β the naming vocabulary, inferred from the call graph, and a verdict on every new name
The first three layers and why an agent needs all of them at once are laid out in Codebase Intelligence Π΄Π»Ρ Π°Π³Π΅Π½ΡΠ° (Habr, in Russian).
TeaRAGs also ships agent skills that know which layer a task needs. The agent stops guessing which code is safe to copy, what is critical, and what a change will break β it reads the dossier instead.
π Documentation Β· π 15-minute quickstart Β· π§ Core concepts
π See It
Three questions an agent asks before touching code, answered by TeaRAGs on its own repository. Every number below is a real response, trimmed.
1. "Find retry logic I can reuse"
semantic_search { query: "retry a failed request with exponential backoff", rerank: "hotspots" }
Similarity alone puts OllamaEmbeddings#retryWithBackoff first. The dossiers of
the top two candidates tell different stories:
π₯ | π₯ | |
Similarity rank | #1 | #2 ( |
Commits to the file | 28 | 1 |
Share that were bug fixes | 54% Β· π΄ concerning | 0% Β· π’ healthy |
Last changed | 2 days ago Β· recent | 86 days ago Β· old |
Callers | 2 | 1 |
The closest match keeps getting fixed. The agent copies the quiet helper's shape β or learns why the first one keeps breaking before it repeats the mistake.
{
"symbolId": "OllamaEmbeddings#retryWithBackoff",
"relativePath": "src/core/adapters/embeddings/ollama.ts",
"startLine": 290,
"endLine": 378,
"preset": "hotspots",
"git": {
"file": {
"commitCount": 28,
"ageDays": { "value": 2, "label": "recent" },
"bugFixRate": { "value": 54, "label": "concerning" },
"relativeChurn": { "value": 2.55, "label": "normal" }
},
"chunk": {
"commitCount": { "value": 11, "label": "extreme" },
"relativeChurn": { "value": 9.09, "label": "high" }
}
},
"codegraph": { "symbols": { "chunk": { "fanIn": 2, "fanOut": 6 } } }
}Labels are computed from this repository's own percentiles, so extreme means extreme for this codebase, not for some global average.
2. "What is risky to touch around vector writes?"
semantic_search { query: "write points to the vector database in batches", rerank: "dangerous" }
Similarity alone ranks PointsAccumulator#flushBatch,
QdrantPointStore#addPointsOptimized and QdrantPointStore#addPoints first.
The dangerous preset reorders by risk and says why:
# | Ranked by risk | Why it moved up |
1 |
| 16 outgoing calls, 77 lines, 5 commits Β· high |
2 |
| file with 45 commits, relative churn 8.09 Β· π΄ high, 4 authors |
3 |
| one author owns 100% of the live lines Β· π deep-silo, 158 days untouched |
3. "Who calls it before I change it?"
get_callers { symbolId: "QdrantManager#addPointsWithSparse" }
Ten exact call sites across eight files β method fan-in 10 Β· central, file transitive impact 47 Β· regional:
ChunkPipeline#createBatchHandler ingest/pipeline/chunk-pipeline.ts
createQdrantPipeline ingest/pipeline/pipeline-manager.ts
storeIndexingMarker (2 sites) ingest/pipeline/indexing-marker.ts
DocumentOps#add api/internal/ops/document-ops.ts
SchemaManager#storeSchemaMetadata adapters/qdrant/schema-manager.ts
EmbeddingModelGuard#readOrCreateMarker adapters/qdrant/embedding-model-guard.ts
IndexStoreAdapter#storeSchemaVersion maintenance/migration/adapters/index-store-adapter.ts
SparseStoreAdapter#rebuildSparseVectors maintenance/migration/adapters/sparse-store-adapter.ts
SparseStoreAdapter#storeSparseVersion maintenance/migration/adapters/sparse-store-adapter.tsNeed the whole chain from an entry point to this call? trace_path enumerates
every AβB path and, with a rerank preset, sorts them by how dangerous each step
is.
Related MCP server: codesteer-atlas
β What It Answers
Ask in plain language. The plugin picks the skill, tools and rerank presets
for every question automatically β it ships a decision table that maps intent
to the right call, so nobody has to know a preset name. Other MCP clients get
the same routing guide as an MCP resource (tea-rags://schema/search-guide).
One question per layer, most distinctive first. Every number is a real response on TeaRAGs' own repository.
1. ποΈ "Is this codebase laid out correctly?"
/tea-rags:architecture-diagnostics β get_architecture_report
Four detectors judge the borders between modules, not the risk of touching them. Each violation carries the file edges or commits that prove it, grouped into root causes:
Detector | What it found here |
Stable Dependencies |
|
Leaking abstraction |
|
Silent coupling | the Go and Java resolvers changed together in 5 sessions β every Java change came with a Go one (lift 86.5) β with no import or call between them: a shared shape waiting to move into the language kernel |
Main sequence |
|
A component is a module with a facade, or a plain directory when there is none. The facade-adoption and coupling-strength cut-offs are derived from the repository itself (Otsu), so the same report reads a small library and a monolith without tuning.
2. π€ "Do the names in my diff speak the project's language?"
review_changes { changes: { base: "main" }, sections: ["naming"] } β also step
D8 of /tea-rags:mr-review and the last check of
/tea-rags:data-driven-generation
The project's vocabulary is read from its call graph: how values of each type are named, which role word each directory gives its types, which word the project already uses for a concept. Only declarations on added lines are judged, and the changed files are left out of the evidence, so a diff never confirms itself:
Draft | Verdict | Why |
| MISFIT β | the project names this type's values after |
| CONFORMS, alt. | the project's word for this concept is |
| NEW_TERM, alt. | same meaning as the project's |
Verdicts are CONFORMS, MISFIT with a suggestion, NEW_TERM with the
project's closest terms, and COLLISION for a name already taken elsewhere. The
rules are checked against the project's own history: every type rename a commit
message records is a case the tool must flag. get_ontology_report audits the
whole vocabulary β synonyms, homonyms, outliers.
3. 𧬠"We have four payment-gateway retries. Which one should I copy?"
semantic_search with proven β /tea-rags:data-driven-generation runs it as
its template step
Similarity finds all four; history decides. proven ranks long-lived,
low-bug-rate, multi-author code first, and every result carries its dossier β
commits, bug-fix share, age, owners β labelled against this repository's own
percentiles. The closest match by text is often the one fixed every sprint; see
See It β 1 for a real pair.
4. πΈοΈ "How does a request get from the API to the card charge, and which step is riskiest?"
trace_path with dangerous β get_callers for a single hop
Every call path between the two symbols, resolved from the call graph rather than guessed from names, with each step ranked by how risky it is to touch. An ambiguous call is reported as ambiguous, never picked at random. See See It β 3 for the ten call sites of a real write path.
5. π "Where do we charge a bill with a saved card?" β when the code says invoice
hybrid_search β dense vectors for the meaning, BM25 for exact names
The code is chunked on AST boundaries, so a result is a whole method with its class, not a window of N lines. Search by meaning crosses the vocabulary gap between the question and the code; BM25 still pins an exact identifier when the question names one.
The right column shows what runs under the hood.
πΊοΈ Understand
Ask your agent | What runs |
"Where do we charge a bill with a saved card, and what will it touch?" |
|
"Onboard me into billing β where are the entry points?" |
|
"Which modules is this whole app built around?" |
|
"What was done under ticket #4521?" |
|
β»οΈ Reuse and generate
Ask your agent | What runs |
"Add partial payments to bill payment β in our style, no duplicates." |
|
"We have four payment-gateway retries. Which one should I copy?" |
|
"Is there already a helper that rounds money amounts?" |
|
π― Change safely
Ask your agent | What runs |
"What should I not touch in this task, and where is it safer to build a parallel implementation?" |
|
"Who calls bill payment, and how does a request get from the API to the card charge?" |
|
"Which code here should never change without a second reviewer?" |
|
"Which tests cover the behaviour I'm about to change?" |
|
π Find problems
Ask your agent | What runs |
"Where are the most dangerous modules in the payments domain?" |
|
"After a retry, a bill gets marked as paid twice. What is most likely to blame?" |
|
"Map the tech debt in invoicing." |
|
"Which files in this domain changed most this month?" |
|
"What here is dead or abandoned?" |
|
π₯ Review, ownership and audit
Ask your agent | What runs |
"What in this merge request should I look at first?" |
|
"Whose code is this, and where is the bus factor one?" |
|
"Which old security-critical code is overdue for an audit?" |
|
β¨ Features
π Git- and codegraph-aware ranking β 23 rerank presets blend churn, bug-fix rate, ownership and age with fan-in, PageRank and transitive impact (
proven,hotspots,techDebt,blastRadius,criticalPath, β¦), plus 12 filter presetsπΈοΈ Call graph β callers, callees, cycles and AβB paths (
get_callers,get_callees,find_cycles,trace_path) for TypeScript, JavaScript, Ruby, Swift and Python at a high tierποΈ Architecture report β Stable Dependencies at component level, imports that leak past an adopted facade, files that change together with no edge between them, and distance from the main sequence (
get_architecture_report)π Change review β one call over a working-tree change or a branch: names off the project vocabulary, co-change partners the change left untouched, per-file cohesion, and new edges that break architecture boundaries (
review_changes)π€ Naming review β the project's vocabulary inferred from the call graph: verdicts on value and type names, the project's own word for a synonym, a review of every name a diff declares (
review_changes), and a whole-code audit of synonyms, homonyms and outliers (get_ontology_report)π§ Agent skills β the plugin routes every question to the right tools and presets on its own; 15 ready-made workflows (
explore,bug-hunt,risk-assessment,data-driven-generation,mr-review, β¦) plusdinopowers, 10 wrappers that feed index signals intosuperpowersπ 100% local β embedded Qdrant and DuckDB, no Docker; embeddings through Ollama β on the laptop or on any machine in your network, with the laptop as an automatic fallback (see Embedding providers) β with OpenAI, Cohere and Voyage optional
π Always fresh β incremental reindex, auto-update on a target branch, per-worktree index clones, and a drift report that names the exact command to run
π’ Built for enterprise monorepos β AST chunking for 9 languages, parallel pipelines, validated on a 3.5M-line production monolith
π¦ Installation
π» System requirements
Requirement | |
OS | macOS (arm64, x64) Β· Linux (x64, arm64) Β· Windows (x64) |
Node.js | 22+ supported, 24+ recommended |
git | Required β churn, ownership and bug-fix signals come from the repository's history |
Embeddings | Ollama with the code-embedding model (322 MB), or an OpenAI, Cohere or Voyage key |
Disk | 66 MB for the Qdrant binary, plus the per-project indexes below |
Disk taken by real indexes (turbo quantization, dense + sparse vectors):
Codebase | Indexed | Vector index (Qdrant) | Call graph (DuckDB) |
Production monolith (Ruby + TypeScript) | 3.3M lines of Ruby + TypeScript, tests included + 162K lines of docs Β· ~34k files Β· 175k chunks | 2.0 GB | ~460 MB |
TeaRAGs itself (TypeScript) | 686K LoC + 290K lines of docs Β· ~3.5k files Β· 45k chunks | 1.2 GB | 40 MB |
The call graph grows with the code; the vector index much less β a codebase five times smaller still takes 1.2 GB.
Pull the code-embedding model:
ollama pull unclemusclez/jina-embeddings-v2-base-code:latestClaude Code β plugins plus a setup wizard that detects your hardware and tunes the pipeline:
/plugin marketplace add artk0de/TeaRAGs-MCP
/plugin install tea-rags-setup@tea-rags
/tea-rags-setup:install
/plugin install tea-rags@tea-ragsAny MCP client (Cursor, Roo Code, Continue, β¦):
npm install -g tea-rags{
"mcpServers": {
"tea-rags": {
"command": "tea-rags",
"args": ["server"],
"env": { "CODEGRAPH_ENABLED": "true" }
}
}
}Qdrant downloads and starts on first use. Cloud embeddings (OpenAI, Cohere, Voyage), an external Qdrant, and the built-in ONNX provider (beta) are covered in the installation guide.
πΈοΈ Enable the call graph
The call graph is off by default while it is in beta. Turn it on with
CODEGRAPH_ENABLED=true in the MCP server's environment β the JSON above
already does β or, in Claude Code:
claude mcp add tea-rags -s user -e CODEGRAPH_ENABLED=true -- tea-rags serverThen reindex. The flag is recorded per project, so later runs from the CLI or auto-update keep the graph on. Details: Codegraph Enrichments.
π Quick Start
tea-rags index-codebase /path/to/repo --name myrepo # first index: register + index
tea-rags prime /path/to/repo # index state, drift, signal thresholdsIn Claude Code, /tea-rags:index does the same. Then ask your agent:
"How does auth work in this project?"
"Find stable examples of retry logic I can copy."
"What breaks if I change the payment module?"
π€ Why TeaRAGs?
| Embedding search | TeaRAGs | |
Finds | Exact text | Similar code | Similar code, ranked by evidence |
Knows history | β | β | Churn, bug fixes, owners, age |
Knows callers | β | β | Fan-in, transitive impact, call paths |
Ranks for the task | β | Similarity only | 23 presets β see What It Answers |
Cost on a large monorepo | Many agent turns | One query | One query |
π Compared to other tools
TeaRAGs is not a coding agent. It is the context layer an agent queries, so the closest comparisons are the tools that hand a codebase to an LLM. Every competitor cell links to that product's own documentation, checked on 2026-09-23; "β" means the capability is not in those docs.
Ranks by git history | Semantic search | Call graph | Serves context over MCP | Runs locally | Rerank presets | |
TeaRAGs | β churn, bug-fix rate, ownership and age, per file and per chunk | β dense + hybrid (BM25) | β callers, callees, cycles, AβB paths | β 28 tools | β embedded Qdrant and DuckDB, local embeddings | β 23 |
β | β | β οΈ internal only: a file dependency graph ranks the repo map it sends to the LLM | β | |||
β οΈ orders files by git change count inside the packed file | β | β | β CLI | β | ||
Sourcegraph (incl. Cody Enterprise) | β οΈ commit and diff search; no ranking by history documented | β
| β οΈ your Sourcegraph instance, self-hosted or cloud | β |
Two names that usually come up here changed shape. Cody Free and Cody Pro shut down on 2025-07-23 (Sourcegraph); Cody Enterprise continues and uses Sourcegraph Search as its context source, which is why it shares the Sourcegraph row. GitHub ended the Copilot Workspace technical preview on 2025-05-30 (GitHub Next).
A wider table against other MCP code-search servers (claude-context, serena, grepai, β¦) lives in the comparison guide.
βοΈ How It Works
%%{init: {"theme": "base", "themeVariables": {"primaryColor": "#fdf8e7", "primaryTextColor": "#2d2d2d", "primaryBorderColor": "#d4af37", "lineColor": "#c4941f", "secondaryColor": "#f5f5dc", "tertiaryColor": "#fafafa", "mainBkg": "#fdf8e7", "secondBkg": "#f5f5dc", "nodeBorder": "#d4af37", "clusterBkg": "#fffdf6", "clusterBorder": "#d4af37", "titleColor": "#2d2d2d", "edgeLabelBackground": "#ffffff", "fontSize": "15px"}}}%%
flowchart LR
User([π€ You])
Agent[π€ Coding agent<br/>+ TeaRAGs skills]
subgraph pkg["π΅ tea-rags"]
MCP[π MCP server<br/>28 tools]
CLI[β¨οΈ CLI<br/>index Β· prime Β· projects Β· auto-update]
Core[βοΈ Core<br/>chunk Β· enrich Β· search Β· rerank]
MCP --> Core
CLI --> Core
end
subgraph storage["π» Local storage"]
Qdrant[(ποΈ Qdrant<br/>embedded Β· vectors + signals)]
DuckDB[(π¦ DuckDB<br/>embedded Β· call graph)]
end
Embeddings[β¨ Embeddings<br/>Ollama Β· OpenAI Β· Cohere Β· Voyage]
Repo[π Your repo<br/>code + git history]
User <--> Agent
Agent <--> MCP
User --> CLI
Core <--> Qdrant
Core <--> DuckDB
Core --> Embeddings
Core --> RepoYour agent calls TeaRAGs over MCP; you run the CLI to index and maintain. Both
drive one core: it chunks code on AST boundaries, embeds each chunk, attaches
git and call-graph signals, and ranks results by the preset the task asks for.
Qdrant and DuckDB run embedded under ~/.tea-rags β no Docker, no servers to
manage.
β‘ Indexing Speed
Measured on a production monolith β 3.3M lines of Ruby and TypeScript, tests included, plus 162K lines of docs in ~34k files β and on TeaRAGs itself. Search is available as soon as the embeddings are stored; git and call-graph enrichment keep filling in behind it.
What | Time | Setup |
π Full | ~35β40 min Β· 3.3M lines of Ruby + TypeScript, tests included + 162K lines of docs Β· ~34k files | LAN GPUΒΉ Β· from the measured embedding throughput |
πΈοΈ TypeScript call graph rebuild on the monolith | 165 s |
|
π Incremental reindex of the monolith | seconds for a commit, 5β8 min for a week of edits | only changed files are re-embedded |
π΅ Full | 9 min wall Β· 686K LoC + 290K lines of docs Β· ~3.5k files | LAN GPUΒΉ Β· 97% of it is embeddings |
π Agent finds a bug's root cause on the monolith | ~40 s vs 10+ min with grep | same question, same agent |
ΒΉ Ollama on a LAN mini-PC with an AMD RX 7800M eGPU (ROCm), measured 2026-09-27. Ollama 0.34.4 (auto-updated from 0.24.0 that day) embeds one chunk per GPU pass through a single llama-server slot, which keeps the GPU about 55% busy β expect roughly a quarter to a third faster once that is fixed. Direct support for the llama.cpp server, as an alternative to Ollama, is coming soon.
For scale: in January 2026, on the smaller monolith of the time, the server TeaRAGs was forked from needed 4β10 hours for a full index and 40+ minutes to catch up on 100 commits.
Why it is fast. Every stage runs in parallel: 50 files in flight,
tree-sitter parser and git-blame worker pools, GPU batches of up to 512 chunks,
and the call graph in its own DuckDB process under a hard 2 GB memory cap, with
SCC and PageRank computed as streams. Embeddings dominate a full rebuild, so a
change to git or call-graph signals is recomputed with --force-enrichments
without re-embedding a single chunk β minutes instead of a full reindex. Each
run picks its embedding batch size and concurrency itself β backing off when the
server fails on a batch, climbing toward the fastest measured size, remembering
the result per endpoint β within bounds that tea-rags tune measures for your
hardware in about 90 seconds; details in
Performance Tuning.
π Measured
Call-graph quality is checked against independent oracles, not eyeballed.
What | Result | Corpus |
π Python call graph vs. jedi + pyright (pyright tie-break) | recall 0.92β1.00 Β· wrong edges β€ 0.29% | flask, httpx, netbox, polar |
π Ruby call graph, YARD-annotated | in-project recall 1.00 Β· 0 fabricated | octokit.rb |
π Ruby call graph, un-annotated Rails | bare-call recall 0.93 | mastodon |
π Ruby call graph, production Rails | in-project recall 87.7% | 3.5M-line production monolith |
π¦ TypeScript call graph vs. the TypeScript type checker | phantom edges 0.32% Β· agreement 72.8%ΒΉ | 17k-file production React frontend |
π¦ TypeScript call graph on TeaRAGs' own source | fabricated edges 93 β 0 | tea-rags |
π§ | +71 pp mean pass rate | 136 eval cases, 10 wrappers |
π©Ή Healing a drifted index instead of recomputing it | 113 ms | 134k-point production index |
ΒΉ About two thirds of the TypeScript gap is callbacks passed through props and dependency injection β the type checker names a function type there, not an implementation, so no static resolver can pin those edges.
The Python oracle harness ships in the repo
(scripts/py-codegraph-jedi-oracle.ts), so those numbers can be reproduced on
your own corpus.
Languages Compatibilities
Support: π maximum Β· π full Β· π high Β· π medium Β· π moderate Β· π partial/low Β· π minimal Β· π none
What tea-rags supports per language and at what level. AST chunking is how
source is split into searchable chunks; Test chunking is how faithfully test
structure is preserved; Codegraph is the call-graph resolution ceiling (the
realized per-project number lives in the tea-rags prime digest, not here).
Rows are ordered by overall capability, richest support first.
Language | AST chunking | Test chunking | Codegraph |
TypeScript | π full Β· tree-sitter (comment attachment, method-body splitting, describe/it scopes) | π high Β· testScopeChunker (describe/it scopes, one addressable chunk per example) | π high β 14-strategy chain (10 tree-sitter + 4 ts.Program/typeChecker) + cone dispatch + typeChecker-backed union-receiver fan-out |
JavaScript | π full Β· tree-sitter (assignment chunking, describe/it scopes, module/class split) | π high Β· testScopeChunker (describe/it scopes, one addressable chunk per example) | π high β 6-strategy; CommonJS/ESM require resolution (dynamic gaps) |
Ruby | π full Β· tree-sitter (RSpec block grouping, comment attachment, spec scope splitting, method-body splitting) | π high Β· RSpec scope chunker (one chunk per example, ancestor setup injected) | untyped π high Β· YARD π maximum Β· RBS/Sorbet π TBD β 15-strategy chain + 4 dispatch components + 20-grammar DSL catalogue + YARD type-source + db/schema.rb column accessors |
Swift | π full Β· tree-sitter | π high Β· XCTest + swift-testing recognition (test cases, setUp/tearDown, @Test/@Suite) plus Quick/Nimble DSL scope chunking (per-scenario chunks with ancestor beforeEach spliced in) | π high β 10-strategy chain + superclass dispatch + field and return-type receiver typing + nested-type and module-value receivers; no import narrowing |
Python | π full Β· tree-sitter | π medium Β· generic AST | π high β 10-strategy chain + C3 MRO + cone/union/callable-param/dict-table dispatch + re-export-aware imports + inferred return types + gated framework vocabularies |
Go | π full Β· tree-sitter (func/type split) | π medium Β· generic AST | π moderate β 7-pass chain + scope-aware typed locals + struct-field chains + embedding promotion + go.mod module-path imports; no interface dispatch |
Java | π full Β· tree-sitter | π medium Β· generic AST | π moderate β 6-strategy + java.lang stdlib whitelist + overload disambiguation |
Rust | π full Β· tree-sitter (named-item extraction) | π medium Β· generic AST (#[test] attrs not preserved) | π moderate β 7-strategy; trait-based dispatch |
Bash | π full Β· tree-sitter | π low Β· generic AST (bats/shunit not recognized) | π minimal β function-call extraction only, no dispatch |
Markdown | π full Β· MarkdownChunker (ToC + smart chunking) | π N/A Β· doc-only | π none β no call graph |
sql | π none Β· CharacterChunker | π N/A | π none |
jsonc | π none Β· CharacterChunker | π N/A | π none |
json | π none Β· CharacterChunker | π N/A | π none |
MCP clients
Every language above works the same way in every client β the client only decides how much of the tooling on top of the MCP server you get.
Client | MCP tools | Routing guide ( | Skills ( | Setup wizard |
Claude Code | β stdio | β plus the plugin's routing rules | β plugins | β
|
Other stdio clients (Cursor, Roo Code, Continue, β¦) | β
| β where the client reads MCP resources | β plugins are Claude Code only | β manual install |
HTTP clients | β
| β where the client reads MCP resources | β | β manual install |
Embedding providers
Set EMBEDDING_PROVIDER; EMBEDDING_MODEL overrides the default model.
Provider |
| Where it runs | Default model | Needs |
Ollama (default) |
| Local |
| A running Ollama |
llama-server |
| Local / LAN GPU host |
| A running llama-server per GPU; recommended for 3M+ lines |
ONNX (beta) |
| Local, built-in runtime |
| Nothing |
OpenAI |
| Cloud |
|
|
Cohere |
| Cloud |
|
|
Voyage |
| Cloud |
|
|
Throughput per provider and how to choose: Embedding Providers.
Ollama on another machine, the laptop as the fallback. Ollama does not have to run where the agent does. Put it on any computer in your local network β a desktop GPU, a home server, a spare Mac β and keep the laptop's own GPU or Apple chip as the fallback. TeaRAGs switches to the fallback when the primary stops answering or fails three embed calls in a row, and switches back once the primary is healthy again. Your code and index stay on the laptop; only chunk text crosses your own network.
EMBEDDING_BASE_URL=http://gpu-box:11434 # primary: any machine on your LAN
EMBEDDING_FALLBACK_URL=http://localhost:11434 # fallback: the laptop itselfDetails: Ollama provider.
Embedding model comparison
Measured on one GPU host and three code corpora (TypeScript, two Ruby); quality is dense MRR on 200 identifier-free natural-language queries per corpus.
Model | Role | Speed vs jina | Quality vs jina v2 code |
CodeRankEmbed (137M) | Recommended for llama-server | 0.93Γ | +0.06 to +0.19 MRR on all three corpora |
jina v2 code (161M) | Ollama default, baseline | 1.00Γ | 0.850 TypeScript Β· 0.832 / 0.683 Ruby |
Muninn-small (47M) | Fast option, no Ruby | 2.37Γ | +0.04 on TypeScript; mixed on Ruby (+0.08 / β0.05) |
Models of 1.5Bβ7B parameters gain mostly at R@1 and run 15β110Γ slower: every model finds the target in the top 10 on 99β100% of the TypeScript queries, and an agent reads the whole top-10 page. Which model to pick: Embedding model choice; the method and all tables: Embedding model comparison.
β¨οΈ CLI
Command | What it does |
| Index or incrementally update a codebase, with live progress |
| Markdown digest of index state, drift and signal thresholds |
| Manage the project registry: |
| Keep a project's index fresh on its target branch |
| Per-worktree index clones for parallel branches |
| Infrastructure and registry health |
| Recover a failed Qdrant optimizer without restarting the daemon |
| Auto-tune performance parameters for your hardware |
| Check for and install a newer version |
| Start the MCP server |
| Run one MCP tool in-process β check the tool surface from a shell |
π FAQ
How is this different from Aider or Copilot? Those are coding agents; TeaRAGs is what an agent asks before it writes. It indexes the repository once, keeps the index fresh, and answers over MCP with code plus its history and call graph. Any agent that speaks MCP can use it. See Compared to other tools.
Does it need the cloud? No. Qdrant and DuckDB run embedded under
~/.tea-rags, and the default embeddings come from a local Ollama (or the
built-in ONNX runtime), so code never leaves your machine. The network is used
to download the Qdrant binary on first run and for a cached npm version check in
tea-rags prime. OpenAI, Cohere and Voyage are opt-in.
How big a repository can it handle? The largest measured index is a production monolith of 3.3M lines of Ruby and TypeScript, tests included: ~34k files, 175k chunks, 2.0 GB of vectors and ~460 MB of call graph (see System requirements). After the first run, reindexing is incremental β only changed files are re-embedded.
Which languages are supported? Nine languages get AST chunking and a call graph of varying depth: TypeScript, JavaScript, Ruby, Python, Swift, Go, Java, Rust and Bash. Markdown is chunked by heading; SQL and JSON fall back to plain character chunks. Per-language depth is in Languages Compatibilities.
How accurate are the git signals? They are read from your real history; the
one heuristic is bug-fix detection. A commit counts as a bug fix when its
message says so (fix:, [Bug], TICKET-123 Fix β¦, fixes #123) or it
arrived through a merged fix/, hotfix/ or bugfix/ branch; "fix typo", "fix
lint" and similar are excluded. Chunk-level history follows each chunk's lines
through diff hunks and looks back 6 months by default (12 for file level); files
over 5,000 lines get file-level signals only. Labels such as high or
concerning are percentiles of your own repository, and signals backed by only
a few commits are dampened before they affect ranking.
π Documentation
I want to⦠| Start here |
Get it running | Quickstart β install, index, first query |
Understand the concept | Core Concepts β vectorization, trajectory enrichment, reranking |
See what my agent can do | Skills β the agent workflows and when each one fires |
Keep the index fresh | |
Look under the hood | Architecture β pipelines, data model, reranker internals |
Learn the theory | Knowledge Base β RAG, code search, software evolution |
π From the Blog
Engineering notes behind the releases, each with the corpus it was measured on β all posts Β· RSS.
β Star History
π€ Contributing
See CONTRIBUTING.md for workflow and conventions.
π Acknowledgments
Started as a fork of mhalder/qdrant-mcp-server β clean architecture, solid tests, open-source spirit β and its ancestor qdrant/mcp-server-qdrant. Code vectorization inspired by claude-context (Zilliz).
Feel free to fork this fork. It's forks all the way down. π’
βοΈ License
This server cannot be deployed
Maintenance
Related MCP Connectors
An MCP server that gives your AI access to the source code and docs of all public github repos
Capability registry for the agentic economy. Semantic search over verified MCP server listings.
Repository knowledge graph MCP server for codebase understanding and debugging.
Related MCP Servers
- AlicenseAqualityDmaintenanceMCP server with local vector search for your codebase. Smart indexing, semantic search, Git history β all offline.7121 PyPI49MIT
- AlicenseAqualityAmaintenanceLocal MCP server for semantic code search using Tree-sitter AST parsing, local embeddings, and hybrid search; enables indexing and querying codebases entirely offline.5MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for semantic code search and dependency graph analysis. Indexes codebases into a knowledge graph with vector embeddings for AI-powered code understanding.19 npmMIT
- AlicenseAqualityDmaintenanceMCP server for semantic code search that indexes your codebase and allows AI editors to search using natural language queries.958 npm53MIT