defrost-ai
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@defrost-aihow does the retry policy work and where is it documented?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ask a question in plain words and get back the doc sections that answer it, quoted with file and line numbers, together with the code each section names. Agents use it through MCP or the CLI. Nothing leaves your machine and there is no per-query API cost.
$ defrost search "how do I bind values to a structlog logger so they show up in every message?"
3 sections (accurate: reranked, the retrievers disagreed)
[1] e2e:structlog/docs/processors.md:L39-90 processors.md > Processors > Chains > Examples
### Examples If you set up your logger like: structlog.configure(processors=[f1, f2, f3])
log = structlog.get_logger().bind(x=42) and call log.info("some_event", y=23), it results in …Finds the answer, not the keyword | The answering doc section is in context for 94% of questions, against 6% for stock graphify (E2E) |
Free and private | Builds and searches on your laptop: no API calls, no tokens, no data leaving the machine |
Stays fresh by itself | Refreshes on merge to main or on a schedule, and re-embeds only what changed |
Coming fromkev-memory or defrost-ai 1.1? Old commands, MCP tool names, KEV_MEMORY_* variables and installed
hooks keep working; they are just no longer listed. Run defrost setup again to switch a repository to the new hooks.
One change in meaning: the mode 1.1 called fast (rerank on disagreement) is now accurate, the default; fast
now means no reranker.
Quick start
One command installs the CLI, downloads the weights and registers the Claude Code integration for all projects:
curl -fsSL https://raw.githubusercontent.com/Signaturi4/defrost-ai/main/install.sh | shIt needs uv. It installs defrost-ai (the defrost command) as a uv tool (Python 3.12),
fetches the adapters (about 140 MB, checked by sha256) to ~/.cache/defrost-ai, and registers the MCP server and
slash commands with Claude Code. It installs the release named in the script (a pinned tag, not the moving main); set
DEFROST_VERSION=x.y.z for another release or DEFROST_REF=main for the development branch. defrost --version
shows the installed code and weights versions.
Then open Claude Code in any repository and type /defrost-setup, or run it in a terminal:
defrost setup # asks 3 questions; press Enter for the recommended answer
defrost setup --yes # no questions, recommended answersThe three questions:
How much should defrost do? (
--profile)profile
what you get
minimalthe index refreshes after every merge or commit to main. Nothing else.
standard(recommended)minimal + the memory searched for each question + doc-writing rules in
CLAUDE.md+ a handoff note shown after/clear+ a reminder at session start of commits whose docs need an updatefullstandard + Claude is asked to update docs before it finishes and before it commits + a handoff note is written automatically before compaction
How far should Claude trust your docs? (
--doc-trust). Required: setup asks it for every new project and never picks it silently;--yeswithout--doc-truststops and prints the suggestion for the repository.defrost setup --suggest-trustshows it:highfor a docs repository (at least 10 doc files per code file), elselow. Re-running setup keeps the stored level.setting
for
what Claude does with a hit
low("code is the truth")code that changes daily, few docs
treats the section as a hint, reads the
verify in:files, answers from the code and lists doc/code conflictshigh("docs are reliable")docs repositories, legacy or well-documented projects
answers from the section; reads code only when a hit carries a
!stale or conflict lineSearch the memory automatically for each question? (
--prompt-context, on instandardandfull;--no-prompt-contextturns it off). A Claude CodeUserPromptSubmithook (defrost hook prompt) searches the memory (fast mode, top 3) and adds the sections to the prompt, so a lookup is answered in one model turn instead of three to five (measured on a docs repository: ~4 s with Sonnet instead of 18–30 s). It skips slash commands, prompts under 3 words and prompts whose best section is less similar thanprompt_context.min_cosine(0.34), and it never waits for a cold service.
The search mode is not asked: it defaults to accurate (see Search modes); --mode or
defrost config search.mode fast changes it.
You never run a build by hand. Refreshes run in the background, are incremental (only new or edited sections are
re-embedded) and log to ~/.defrost-ai/<domain>.refresh.log. The search index lives in ~/.defrost-ai/<domain>,
never in your repo. The only folder setup adds to your project is defrost-memory/ (handoff notes and decisions,
see below). Indexing follows .gitignore, and secret-like files (.env, *secret*, keys) are never indexed.
Setup also adds a short "Project memory" block to CLAUDE.md that tells Claude to use the memory; without it, agents
in our evals mostly ignored the MCP tools.
Want something between the profiles? Each option overrides its profile:
defrost setup --profile minimal --every-hours 6 # also refresh on a schedule (launchd / cron)
defrost setup --claude-hook # also refresh when a Claude session starts
defrost setup --no-doc-rules # keep CLAUDE.md free of the doc-writing rules
defrost setup --help # every option
defrost setup --remove # remove every hook and schedule defrost installedUpgrade: run the install line again (the resident service restarts itself on the new build).
Teammates. Commit what setup adds (CLAUDE.md, .claude/settings.json, docs/). The Claude hooks call
defrost from PATH and do nothing where defrost is not installed, so a fresh clone works in Claude Code right away.
Each teammate then runs the install line once and defrost setup --yes --domain <name> in their clone: the index and
the git hooks are per machine and never committed.
Related MCP server: MCP Codebase RAG Server
Search modes
There are two. Pick per question, or set your default once.
mode | speed | quality (locked test, nDCG@10) | how |
| ~1-2 s when it reranks, ~0.1 s when it does not | 0.838 | keyword (BM25) and meaning (Defrost-Ret-B) search run together. When their top hits agree, that answer is returned at once. When they disagree, the Defrost-Rerank model reads both top-20 lists and picks. |
| ~0.1 s | 0.776 | the same two searches, merged by rank. No reranker. |
Use accurate when the answer matters (an agent about to change code). Use fast for interactive lookups,
autocomplete or chatbots where 0.1 s matters more than the last few points. Times are for Apple Silicon (MLX).
defrost search "how is a refund issued?" # your default mode
defrost search "how is a refund issued?" --fast # this question only
defrost config search.mode fast # change your defaultResearchers can still ask for one retriever with --mode bm25|dense|hybrid|rerank|all.
Commands
command | what it does |
| set up this repository (3 questions) |
| the doc sections that answer it, with the code they name ( |
| what is indexed, how fresh it is, which settings are active |
| update the index now ( |
| which doc sections to update for your change ( |
| save a handoff note ( |
| show or change personal settings |
| the MCP server and the local HTTP service (started for you) |
Settings
Personal settings live in one commented file, ~/.defrost-ai/config.toml. defrost config shows every setting,
its value and where the value comes from; an environment variable, if set, wins over the file.
$ defrost config
search.mode accurate (default)
search.k auto (default)
project.doc_trust low (default)
models.backend auto (default)
...
$ defrost config search.k 3
$ defrost config search.k --resetsetting | values | env override |
|
|
|
|
|
|
| default for new projects: |
|
|
|
|
|
|
|
|
|
|
| reranker scores kept for repeated questions ( |
|
| port of the local search service (8765) |
|
Per-project doc trust lives in ~/.defrost-ai/<domain>.workspace.json; change it with
defrost setup --doc-trust high --build skip (no rebuild needed).
Use with Claude (MCP server + slash commands)
install.sh already did this for all projects. To register the server in another MCP client (Claude Desktop,
Cursor):
claude mcp add defrost -- defrost mcpFour MCP tools:
tool | what it does |
| doc sections with |
| the doc sections that describe the changed files, so docs are updated in the same change |
| save a handoff note, or your decision on a doc/code conflict |
| update or build an index; |
Slash commands:
in Claude | what it does |
| the 3 setup questions, then setup |
| searches, then answers with |
| saves a handoff note before |
| updates the docs that describe your change |
The MCP server is a thin stdio process with no ML dependencies. Searches go to one resident local service
(defrost serve, started on first use), so the models load once for all clients.
The service listens on 127.0.0.1 only and needs a token: on start it writes ~/.defrost-ai/service-<port>.json
(readable by you only), and the CLI, the MCP server and the graphify patch send it as Authorization: Bearer.
Requests without it get 401; requests from a web page (an Origin other than localhost, or a Host other than
127.0.0.1/localhost) get 403. GET /health needs no token. Rebuilding a domain does not stop searches: builds take
a per-domain lock and share the model with searches in small chunks.
There is one service per install. After an upgrade, the first call replaces a service that runs older code of the
same install. A service started by another install (say a repo .venv next to the uv tool) is used as it is and
never stopped. An MCP server started before the upgrade gets 401 until you restart Claude;
DEFROST_SERVICE_AUTH=0 on the service turns the token check off for that transition.
Working memory: a git-backed context repository
Handoff notes (/handoff, remember(kind="note"), defrost note) and your doc/code conflict decisions are stored
as small Markdown files in defrost-memory/ inside your project, so you can read them next to your code. The
folder is its own git repo: your project's git does not see it (setup adds it to .git/info/exclude), and the
project's index skips it. Every write is one commit, so you can audit what the agent remembered and why
(defrost note --history). Search covers it as the domain <domain>-context, and defrost status shows it under
its project. The layout and the pre-commit validation follow Letta Code's context repositories.
Setup options --memory-dir DIR and --memory-home move it. Details: docs/LETTA_CONTEXT_REPOS.md.
With graphify
graphify users can get the same search inside graphify's own CLI and MCP server (patch on v0.4.32):
git clone https://github.com/safishamsi/graphify && cd graphify && git checkout v0.4.32 \
&& git apply ../defrost-ai/integrations/graphify/graphify-0.4.32-defrost-ai.patch && pip install -e ".[mcp]"
export DEFROST_SERVE_CMD="defrost serve"
graphify memory init . --domain my-repo && graphify claude install
graphify memory search "how are refunds issued?" -k autoPython:
from defrost_ai import Library
res = Library().search("how does the feed reach mobile") # accurate
res = Library().search("how does the feed reach mobile", mode="fast")
for hit in res["hits"]:
print(hit["domain"], hit["path"], hit["lines"], hit["heading"], [c["label"] for c in hit["code"]])Grounded in the code, not just the docs. Docs go stale, so every hit carries the files to check it against:
[2] my-repo:docs/DEVOPS.md:L51-64 DEVOPS.md > Known problem: the CI deploy never runs
...
-> code deploy.yml (my-repo/.github/workflows/deploy.yml)
! doc may be stale: my-repo/.github/workflows/deploy.yml changed 2026-09-29, after this doc (2026-09-26)
! doc/code conflict: names not found in the code: `scripts/old_deploy.sh`
verify in: my-repo/.github/workflows/deploy.ymlWhat gets linked: config, CI and infra files (compose files, Dockerfiles, workflows, crontabs, shell scripts), as well as code symbols.
Stale docs: a hit is flagged when a file it names was committed after the doc was.
Conflicts: a hit is flagged when the doc names a file or function that no longer exists. The agent does not pick a side: it shows the doc and the code and asks you ("code is right", "doc is right", "not a conflict", "not sure"). Your answer is recorded (
remember(kind="decision"),defrost note --conflict DOC --verdict code) and shown on later hits as aresolved:line, so each conflict is asked once.Claude's instructions: the CLAUDE.md block tells Claude to read the
verify in:files before stating how something behaves, to trust the code when the two disagree, and to list the doc/code conflicts it found.Docs to update: after a change,
docs_for(MCP) ordefrost docslists the doc sections that describe the changed files.
FAQ / support chatbots: see docs/FAQ_CHATBOT.md (use fast, or --mode hybrid directly).
How it works
The agent keeps its usual loop (think, call a tool, observe). defrost-ai adds the tools for how/why questions, checks every doc hit against the code, asks you when the two disagree, and keeps working notes in a git repo inside your project. Mermaid sources and the rules behind each step: docs/AGENT_LOOP.md.
Why
Code graphs such as graphify are good at structure: which function calls which, what lives in which module. They are weak at the question developers ask most: "how do I…", "why does…", "what happens when…". The answer to those is usually a paragraph in the docs, and a keyword query over node labels rarely reaches it. On three repos our models never saw, graphify's query put the answering doc section in its context for 6% of questions; this memory did for 94% (details below).
The usual fix is a hosted embedding API or an LLM pass over every document, which costs money on every refresh and sends your code out. This project trains small models instead (a 0.5B-parameter backbone) that run on a laptop.
Results
Every choice was made on dev splits; the test splits were scored once. Intervals are paired-bootstrap 95% CIs. Protocol: docs/EVALUATION.md. Everything that worked and did not: docs/RESULTS.md.
Same comparison, agent and cost (docs/E2E_GRAPHIFY.md):
stock graphify | graphify + defrost-ai | |
build cost for the 3 repos | $11.87 of Claude usage | $0, runs locally |
Claude Sonnet agent with file tools: accuracy | 0.906 | 0.922 (n.s.) |
Claude Sonnet agent: cost per question | $0.094 | $0.071 (−25%, CI excludes 0) |
A 13-gram gate removed generated training questions that overlap an eval suite. It did not cover the unsupervised
pretraining text: part of the private product repos' docs was in that corpus, so the private-repo numbers are
in-domain, not held-out. The e2e repos and the held-out OSS repos were not in any training corpus we built. The e2e
and RAGAS numbers were measured with Defrost-Rerank v1; v2 adds +0.024 [+0.010, +0.041] nDCG@10 to accurate on the
locked test.
On repos this small, a strong agent with plain grep also scores 0.906 and is the cheapest arm. The memory matters most where grep stops working: large or multi-repo corpora, docs kept apart from code, weaker or cheaper answer models, and fixed context budgets.
Why not a pure knowledge graph? The graph is reliable for structure and doc→code links, not as the place answers come from: instructions keep their conditions and exceptions in paragraphs. See docs/DESIGN_NOTES.md.
Known weaknesses: the dense retriever alone loses to BM25 on private product docs; a reranked search takes about 1.7 s on an Apple M5 with MLX (2.3 s on the torch path); a 568M public reranker still beats ours on long narrative prose (books 0.949 vs 0.902); multi-hop questions are unsolved.
Use cases
Coding agents (Claude Code, Codex, Cursor…):
searchas an MCP tool next to graphify's graph tools. The agent gets the doc section and the code it names in one call.RAG over internal docs: handbooks, runbooks, ADRs, READMEs across many repos, with citations to path and lines.
Company knowledge base, several domains at once: register each repo or folder as a domain; one query searches all of them and merges the results with the reranker.
Always fresh: graphify's git hooks refresh the memory after every commit and branch switch. Only changed sections are re-embedded, and the previous build is kept for rollback.
Reviewable updates and memory-grounded agents: two Shepherd tasks keep a memory refresh as a reviewable result (accept, or discard and roll back) and answer questions from cited context in a sandbox.
Built on
component | origin | used for |
Qwen2.5-0.5B (rev | Alibaba Qwen, Apache-2.0 | the backbone of both models |
MNTP + CGSA recipe (KG-BiLM / LLM2Vec) | McGill NLP, MIT ( | turning the causal decoder into a bidirectional text encoder: masked next-token prediction, then contrastive sentence alignment |
Defrost-Ret-B (ours) | LoRA r16 on the backbone, contrastive training on 57k (query, passage, hard negative) rows: MS MARCO, NQ, HotpotQA, AllNLI, Quora, StackExchange + 9.7k tech-doc questions | dense retrieval of doc sections |
Defrost-Rerank v2 (ours) | same backbone + LoRA + score head, listwise loss over 1 positive + 7 negatives (BM25, same-file siblings, changelog sections) | reordering the top 40 candidates |
SQLite FTS5 BM25 | SQLite | keyword retrieval: exact identifiers, flags, error strings |
| no parameters | uses the cheap fusion when BM25 and the dense retriever agree on the top section, the reranker when they disagree (about half the queries) |
adaptive k (ours, optional | temperature-scaled dense confidence | sends 1–5 sections: 18% fewer context tokens at the same hit rate |
graphify 0.4.32 | safishamsi/graphify, MIT | tree-sitter AST code graph, git hooks, MCP server, CLI. We add a patch ( |
Shepherd ( | shepherd-agents | sandboxed agent tasks with retained, reviewable outputs |
RAGAS 0.4.3 (NVIDIA metrics) | explodinggradients/ragas | answer-level evaluation: accuracy, context relevance, groundedness |
What is different from the parts it is built on:
graphify indexes code structure and, in its paid semantic tier, uses Claude to extract concepts from docs. This project indexes every doc section locally, links each one to the exact code nodes it names (from graphify's own AST graph), and ranks sections with trained models. It reuses graphify's graph, hooks and MCP server rather than replacing them.
Off-the-shelf embedders (bge-small and similar) are trained on web text. Defrost-Ret-B starts from a backbone adapted to technical prose and is trained on developer questions about documentation. Same-sized rerankers trained on web data scored lower on our held-out repos (0.717 for bge-reranker-base vs 0.890 for Defrost-Rerank v2).
Writing docs the memory reads well
/defrost-setup offers this kit, or run defrost setup . --doc-rules (the kit is
defrost_ai/assets/doc_rules/, also linked as templates/doc-rules/). It appends a short, highlighted rule block to the end of CLAUDE.md and installs:
the full rules, read only when docs are written, so they don't fill every session;
a glossary template;
a linter;
a Facts extractor that turns
- Subject → relation → Objectlines into triples.
The research behind the rules: docs/WRITING_FOR_EXTRACTION.md.
Weights
The LoRA adapters (MNTP, CGSA, Defrost-Ret-B, Defrost-Rerank v2 + score head, about 140 MB) are attached to the
v1.1.0 release. install.sh (or
defrost download-weights, or the first search) fetches them to ~/.cache/defrost-ai/models and checks the
archive's sha256. If they cannot be fetched, search stops with an error that names the version and the fix
(it does not fall back to older weights). defrost status shows the version in use and the size of the
merged-weights cache (~/.cache/defrost-ai/merged, fp32, about 1.8 GB per model; caches for other weights are
removed after 7 days unused, on download-weights and on service start). By hand:
curl -L -o weights.tar.gz https://github.com/Signaturi4/defrost-ai/releases/download/v1.1.0/defrost-ai-weights-v1.1.0.tar.gz
tar xzf weights.tar.gz # -> models/ (sha256 of the archive: 494fab8997ce01c4…)
defrost verify-weights # checks every file against models/MANIFEST.jsonThey load on top of Qwen/Qwen2.5-0.5B at revision 060db649 (downloaded from Hugging Face on first use).
You can also keep them elsewhere and set DEFROST_MODELS. training/ holds the exact scripts and configs
used to train them, stage by stage.
Layout
defrost_ai/ library: ingest (sections, code graph, doc->code links), models, retrieval, builder, service, CLI
integrations/ the graphify patch
benchmarks/ frozen question suites (held-out repos, books, e2e) + the e2e harness
training/ training scripts and configs (MNTP -> CGSA -> Defrost-Ret-B / Defrost-Rerank), Kaggle notebooks
scripts/ parity check, benchmark source fetcher, weight export, reranker efficiency, README charts, brand assets
docs/ EVALUATION, RESULTS, ARCHITECTURE, AGENT_LOOP, E2E_GRAPHIFY, DESIGN_NOTES, WRITING_FOR_EXTRACTION
docs/brand/ app icon, logo, favicons, social preview (BRAND.md has the rules)
templates/ doc-rules kit for CLAUDE.md (rules, glossary, linter, Facts extractor)Reproduce
python scripts/fetch_benchmark_sources.py # pinned commits of the benchmark repos and books
defrost build benchmarks/heldout/workspace.json
defrost benchmark --suite benchmarks/heldout/questions.jsonl --memory ~/.defrost-ai/benchmark-heldout --split devExpected on held-out dev (nDCG@10): bm25 0.711, dense 0.885, hybrid 0.789, rerank 0.890, accurate 0.887 (weights v1.1.0; with v1.0.0: rerank 0.870, accurate 0.870).
License
Code: MIT; parts of the context repository are ported from Letta Code (Apache-2.0, see NOTICE). Model adapters: LoRA weights on Qwen2.5-0.5B (Apache-2.0). training/source/kg_bilm_experiments: MIT
(McGill NLP). The benchmark books keep their own licenses (Pro Git CC BY-NC-SA 3.0, Eloquent JavaScript CC BY-NC,
500 Lines or Less CC BY 3.0); only questions and line references are included, not the texts.
This server cannot be deployed
Maintenance
Related MCP Connectors
A cited wiki of your GitHub repo: search, read pages, find symbols and ask, with line citations.
Search indexed code, trace dependencies, assess change impact, and recall repository memory.
- SeturosOAuthcom.seturos
Shared work memory for Claude Code, Codex, Cursor and chat, scoped to each repository.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA local, SQLite-backed code index for Claude Code, exposed over MCP, enabling targeted code retrieval without external APIs.101 PyPI1MIT
- FlicenseAqualityDmaintenanceProvides semantic vector search over local codebases via MCP, enabling hybrid search (dense + sparse + RRF) for any MCP client like GitHub Copilot or Claude Desktop.58-
- AlicenseNot gradedqualityDmaintenanceProvides local-first, cross-session memory for Claude Code, enabling semantic search across past sessions to retrieve procedures, decisions, or answers without exposing secrets.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceMCP server for local codebase memory: semantic indexing, hybrid search, and clean code analysis (static + LLM) with retrievable project notes, designed to run fully local on modest hardware.MIT