Skip to main content
Glama
phense
by phense

agentic-rag

Provider-neutral long-term memory and compaction continuity — in a real database.

Coding sessions end and long contexts compact. agentic-rag preserves both: a canonical, searchable knowledge base in local PostgreSQL + pgvector, plus bounded checkpoints that let Claude Code, Codex, and Antigravity resume after compaction. Hybrid vector + full-text search, lifecycle hooks, and a provider CLI you control — without a hosted RAG service.

License: MIT version Python 3.13 PostgreSQL + pgvector

New in v0.5.0: Antigravity CLI (agy, Gemini) compaction continuity — rag install --agy, SessionStart/PreInvocation/Stop hooks, /compact handoff, automatic-compaction detection, Antigravity transcript mining. Read What’s New in 0.5.0.

0.4.0: Claude compaction continuity — six Claude hooks, the managed 1M/500K autoCompactWindow policy, the compact_summary handoff, and rag install --check/--restore. Read What’s New in 0.4.0.

0.3.0: Codex compaction continuity, native-memory policy, recoverable global installation, and provider-neutral mining. Read What’s New in 0.3.0.

Most "RAG memory" tools are a cloud retrieval layer you feed documents to: you push, you query, you pay per call. agentic-rag flips both halves. It stores knowledge in local Postgres + pgvector — real HNSW approximate-nearest-neighbour search blended with bilingual full-text — and it populates itself from supported coding sessions. It also stores compact, audited continuation checkpoints at Claude Code, Codex, and Antigravity compaction boundaries — on Claude including Claude's own compact summary as a bounded handoff.

Every content write funnels through one gateway: it strips secret-shaped tokens, chunks and embeds the text with a local model, resolves the document's links into a typed knowledge graph, and logs the change — all in one transaction. When a session ends, a single-writer worker uses the configured Codex or Claude CLI to turn the bounded transcript digest into durable, findable memories. The core data model and provider seam are provider-neutral; integrations adapt each coding agent's lifecycle and output contracts.

It runs on your machine and uses your configured CLI account for LLM-assisted mining, curation, and bounded checkpoint enrichment: Codex with ChatGPT login or Claude with its supported OAuth or API-key authentication. Mining prompts can also include all matching pin bodies; mining secret-strips the provider-bound copies without mutating stored pin text. Embeddings are always local (Ollama), so retrieval does not call either provider. It's RAM-lean by design: no always-on daemon beyond Postgres and Ollama, and an idle footprint near zero between sessions. And it's built data-safety-first — it archives rather than deletes, writes through a least-privilege role matrix, audits every change, and periodically restore-tests its own backups.

Your data stays under your control, with explicit provider calls. This repository is code only — it ships no content. The canonical store lives in your PostgreSQL database, but the configured CLI intentionally sends these provider inputs: mining sends a bounded, secret-stripped transcript digest, live domain names, and secret-stripped copies of all matching pin bodies without mutating stored pin text; curation sends selected stored documents and contradiction evidence; checkpoint enrichment sends a secret-stripped transcript delta and validates the returned checkpoint content before persistence. Optional synced backups copy data only to a directory you configure. agentic-rag has no separate hosted RAG backend.

Why · Quick start · What's different · How it works · Comparison · Configuration · 📖 Handbook · Status · Acknowledgments · License


Why agentic-rag

🔎 Hybrid search that actually ranks. Vector ANN over pgvector (HNSW, cosine) — multilingual by way of bge-m3 embeddings — blended with GIN keyword full-text into one ranked query. Search in any language; not a file scan, not lexical-only.

🌱 It turns sessions into durable knowledge and continuation state. Mining uses your configured Codex or Claude CLI. On Claude Code and Codex, PreCompact also captures a fast checkpoint so SessionStart(source="compact") can restore the goal, blockers, next action, repository state, and evidence references; on Claude the checkpoint also carries Claude's compact summary.

♻️ It curates itself. A near-duplicate gate stops the store from bloating; rag review surfaces duplicates, dangling links, and stale pins; refuting a fact archives it (with a reason and evidence), never hard-deletes it.

🔒 Local-first, on your own account. Canonical knowledge and checkpoints live in your Postgres. LLM-assisted work runs through the local Codex or Claude CLI you configured. Embeddings are always local (Ollama), so search and retrieval do not call either provider.

RAM-lean. A single-writer worker (flock singleton), no long-lived daemon of its own. Between sessions the footprint is essentially Postgres + Ollama idling — nothing else.


Related MCP server: rawthink

Quick start

OpenCode / DeepSeek (unreleased source addition): install the native RAG adapter on the execution host with uv run rag install --opencode. It supports startup context, selective recall, checkpoint/handoff and idle mining. For T3 on Windows, install on the Mac execution host; canonical context and a real read-only RAG tool call were verified through T3. See the documented limits.

agentic-rag is a rag command-line tool with provider integrations. The no-option install wires two MCP servers, six lifecycle hooks, and the managed compaction window into Claude Code; the explicit Codex and Antigravity (--agy) targets install continuity configuration and hooks.

Prerequisites:

  • PostgreSQL 17 with the pgvector extension (the schema uses halfvec, pgvector ≥ 0.7).

  • Ollama with the embedding model pulled — ollama pull bge-m3 (1024-dim, fixed to the schema).

  • An authenticated LLM CLI: Codex (codex login) or Claude (claude -p). Claude/Haiku remains the package default for compatibility; select the provider in [llm].

  • uv and Python ≥ 3.13.

Install the common foundation, then choose an integration:

uv sync
uv run rag init-db          # creates the DB + schema + roles, seeds the 'general' domain
uv run rag domain add programming --description "Software engineering notes"
uv run rag install --check  # preview the Claude settings merge; writes nothing
uv run rag install          # Claude MCP/hooks + macOS backup schedule; omit for Codex-only
uv run rag install --agy    # Antigravity CLI (agy) hooks.json target; --check previews
  • rag init-db creates the database if needed, applies the migrations in sql/, creates the three least-privilege roles, and seeds the built-in general domain. Run it firstrag install does not create the database.

  • rag domain add <name> adds any domains you want to organize documents under (general always exists; add more anytime).

  • rag install --check previews the Claude merge (managed: autoCompactWindow=500000, the would-change path, policy warnings) and writes nothing.

  • The no-option rag install is the Claude target: it registers the agentic-rag (read-write) and agentic-rag-ro (read-only) MCP servers, merges six hooks (SessionStart, UserPromptSubmit, Stop, PreCompact, PostCompact, SessionEnd) plus autoCompactWindow = 500000 into ~/.claude/settings.json, backs the file up to a unique settings.json.bak.<id>, and prints a rag install --restore <record> rollback command. On macOS it also schedules nightly backup; omit this command for a Codex-only setup.

If you ran the Claude target, hooks reload live; start a new Claude Code session so it picks up the MCP servers, then review the handlers with /hooks and confirm /autocompact reports 500000 tokens from settings. For humans the same store is available through the CLI:

rag save --title "Postgres VACUUM tuning" --domain programming \
    --dtype lesson --body "autovacuum_vacuum_scale_factor tradeoffs..."
rag search "vacuum tuning" --domain programming
rag get <slug-or-id>          # body + incoming/outgoing graph edges
rag status                    # counts, queue health, last backup/curation

For Codex continuity, preview before touching your user configuration, install, then inspect and trust the changed commands in Codex:

uv run rag install --codex --check
uv run rag install --codex
# Start Codex, run /hooks, inspect all six agentic-rag commands, then trust them.
uv run rag status

The Codex transaction manages only ~/.codex/config.toml, ~/.codex/hooks.json, and ~/.codex/compact_prompt.md. It prints every changed path, backup, validation result, and a ready-to-run rollback command:

uv run rag install --codex --restore /absolute/path/to/codex-rollback-<id>.json

Use the exact rollback-record pathname printed by the successful install. Check mode writes nothing. Repository support does not mean a particular machine has completed hook trust or live compaction verification; review /hooks and confirm rag status after every installation.

The Codex target never installs a scheduler. On a Codex-only macOS setup, use uv run rag backup --install-launchd for scheduled database backups; Linux uses the handbook's cron/systemd recipes.


What makes it different

Four things that, together, set it apart from both file-based knowledge wikis and hosted RAG stacks:

1. Real hybrid search over a curated graph

Documents are chunked and embedded into halfvec(1024) columns indexed with HNSW, and each chunk is embedded with the multilingual bge-m3 model — so semantic recall works in any language — and also carries generated tsvectors for English/German keyword full-text. A single query runs vector ANN and full-text together and returns one ranked list. Documents aren't an undifferentiated pile: they carry a type (concept, lesson, signal, synthesis, reference, …) and connect through a typed edge graph (references, extends, depends_on, supersedes, contradicts, …), so rag get shows you not just a document but its neighbourhood.

2. Automatic session-mining — the star feature

This is what makes agentic-rag feel like it grows rather than sits there. When a supported coding session ends, a lifecycle hook enqueues the transcript. A single-writer worker drains the queue and calls the configured Codex or Claude CLI to pull out durable memories, lessons, and signals, each saved through the write gateway behind a near-duplicate gate. A fix you discovered today becomes something a future session can recall, with no "remember to write this down" step. It reads only your local session transcripts.

3. Knowledge domains you grow and curate

Domains are just data — a label for where to look (general is seeded at init). Add them with rag domain add, scope any search with --domain, and let the importer derive them from an existing store's topics. Curation is first-class: rag review reports near-duplicates, dangling links, and stale pins; refuting a fact archives it with a required reason + evidence; rag purge removes only already-refuted documents, and only as rag_admin.

4. Built-in maintenance, backup, and restore-testing

rag backup runs pg_dump -Fc locally (plus an optional copy to a synced directory you configure), with rotation. rag maintenance is a tiny, single-flight, always-exit-0 job that ticks the worker, rotates logs, and — weekly — runs a report-only restore-test: it restores your newest dump into an isolated scratch database, compares row counts, and drops it. A backup you've never restored isn't a backup; this one checks itself.


How it works

   Supported coding sessions                rag save · migrate import · MCP write tools
   (queued & mined on session end)          (you, or an agent)
            │                                               │
            └───────────────────┬───────────────────────────┘
                                │
                    one audited write gateway
       strips secret-shaped tokens · chunks + embeds (local Ollama) · resolves edges · logs
                                │
                ┌───────────────┴────────────────┐
        PostgreSQL + pgvector              typed knowledge graph
        documents · chunks halfvec(1024)   edges: references · extends · …
        HNSW ANN  +  EN/DE full-text
                                │
                one hybrid ranked search  ──  vector ⊕ full-text
                                │
   session-start context · prompt-time recall · read-only MCP behind a privilege boundary

Claude Code and Codex continuity use a separate operational path. The Claude flow:

PreCompact ──► bounded deterministic snapshot ──► audited checkpoint
     │                    └──► priority enrichment job ──► provider CLI
     └──► stdout: versioned compact instructions (+ checkpoint id)
Claude compacts (instructions appended)
     │
PostCompact ──► mark boundary + store compact_summary as bounded handoff
     │
SessionStart(source="compact") ──► checkpoint + handoff, ≤ 10,000 chars ──► next request

The Codex flow (PreCompact stays silent; PostCompact stores no handoff):

PreCompact ──► bounded deterministic snapshot ──► audited checkpoint
     │                    └──► priority enrichment job ──► provider CLI
     ▼
Codex compacts
     │
PostCompact ──► mark boundary only (cannot inject context)
     │
SessionStart(source="compact") ──► bounded checkpoint context ──► next request

Native Codex memories are complementary, not the canonical record. With the installed policy they remain enabled and can be inspected with /memories; agentic-rag is canonical for durable searchable knowledge, audit history, and explicit continuation checkpoints.

  • Your chosen CLI provider. Every LLM call goes through the single agentic_rag.llm seam and the configured local Codex or Claude command. Transcript digests and checkpoint deltas are character-bounded; each curation call covers one selected document/evidence set. Mining also sends secret-stripped copies of all matching pin bodies without changing the stored pins. Embeddings never leave the box (local Ollama), so retrieval is independent of provider authentication.

  • One audited write path. Every change — a manual save, a mined memory, an import — funnels through a single gateway that strips secret-shaped tokens, regenerates chunks + embeddings in one transaction, resolves dangling edges, and writes an audit row. Embeddings fail open (queued for retry if Ollama is down); nothing else does.

  • Least privilege, by role. Three login roles enforce a destruction-protection matrix: rag_reader (SELECT only, used by search and the read-only MCP), rag_writer (INSERT/UPDATE but no DELETE/TRUNCATE/DROP), and rag_admin (migrate, purge, restore).

  • It steps aside, not in front. If Ollama is down, search degrades to full-text-only and returns a warning rather than failing; the maintenance job always exits 0.

The full story is in the handbook — the mental model, everyday use, configuration, importing an existing wiki, and the architecture and design rationale.


Comparison

agentic-rag sits between two worlds: the file-based LLM-Wiki family (human-readable Markdown with a lint/graph layer) and typical RAG stacks (hosted or API-driven retrieval you feed documents to). Every cell below is marked honestly — including the rows where each of them beats us.

Legend: ✅ shipped in code and operationally established · 🧪 shipped and installed, live verification pending · ⚠️ partial / caveated · ❌ absent

vs LLM-Wiki systems (file-based knowledge wikis)

Capability

agentic-rag

File-based LLM-Wiki

Hybrid vector + full-text ranked search (ANN at scale)

⚠️ lexical/graph, file-scan

Bilingual full-text (EN + DE) + semantic recall

⚠️

Transactional, audited writes through one gateway

⚠️

Scales to a large corpus (HNSW index)

⚠️ file-scan slows

Human-readable, git-diffable plain-text store

⚠️ import/export MD; store is Postgres

Zero-infrastructure (no DB/service to run)

❌ needs Postgres + Ollama

Imports an existing llm-wiki store

rag migrate

✅ it is one

Bottom line: if you want a git-tracked pile of Markdown, a file-wiki wins on its home turf. If you want fast hybrid recall over a growing corpus with transactional safety, agentic-rag wins — and it can import your existing llm-wiki to get you there.

vs typical RAG systems (retrieval frameworks / hosted memory)

Capability

agentic-rag

Typical RAG stack

Local-first store, provider CLI under your control, no hosted RAG service ¹

⚠️ usually a hosted service

Auto-populates from Claude Code sessions (mining)

❌ you feed it

Codex session mining and continuity

🧪

❌ you feed it

Claude Code compaction continuity (checkpoint + handoff)

🧪

Self-curation (dedup, near-dup gate, refute/archive)

⚠️

Typed knowledge graph (edges) alongside vector search

⚠️

One audited write gateway with secret stripping

Read/write privilege boundary for subagents (RO MCP)

⚠️

Turnkey managed hosting / large ecosystem ²

⚠️ self-host, young

Bottom line: a hosted RAG stack wins on turnkey scale and ecosystem. agentic-rag wins on being local-first, using a provider CLI under your control with no hosted RAG service in the loop, self-populating-from-your-own-work, and self-curating — a memory that fills and tidies itself instead of one you have to keep feeding.


Configuration

Config lives in one TOML file at ~/.agentic-rag/config.toml. Every key is optional — omit a section to keep its defaults.

Setting

Default

What it does

[db] name

agentic_rag

Database name.

[db] host

"" (local socket)

Empty = local unix socket; set it for a networked/remote server.

[embed] model

bge-m3

Ollama embedding model tag.

[embed] dim

1024

Fixed to the schema (halfvec(1024)); init-db refuses a mismatch.

[ollama] url

http://localhost:11434

Local Ollama endpoint.

[backup] local_dir

~/.agentic-rag/backups

Where pg_dump archives are written.

[backup] cloud_dir

— (unset)

Opt-in copy to a synced/cloud directory; unset means backups remain local. Provider-bound LLM inputs are disclosed above.

[pg] bin_dir

auto-resolved

Only needed if pg_dump/pg_restore/psql aren't on PATH (e.g. Postgres.app).

Roles are created passwordless by default, relying on local peer/trust auth (Postgres and agentic-rag on the same machine). For a networked or shared instance, set role passwords with ALTER ROLE … and let libpq authenticate via ~/.pgpass or PGHOST/PGPORT/PGPASSWORD — see the handbook's privacy chapter.

The Claude target separately manages autoCompactWindow = 500000 in ~/.claude/settings.json — a 1M context compacting at 500K with a [1m] model; model is reported, never rewritten, and long-context requests above 200K input tokens cost more on API billing. The Codex target separately manages a 350000 context window and a 250000 total-token compaction threshold, leaving a 100K reserve and a nominal 22K buffer below the higher-pricing boundary, plus native memories and the compact prompt. Official GPT-5.6 capacity is 1.05M, but inputs above 272K are subject to higher provider pricing and may add latency; see Configuration and Privacy, cost & control.


📖 Documentation / Handbook

The full story lives in the agentic-rag Handbook — a single, progressively-ordered read from the mental model through everyday use, configuration, importing, and the engine's architecture and design rationale. A few key chapters:

Start at the handbook index for the one-line "what you'll learn" map of every chapter.


Status

agentic-rag is young but solid — a real engine, openly developed. Repository and rollout state are listed separately below:

  • Storage & search: PostgreSQL + pgvector schema, HNSW ANN blended with EN/DE full-text into one ranked list, the typed edge graph, the three-role destruction-protection matrix.

  • The document write gateway: secret stripping on document inputs and generated document writes, one-transaction chunk + embed + edge-resolve + audit, embeddings that fail open with a retry queue.

  • Session mining: hooks → queue → single-writer worker → configured Codex/Claude CLI → gateway, with a near-duplicate gate and provider-outage circuit breaker.

  • Curation & safety: rag review, refute-as-archive, and admin-only rag purge (removes only already-refuted documents, as rag_admin).

  • Maintenance & backups: pg_dump backups with rotation, the tiny always-exit-0 maintenance job, and the weekly report-only restore-test. macOS auto-schedules via launchd; Linux uses the documented cron/systemd recipes.

  • Claude integration: two user-scope MCP servers (read-write + a read-only server behind a privilege boundary for subagents), idempotent install that preserves foreign hooks.

  • Claude continuity in code: six Claude hooks, the PreCompact stdout compact prompt, the bounded compact_summary handoff, the 10,000-character SessionStart cap, the managed 1M/500K policy, and rag install --check/--restore.

  • 🔒 Claude continuity live rollout: rag install, /hooks review, /autocompact, manual/automatic compaction, and SessionEnd tail capture on the maintainer machine remain open (backlog 0.3).

  • Codex continuity in code: audited checkpoints, bounded capture and restoration, asynchronous enrichment, all six lifecycle handlers, a versioned compact prompt, recoverable installer/check mode, and checkpoint health in rag status.

  • 🔒 Codex continuity live rollout: the pre-install whole-diff/security review and provider-bound pin hardening are complete. The live global install, /hooks trust, manual/automatic compaction, provider-recovery, and SessionEnd smoke tests remain open. See FEATURES.md and blocker-first BACKLOG.md.

  • Quality: a content-free repository with a comprehensive local test suite; exact verification counts belong in rollout evidence, not a static badge.

The clearest gap relative to the field is maturity: it's newly public and self-hosted, without the turnkey hosting or large ecosystem of established RAG stacks.


Acknowledgments

agentic-rag builds on other people's ideas and tools:

  • Andrej Karpathy — the LLM-Wiki idea that shaped the durable-knowledge model this import path speaks to.

  • The llm-wiki format — topic-partitioned Markdown with an optional memory store; rag migrate imports it wholesale, so an existing wiki carries straight over.

  • pgvector and Ollama (bge-m3) — the local vector search and embeddings underneath everything.

  • Anthropic — Claude Code, its hooks, and the MCP integration agentic-rag plugs into.


Project and global applicability now share one explicit scope policy: rag search "your question" --project /absolute/repository.

Contributing

Tests come first (TDD), and docs/ is kept in step with the code. A warn-only doc-reminder hook ships under .githooks/: if a commit touches agentic_rag/ or sql/ without touching docs/, it prints a reminder — it never blocks. Enable it once per clone:

git config core.hooksPath .githooks

Search now returns diverse sources and query-centered spans with chunk citations. See retrieval quality, optional graph expansion and limits.

Measure retrieval and memory quality with the synthetic EN/DE benchmark: rag benchmark run --output /tmp/rag-bench-baseline.

Run the suite with uv run pytest. See the handbook's Contributing chapter for dev setup, the test database, and code layout.

License

MIT. The repository is code-only and content-free — your documents, embeddings, and config stay in your own PostgreSQL database, on your own machine.

Bounded project context

Source-backed advisory profiles and selective EN/DE project recall share one local context service across hooks, CLI and MCP. Exact pins and checkpoint restoration retain priority. See usage and limits.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Persistent memory for Claude Code. Automatically indexes every conversation and provides production-grade hybrid search (BM25 + vectors + reranker) via MCP tools. 100% local, zero config, zero API keys, zero invoice.
    16
    28 npm
    7
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Persistent memory for Claude Code — hybrid search, knowledge graph, session lifecycle.
    17
    30 PyPI
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A local, persistent, semantically-aware knowledge graph for AI coding agents like Claude Code, providing efficient session memory with minimal token cost and zero runtime network calls.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides persistent, searchable memory for Claude Code using local SQLite, semantic embeddings, and full-text search, enabling Claude to recall and retrieve context across sessions and projects without external services.
    8 npm
    4
    MIT