Paparats MCP
by IBazylchuk
README.md
# Paparats MCP
<img src="docs/paparats-kvetka.png" alt="Paparats-kvetka (fern flower)" width="200" align="right">
[](https://www.npmjs.com/package/@paparats/cli)
[](LICENSE)
[](https://www.pulsemcp.com/servers/paparats-mcp)
[](https://lobehub.com/mcp/ibazylchuk-paparats-mcp)
[](https://codespaces.new/IBazylchuk/paparats-mcp) <sub>β try the full stack in your browser, no install ([details](#try-it-in-the-browser-no-install))</sub>
**Paparats-kvetka** β a magical flower from Slavic folklore that blooms on Kupala Night
and grants whoever finds it the power to see hidden things. Likewise, paparats-mcp
helps your agent see the right code across a sea of repositories.
> πΏ Works with **Claude Code Β· Cursor Β· Windsurf Β· Copilot Β· Codex Β· Antigravity** Β· any MCP-compatible agent
**Give your AI coding assistant deep, real understanding of your entire workspace.**
Paparats indexes every repo you care about β semantically, with AST-aware chunking and
a cross-chunk symbol graph β and exposes it through the Model Context Protocol. Search
by meaning, follow `who-uses-what` through real symbol edges, see who last touched a
chunk and which ticket it came from β all without your code ever leaving your machine.
<a href="docs/dashboard.png"><img src="docs/dashboard.png" alt="Paparats operator console β ROI, top queries, cross-project usage, indexer health, embedding latency (synthetic data)" width="100%"></a>
<sub>π The built-in `/ui` operator console β ROI, query quality, cross-project usage, per-user activity, indexer health. <em>Screenshot uses synthetic data (<code>?demo=1</code>) β no real queries, users, or project names.</em></sub>
- β‘ **One install, one config.** `paparats install` β `paparats add ~/code/repo` β done.
- π³ **AST-aware chunking and symbol extraction.** Tree-sitter parses every supported
file once and feeds both chunking and the cross-chunk symbol graph (calls /
called_by / references / referenced_by) β 11 languages including TypeScript, Python,
Go, Rust, Java, Ruby, C, C++, C#.
- π§ **Architectural memory that the agent maintains itself.** A second vector store
per group holds **components, decisions (ADRs) and lessons learned** β your agent
writes them as it works and reads them before answering. Bootstrap on day one with
the `init_arch_memory` MCP prompt (the `/init` of architectural memory). Server-side
similarity gate prevents duplicates, supersedes links replace stale decisions, a
`min_score` threshold gates low-confidence reads, every card carries an
"updated N ago" stamp, and Prometheus metrics tell you whether your memory is
actually being used.
- πΈ **Saves tokens.** Returns only the chunks that matter, with token-savings telemetry
to prove it (per-query, per-user, per-anchor-project).
- π **Production-ready observability.** Prometheus `/metrics`, OpenTelemetry traces
(Tempo, Jaeger, Honeycomb, Datadog, Grafana Cloud, **Elastic APM**), local SQLite
analytics, and a built-in `/ui` operator console that visualises ROI, query quality,
cross-project usage and indexer health in one screen.
- π **100% local by default.** Qdrant + a local embed server (llama.cpp llama-server +
llama-swap) on your machine. No cloud, no API keys, no telemetry leaving the box. Bring
your own Qdrant Cloud / embed server URL if you want.
---
## Table of Contents
- [Why Paparats?](#why-paparats)
- [Quick Start](#quick-start)
- [How the install works](#how-the-install-works)
- [Install variants](#install-variants)
- [Migrating from a v1 install](#migrating-from-a-v1-install)
- [Support agent setup](#support-agent-setup)
- [How It Works](#how-it-works)
- [Key Features](#key-features)
- [Use Cases](#use-cases)
- [Architectural memory](#architectural-memory-agent-maintained-adrs-components-lessons)
- [Configuration](#configuration)
- [MCP Tools Reference](#mcp-tools-reference)
- [Connecting MCP](#connecting-mcp)
- [CLI Commands](#cli-commands)
- [Monitoring](#monitoring)
- [Analytics & Observability](#analytics--observability)
- [Architecture](#architecture)
- [Embedding Model Setup](#embedding-model-setup)
- [Comparison with Alternatives](#comparison-with-alternatives)
- [Token Savings Metrics](#token-savings-metrics)
- [Contributing](#contributing)
- [Links](#links)
---
## Why Paparats?
AI coding assistants are smart, but they can only see files you open. They don't know your codebase structure, where the authentication logic lives, or how services connect. **Paparats fixes that.**
### What you get
- **Semantic code search** β ask "where is the rate limiting logic?" and get exact code ranked by meaning, not grep matches
- **Real-time sync** β edit a file, and 2 seconds later it's re-indexed. No manual re-runs
- **Cross-chunk symbol graph** β `find_usages` walks AST-derived edges (calls, called_by, references, referenced_by) so the agent can trace dependencies without re-grepping
- **Token savings** β return only relevant chunks instead of full files to reduce context size
- **Multi-project workspaces** β search across backend, frontend, infra repos in one query
- **100% local & private** β Qdrant vector database + local llama-server embeddings. Nothing leaves your laptop
- **AST-aware chunking** β code split by AST nodes (functions/classes) via tree-sitter, not arbitrary character counts (TypeScript, JavaScript, TSX, Python, Go, Rust, Java, Ruby, C, C++, C#; regex fallback for Terraform)
- **Rich metadata** β each chunk knows its symbol name (from tree-sitter AST), service, domain context, and tags from directory structure
- **Git history per chunk** β see who last modified a chunk, when, and which tickets (Jira, GitHub) are linked to it
- **Architectural memory** β a living knowledge base of components, decisions (ADRs) and lessons learned, written by the agent as it learns, deduplicated server-side by vector similarity, and consulted on every support query so the agent stays consistent across sessions
### Who benefits
| Use Case | How Paparats Helps |
| --------------------------- | ------------------------------------------------------------------------------------------------------ |
| **Solo developers** | Quickly navigate unfamiliar codebases, find examples of patterns, reduce context-switching |
| **Multi-repo teams** | Cross-project search (backend + frontend + infra), consistent patterns, faster onboarding |
| **AI agents** | Foundation for product support bots, QA automation, dev assistants β any agent that needs code context |
| **Legacy modernization** | Find all usages of deprecated APIs, identify migration patterns, discover hidden dependencies |
| **Contractors/consultants** | Accelerate ramp-up on client codebases, reduce "where is X?" questions |
---
## Quick Start
### Try it in the browser (no install)
[](https://codespaces.new/IBazylchuk/paparats-mcp)
Spin up a full Qdrant + embed server + paparats stack in a Codespace.
A small slice of the repo (`packages/shared/src`) is auto-indexed on first start so
you can run
```bash
paparats search -g demo 'gitignore filter'
```
within a few minutes. Codespace forwards port 9876 for MCP β point Cursor/Claude Code
at it via the URL VS Code shows in the Ports panel.
> Note: Codespaces is for demo only. With CPU embedding the full repo
> would take 15+ minutes and can hit batch timeouts on large files. For real
> workloads run locally β or set `OPENAI_API_KEY` (or `VOYAGE_API_KEY`) as a
> Codespaces user secret and indexing drops to a couple of seconds; see the
> Embedding providers section below.
### Run locally
You need **Docker** and **Docker Compose v2**. On macOS, also install the **embed server
natively** β running it inside Docker on macOS is significantly slower because the Docker
VM cannot use Apple Silicon GPU (Metal) acceleration.
```bash
# 1. Install the CLI.
npm install -g @paparats/cli
# 2. macOS only β install the native embed server (Linux uses the Docker embed
# image by default). Metal-accelerated.
brew install llama.cpp mostlygeek/llama-swap/llama-swap
# 3. One-time bootstrap. Generates ~/.paparats/{docker-compose.yml,projects.yml},
# starts the stack, downloads the embedding model, wires Cursor/Claude Code MCP.
paparats install
# 4. Add the projects you want indexed. Local paths bind-mount read-only into the
# indexer; git URLs and owner/repo shorthand get cloned.
paparats add ~/code/my-project
paparats add git@github.com:acme/billing.git
paparats add acme/widgets
# 5. Watch it work.
paparats list
```
That's it. Your IDE is already wired (`~/.cursor/mcp.json`, `~/.claude/mcp.json`) to
`http://localhost:9876/mcp`. Open Cursor or Claude Code and ask:
> "Search this workspace for the auth middleware and show me everything that calls it."
### Existing v1 user?
Just run `paparats install` again. The installer detects the legacy per-project
compose, asks once before swapping it for the new global setup, and **preserves your
indexed data** (Qdrant collections, SQLite metadata, embedding cache). Your in-repo
`.paparats.yml` files keep working as per-project overrides.
---
## How the install works
`paparats install` is the only setup command. It creates a single global home at
`~/.paparats/`, brings up a Docker stack, and wires your MCP clients. Re-run it any time
to reconfigure β it diffs the existing compose and asks before overwriting hand edits.
```
~/.paparats/
βββ docker-compose.yml generated; hand-editable; install asks before overwriting
βββ projects.yml project list (CLI rewrites it; comments survive your manual edits)
βββ install.json install flags persisted so add/remove can regenerate compose
βββ .env secrets β Qdrant API key, GitHub token; chmod 600
βββ models/ bge-code-v1 + qwen3-embedding-0.6b GGUF (native embed mode)
βββ data/ Docker volumes (mounted by name from compose)
βββ qdrant/ vector index
βββ sqlite/ metadata.db, embeddings.db, analytics.db
βββ repos/ cloned remote projects
```
Inside the Docker stack:
| Service | Image | Port | Role |
| ------------------ | ------------------------------ | ----- | -------------------------------------------------------- |
| `paparats-mcp` | `ibaz/paparats-server:latest` | 9876 | MCP HTTP/SSE endpoints, search, metadata API |
| `paparats-indexer` | `ibaz/paparats-indexer:latest` | 9877 | Cron + on-demand indexing, hot-reload of project list |
| `qdrant` | `qdrant/qdrant:latest` | 6333 | Vector DB (skipped when you pass `--qdrant-url`) |
| `embed` | `ibaz/paparats-embed:latest` | 11434 | Embed server β llama-server + llama-swap, `bge-code-v1` + `qwen3-embedding-0.6b` pre-baked (Linux default; macOS uses native embed server). llama-swap listens on 8080 inside the container |
The indexer hot-reloads `projects.yml`. Edits that **change project metadata
only** (group, language, indexing tweaks) reindex in place. Edits that **add or remove
local-path projects** require a stack restart so Docker picks up the new bind-mount β
the CLI does this for you on `paparats add` and `paparats remove`.
---
## Install variants
### Default (recommended)
```bash
paparats install
```
On macOS prefers the native embed server and dockerized Qdrant. On Linux defaults to
Docker for both.
### Bring your own Qdrant
```bash
paparats install --qdrant-url https://qdrant.example.com
# Asks for an API key after; stored in ~/.paparats/.env as QDRANT_API_KEY.
```
When `--qdrant-url` is set the Qdrant container is omitted from the stack entirely.
### Bring your own embed server
```bash
paparats install --embed-url http://10.0.0.5:11434
```
Skips both the native and Docker embed server.
> **The remote endpoint must serve the `bge-code-v1` model (and `qwen3-embedding-0.6b` for the
> arch-memory layer) over an OpenAI-style `/v1/embeddings` API.** The installer will not
> touch a remote instance. The simplest way is to run the pre-baked image on that host β
> no model registration needed, llama-swap loads GGUF by name on first request:
>
> ```bash
> docker run -d -p 11434:8080 ibaz/paparats-embed:latest
> ```
>
> Then `paparats install --embed-url http://that-host:11434` and Paparats will use it.
### Force Docker embed server on macOS
```bash
paparats install --embed-mode docker
```
Slower on Apple Silicon (no Metal GPU), but useful for parity testing or laptops without
brew.
### Scripted / CI
```bash
paparats install --non-interactive --force
```
Fails on any prompt; `--force` answers Y to compose-overwrite and migration prompts.
---
## Migrating from a v1 install
When `paparats install` finds a legacy `~/.paparats/docker-compose.yml` (the one from the
old per-project flow with no `paparats-indexer` service), it prints a one-screen
migration notice and asks before tearing the legacy stack down.
**What survives:** Qdrant collections, SQLite metadata, indexer repos, and any
`.paparats.yml` files inside your repos (those still take precedence over
`projects.yml` overrides).
**What's deleted:** the legacy `docker-compose.yml` and `.env`. They are regenerated on
the spot under the new schema.
**No re-indexing needed** β the data volumes are referenced by the same names in the new
compose. Add your projects with `paparats add` and they re-appear in `paparats list` with
their existing chunks.
If your install predates the `paparats-indexer.yml` β `projects.yml` rename, the
installer migrates the file in place on first run and prints a one-line notice.
The indexer also reads the legacy name as a fallback, so nothing breaks if you
roll out the indexer before re-running `paparats install`.
Pass `--force` to skip the migration prompt in scripts.
---
## Support agent setup
For bots and support teams that consume an existing Paparats server β no Docker, no
embed server needed on this side.
```bash
# Connect to a running server (default: localhost:9876)
paparats install --mode support
# Connect to a remote server
paparats install --mode support --server http://prod-server:9876
```
The installer verifies the server is reachable, then wires Cursor MCP
(`~/.cursor/mcp.json`) and Claude Code MCP (`~/.claude/mcp.json`) to the support
endpoint. Tools available on `/support/mcp`: `search_code`, `get_chunk`, `find_usages`,
`list_projects`, `health_check`, `get_chunk_meta`, `search_changes`, `explain_feature`,
`recent_changes`, `impact_analysis`, **`arch_context`**, **`arch_record_component`**,
**`arch_record_decision`**, **`arch_record_lesson`** (architectural memory β see
[Key Features](#architectural-memory-agent-maintained-adrs-components-lessons)), plus
the analytics tools described in **Observability** below.
---
## How It Works
```
Your projects Paparats AI assistant
(Claude Code / Cursor)
backend/ ββββββββββββββββββββββββ
.paparats.yml βββββββββΊβ Indexer β
frontend/ β - chunks code β ββββββββββββββββ
.paparats.yml βββββββββΊβ - embeds via llama βββββββββββΊβ MCP search β
infra/ β - stores in Qdrant β β tool call β
.paparats.yml βββββββββΊβ - watches changes β ββββββββββββββββ
ββββββββββββββββββββββββ
```
### Indexing Pipeline
During each indexer cycle (cron-driven, on-demand via `paparats add`, or triggered by
the indexer's chokidar file watcher), every file in scope flows through this pipeline:
```
Source file
β
βΌ
βββββββββββββββββββ
β 1. File discoveryβ Collect files from indexing.paths, apply
β & filtering β gitignore + exclude patterns, skip binary
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 2. Content hash β SHA-256 of file content β compare with
β check β existing Qdrant chunks β skip unchanged
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 3. AST parsing β tree-sitter parses the file once (WASM)
β (single pass) β β reused for chunking AND symbol extraction
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 4. Chunking β AST nodes β chunks at function/class
β β boundaries. Regex fallback for unsupported
β β languages (brace/indent/block strategies)
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 5. Symbol β AST queries extract module-level defines
β extraction β (function/class/variable names) and uses
β β (calls, references) per chunk. 11 languages
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 6. Metadata β Service name, bounded_context, tags from
β enrichment β config + auto-detected directory tags
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 7. Embedding β Jina Code Embeddings 1.5B via llama-server
β β SQLite cache (content-hash key) β skip
β β already-embedded content
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 8. Qdrant upsert β Vectors + payload (content, file, lines,
β β symbols, metadata) β batched upsert
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β 9. Git history β git log per file β diff hunks β map
β (post-index) β commits to chunks by line overlap β
β β extract ticket refs β store in SQLite
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β10. Symbol graph β Cross-chunk edges: calls β called_by,
β (post-index) β references β referenced_by β SQLite
βββββββββββββββββββ
```
Step 5's symbol extractor only emits **module-level** definitions β locals declared
inside function bodies, callback args, and hook closures stay out of the graph because
they're not addressable from another chunk anyway.
### Search Flow
AI assistant queries via MCP β server detects query type (nl2code / code2code / techqa) β expands query (abbreviations, case variants, plurals) β all variants searched in parallel against Qdrant β results merged by max score β only relevant chunks returned with confidence scores and symbol info.
### Watching
The indexer container watches the projects mounted into it via chokidar with debouncing
(1s default). On change, only the affected file re-enters the pipeline. Unchanged content
is never re-embedded thanks to the content-hash cache. The indexer also hot-reloads
`~/.paparats/projects.yml` itself: metadata-only edits reindex in place;
add/remove of local-path projects triggers a stack restart through the CLI.
---
## Key Features
### Better Search Quality
**Task-specific embeddings** β Jina Code Embeddings supports 3 query types (nl2code, code2code, techqa) with different prefixes for better relevance:
- `"find authentication middleware"` β `nl2code` prefix (natural language β code)
- `"function validateUser(req, res)"` β `code2code` prefix (code β similar code)
- `"how does OAuth work in this app?"` β `techqa` prefix (technical questions)
**Query expansion** β every search generates 2-3 variations server-side:
- Abbreviations: `auth` β `authentication`, `db` β `database`
- Case variants: `userAuth` β `user_auth` β `UserAuth`
- Plurals: `users` β `user`, `dependencies` β `dependency`
- Filler removal: `"how does auth work"` β `"auth"`
All variants searched in parallel, results merged by max score.
**Confidence scores** β each result includes a percentage score (β₯60% high, 40β60% partial, <40% low) to guide AI next steps.
### Performance
**Embedding cache** β SQLite cache with content-hash keys + Float32 vectors. Unchanged code never re-embedded. LRU cleanup at 100k entries.
**AST-aware chunking** β tree-sitter AST nodes define natural chunk boundaries for 11 languages. Falls back to regex strategies (block-based for Ruby, brace-based for JS/TS, indent-based for Python, fixed-size) for unsupported languages.
**Real-time watching** β the indexer's `chokidar` watcher reindexes a project on file
changes with debouncing (1s default). For local-path projects bind-mounted into the
indexer, edits on your host show up in MCP queries within seconds.
### Cross-chunk symbol graph
The post-index pass walks every chunk's `defines_symbols` / `uses_symbols` lists and
materializes edges into SQLite β `calls`, `called_by`, `references`, `referenced_by`.
`find_usages` returns those edges grouped by direction so the agent can traverse the
graph without re-searching. Because extraction is AST-driven, function locals don't
pollute the graph.
### Architectural memory (agent-maintained ADRs, components, lessons)
Code search tells the agent **what the code does**. Architectural memory tells it
**why** β and the agent maintains that knowledge itself, across sessions, without you
authoring a single doc.
**Three card kinds, structured by design:**
| Kind | Captures | Fields |
| :------------ | :---------------------------------------------------------------------------------------- | :---------------------------------------------------------------- |
| **Component** | A unit with a clear responsibility (service, module, subsystem) | `name`, `summary` with `Does / Owns / Does not / Touched when` |
| **Decision** | An architectural choice (ADR-style) | `title`, `context`, `decision`, `alternatives_rejected`, `consequences` |
| **Lesson** | A rule learned from an incident, a code review, a bug, or a user correction (Reflexion-style) | `rule`, `why`, `when` |
The agent reads them via **`arch_context`** before any architectural answer, and
writes them via **`arch_record_component`**, **`arch_record_decision`**, and
**`arch_record_lesson`** whenever it discovers something new or learns from a
correction. Each card carries an `updated N ago` stamp in the read tool so the agent
can spot stale memory and verify against current code.
**Server-side similarity gate** (cosine over [Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B)
text embeddings, 1024d):
- `β₯ 0.85` is a **duplicate** β decisions are refused (the agent must reconcile or
supersede); lessons bump `updatedAt` (Reflexion-style "rule confirmed").
- `0.70 β 0.85` is **similar** β surfaced to the agent so it can refine the wording
or chain a supersede.
- `< 0.70` is **new** β accepted as a fresh card.
`supersedes` links bypass the gate and mark the prior decision as `status=superseded`
so it disappears from default search but remains in history.
**Why this matters:**
- π§ **Cross-session continuity** β what the agent learned last week, today's agent
still knows.
- π **ADRs without the ceremony** β no markdown files to maintain, no review process,
no doc drift. The agent writes when it learns.
- π **Reflexion built in** β corrections become lessons; repeated mistakes get caught.
- π¦ **No memory rot** β similarity gate kills duplicates, supersedes link replaces
stale decisions, age stamps trigger verification against code.
Lives in a separate Qdrant collection per group (`paparats_<group>_arch`). Reading
(`arch_context`) is available on **both** endpoints β coding agents need to know
about prior decisions before refactoring. Writing (`arch_record_*`) is **support-only**:
recording belongs to the architectural-review workflow, not to every line edit.
**`arch_context` accepts a `min_score` parameter** (default `0.45`, cosine over
[Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B)). Lower it to broaden recall on a sparse
arch memory; raise it to demand only high-confidence cards. The tool also emits an
explicit low-confidence hint when the question matched nothing above the threshold,
so the agent knows to either rephrase or lower `min_score` instead of inventing
context.
**Initialise the arch layer on day one.** Two purpose-built MCP **workflow prompts**
make the boring scaffolding work disappear:
- **`init_arch_memory`** β the `/init` of architectural memory. Walks the repo,
identifies 8-20 components by domain boundary, writes them, and captures any
obvious decisions inferable from comments or README. Run it once per group, right
after installing.
- **`audit_architecture`** β sweeps the memory of one group, flags cards older than
90 days, verifies anchors against the live code, and surfaces a punch list of
updates / supersedes for your approval.
- **`record_lesson_from_correction`** β converts a user correction into a structured
lesson card (rule / why / when) without overrecording typos.
**MCP resources for live introspection:**
- **`arch://schema`** β the full card-schema reference (fields, similarity-gate
thresholds, write semantics). Cite it from the agent when explaining the model.
- **`arch://stats/{group}`** β live counts (total / by kind / by status) and the
oldest/newest `updatedAt` per group. The same numbers are also pushed to
Prometheus.
**Observability built in.** When `PAPARATS_METRICS=true`, every read/write hits a
counter and the cosine score of returned cards lands in a histogram:
- `paparats_arch_context_calls_total{group}` β calls per group
- `paparats_arch_write_total{kind, status}` β writes by card kind and gate outcome
- `paparats_arch_search_score` β histogram of cosine scores in `arch_context`
results (post `min_score`)
- `paparats_arch_collection_size{group, kind, status}` β gauge updated whenever
`arch://stats/{group}` is read
These let you spot a memory that's not being written to, a similarity gate that's
too aggressive, or a sparse group where every query returns low-confidence hits.
---
## Use Cases
### For Developers (Coding)
Connect via the **coding endpoint** (`/mcp`):
| Use Case | How |
| ---------------------------- | ---------------------------------------------------------------------- |
| **Navigate unfamiliar code** | `search_code "authentication middleware"` β exact locations |
| **Find similar patterns** | `search_code "retry with exponential backoff"` β examples |
| **Trace dependencies** | `find_usages {chunk_id, direction: "incoming"}` β callers via the graph |
| **Explore context** | `get_chunk <chunk_id> --radius_lines 50` β expand around |
| **Manage projects** | `list_projects` and `delete_project` for index hygiene |
### For Support Teams
Connect via the **support endpoint** (`/support/mcp`):
| Use Case | How |
| ------------------------------ | ---------------------------------------------------------------------- |
| **Explain a feature** | `explain_feature "rate limiting"` β code locations + changes |
| **Recent changes** | `recent_changes "auth" --since 2024-01-01` β timeline with tickets |
| **Trace usages** | `find_usages {chunk_id}` β who calls/references this chunk |
| **Change history** | `get_chunk_meta <chunk_id>` β authors, dates, linked tickets |
| **Blast radius** | `impact_analysis <chunk_id>` β cross-chunk + cross-project impact |
| **Architectural Q&A** | `arch_context "why X"` β components / decisions / lessons (with age) |
| **Capture decisions & lessons**| `arch_record_decision` / `arch_record_lesson` β agent writes as it learns, server-side dedup |
**Support chatbot example:**
```
User: "How do I configure rate limiting?"
Bot workflow (via /support/mcp):
1. explain_feature("rate limiting", group="my-app")
β returns code locations + recent changes + related modules
2. get_chunk_meta(<chunk_id>)
β returns who last modified it, when, linked tickets
3. Bot synthesizes response in plain language with ticket references
```
---
## Configuration
Paparats uses two config files. Both are optional β defaults work for the common case.
### `~/.paparats/projects.yml` β global project list
Lives outside your repos. Edited by `paparats add` / `paparats remove` or by hand via
`paparats edit projects`. Every entry has either `path:` (local bind-mount) or `url:`
(remote git, cloned by the indexer), never both.
```yaml
defaults:
cron: '0 */6 * * *' # global indexer schedule
group: workspace # default group when an entry doesn't specify one
repos:
- path: /Users/alice/code/billing # local bind-mount
group: dev
language: typescript
- url: org/widgets # remote git, cloned by the indexer
group: prod
language: ruby
- url: git@github.com:acme/billing.git
name: billing # override the auto-derived name
group: prod
```
The indexer hot-reloads this file. Adding/removing **local-path** entries causes the CLI
to restart the stack so Docker picks up the new bind-mount; metadata-only edits reindex
in place.
### `.paparats.yml` in your repo β per-project overrides
Drop one at the project root to override anything from the global file.
```yaml
group: my-app
language: typescript
# Indexing tuning (all optional)
indexing:
paths: [src, packages] # restrict to these subdirectories
exclude: [node_modules, dist, '**/*.test.ts']
exclude_extra: ['**/__fixtures__/**'] # added on top of language defaults
chunkSize: 1500 # characters per chunk (default: 1200)
overlap: 100 # chunk overlap (default: 100)
concurrency: 4 # parallel embedding requests
batchSize: 8 # embeddings per llama-server call
# Metadata
metadata:
service: billing
bounded_context: payments
tags: [backend, critical]
directory_tags:
src/api: [public-api]
src/internal: [internal]
# Git history per chunk (Jira / GitHub ticket extraction included)
git:
enabled: true
maxCommitsPerFile: 50
ticketPatterns:
- '\b([A-Z]+-\d+)\b' # Jira-style PROJ-123
- '#(\d+)' # GitHub-style #123
```
In-repo `.paparats.yml` always wins over `projects.yml`. The CLI never
overwrites it.
### Groups
A **group** is a Qdrant collection (`paparats_<group>`). Multiple projects can share a
group to enable cross-project search; each project lives as a `project:` field in the
chunk payload. By default `group` defaults to the project name (one project, one
collection). Set the same `group:` on multiple entries to consolidate them.
### Git history per chunk
When `metadata.git.enabled: true` (default), the indexer maps each chunk to the commits
that touched its line range using diff-hunk overlap. Tickets are extracted from commit
messages using `metadata.git.ticketPatterns` (built-in: Jira `PROJ-123`, GitHub `#42`,
cross-repo `org/repo#99`). Surfaced through MCP tools `get_chunk_meta`, `search_changes`,
`recent_changes`, `explain_feature`. Non-fatal: non-git projects index normally.
---
## MCP Tools Reference
Paparats serves the Model Context Protocol on **two separate endpoints**, each with its
own tool set and system instructions.
### Coding endpoint (`/mcp`)
For developers using Claude Code, Cursor, etc. Focus: search code, read chunks, follow
the cross-chunk symbol graph, manage projects.
| Tool | Description |
| :--------------- | :--------------------------------------------------------------------------------------------------------------------------- |
| `search_code` | Semantic search across indexed projects. Returns chunks with symbol info and confidence scores. |
| `get_chunk` | Retrieve a chunk by ID with optional surrounding context. |
| `find_usages` | Walk the symbol graph from a `chunk_id` β `incoming` (callers/references in), `outgoing` (calls/references out), or `both`. |
| `list_projects` | List indexed projects with chunk counts and detected languages. |
| `delete_project` | Wipe Qdrant chunks + SQLite metadata for a project (CLI's `paparats remove` calls it). |
| `health_check` | Indexing status, chunks per group, running jobs. |
| `arch_context` | Read-only architectural memory. Returns components, decisions, and lessons relevant to the query with `updated N ago` stamps and a `min_score` cutoff. |
### Support endpoint (`/support/mcp`)
For support teams and bots without direct code access. Focus: feature explanations,
change history, cost reporting β all in plain language.
| Tool | Description |
| :--------------------- | :------------------------------------------------------------------------------------- |
| `search_code` | Same as coding endpoint. |
| `get_chunk` | Same. |
| `find_usages` | Same. |
| `list_projects` | Same. |
| `health_check` | Same. |
| `get_chunk_meta` | Git history and ticket references for a chunk β commits, authors, dates. No code. |
| `search_changes` | Semantic search filtered by last-commit date. Each result shows when it last changed. |
| `explain_feature` | Comprehensive feature analysis: locations + recent changes for a question. |
| `recent_changes` | Timeline grouped by date with commits, tickets, affected files. `since` filter. |
| `impact_analysis` | Cross-chunk impact for a `chunk_id` β symbol graph traversal + cross-project blast radius. |
| `arch_context` | Read the architectural memory for a group β top-matching components, decisions and lessons, each stamped with "updated N ago" and a cosine score. Accepts a `min_score` parameter (default 0.45) to gate low-confidence hits. Call before any architectural answer. **Also available on `/mcp`.** |
| `arch_record_component` | Record a component with `Does / Owns / Does not / Touched when` fields. Idempotent by `name`. |
| `arch_record_decision` | Record an ADR-style decision (`context / decision / alternatives_rejected / consequences`). Server-side similarity gate refuses duplicates and surfaces near-matches; `supersedes` links replace prior decisions. |
| `arch_record_lesson` | Record a lesson as `rule / why / when`. Duplicates bump `updatedAt` (Reflexion confirmation) instead of overwriting. |
| `token_savings_report` | Aggregate token-savings stats (naive baseline vs search-only vs actually consumed). |
| `top_queries` | Most frequent queries by user/session/project anchor. |
| `slowest_searches` | Top-N slowest searches with timing + chunk counts. |
| `cross_project_share` | Off-anchor result share per user β indicator of search noise. |
| `retry_rate` | Tool-call retry rate per user β indicator of unhelpful results. |
| `failed_chunks` | AST parse failures, regex fallbacks, zero-chunk files, binary skips. |
### Typical workflows
**Drill-down (coding agent):**
```
1. search_code "authentication middleware" β relevant chunks with symbols
2. get_chunk <chunk_id> --radius_lines 50 β expand context around a hit
3. find_usages {chunk_id, direction: "incoming"} β who calls / references this chunk
```
**Single-call (support agent):**
```
1. explain_feature "How does authentication work?" β locations + recent changes
2. recent_changes "auth" --since 2024-01-01 β timeline with tickets
3. token_savings_report β cost report for the last 7 days
```
**Architectural memory (support agent):**
```
1. arch_context "why do we use qwen3-embedding-0.6b for the arch layer?"
β top components / decisions / lessons,
each with an "updated N ago" stamp
2. arch_record_decision { title, context, decision, alternatives_rejected, consequences }
β status=created | duplicate | similar
(gate refuses duplicates server-side)
3. arch_record_lesson { rule, why, when } β status=created | updated (Reflexion bump)
```
---
## Connecting MCP
`paparats install` already wires Cursor (`~/.cursor/mcp.json`) and Claude Code
(`~/.claude/mcp.json`) to `http://localhost:9876/mcp`. The sections below are for
manual setup or for adding the **support** endpoint alongside the default coding one.
### Cursor
Create or edit `~/.cursor/mcp.json` (global) or `.cursor/mcp.json` (project):
```json
{
"mcpServers": {
"paparats": {
"type": "http",
"url": "http://localhost:9876/mcp"
}
}
}
```
For support use case (feature explanations, change history, impact analysis):
```json
{
"mcpServers": {
"paparats-support": {
"type": "http",
"url": "http://localhost:9876/support/mcp"
}
}
}
```
Restart Cursor after changing config.
### Claude Code
```bash
# Coding endpoint (default)
claude mcp add --transport http paparats http://localhost:9876/mcp
# Support endpoint (for support bots/agents)
claude mcp add --transport http paparats-support http://localhost:9876/support/mcp
```
Or add to `.mcp.json` in project root:
```json
{
"mcpServers": {
"paparats": {
"type": "http",
"url": "http://localhost:9876/mcp"
}
}
}
```
### Verify
- `paparats status` β check stack is up
- **Coding endpoint** (`/mcp`): `search_code`, `get_chunk`, `find_usages`,
`list_projects`, `delete_project`, `health_check`
- **Support endpoint** (`/support/mcp`): `search_code`, `get_chunk`, `find_usages`,
`health_check`, `list_projects`, plus the support-specific tools `get_chunk_meta`,
`search_changes`, `explain_feature`, `recent_changes`, `impact_analysis`, and the
analytics tools listed in **Observability** (`token_savings_report`, `top_queries`,
`slowest_searches`, `cross_project_share`, `retry_rate`, `failed_chunks`)
- Ask the AI: _"Search this workspace for the auth middleware"_
---
## CLI Commands
```text
paparats install [flags] Bootstrap or reconfigure the global stack.
paparats add <path-or-repo> [flags] Add a project (local path or git URL/shorthand).
paparats list [--json] [--group g] Show indexed projects with status from the indexer.
paparats remove <name> [--yes] Remove a project β deletes Qdrant + SQLite data.
paparats start [--logs] Start the Docker stack (with `--logs` follows them).
paparats stop Stop the stack (preserves data volumes).
paparats restart Recreate containers (applies new compose changes).
paparats edit compose|projects Open the file in $EDITOR; on save, validate +
regenerate compose + restart + reindex (projects).
paparats search <query> [flags] Semantic search from the terminal.
paparats status Stack health: Docker, embed server, server, indexer.
paparats groups [--json] List groups and their projects.
paparats doctor Diagnostic checks (Docker, embed server, ports, configs).
paparats update Update CLI from npm + pull latest Docker images.
```
The legacy per-project commands (`paparats init`, `paparats index`, `paparats watch`) are
gone β adding a project is now `paparats add`, indexing is automatic in the indexer
container, watching is the `chokidar` watcher inside the indexer.
### Common flags
**`paparats install`**
- `--embed-mode <native|docker>` β force embed server mode (default: native on macOS, docker on Linux)
- `--embed-url <url>` β external embed server; skips both native and docker embed server
- `--qdrant-url <url>` β external Qdrant; skips the Qdrant container
- `--qdrant-api-key <key>` β for authenticated Qdrant (e.g. Qdrant Cloud); written to `~/.paparats/.env`
- `--mode support` β wire MCP clients only, no Docker stack
- `--server <url>` β server URL for support mode (default: `http://localhost:9876`)
- `--force` β skip overwrite/migration prompts
- `--non-interactive` β fail on any prompt instead of asking
- `-v, --verbose` β stream Docker output
**`paparats add <path-or-repo>`**
- `--name <name>` β override the auto-derived project name (basename of path / repo)
- `--group <group>` β override group (default: project name)
- `--language <lang>` β override language (default: auto-detect)
- `--no-restart` β skip the Docker restart for local-path adds (useful in scripts)
- `--no-reindex` β skip the per-project reindex trigger
- `--force` β drop the project's existing chunks before reindexing (destructive, use after schema/config changes)
**`paparats remove <name>`**
- `--yes` β skip the confirmation prompt
**`paparats search <query>`**
- `-n, --limit <n>` β max results (default: 5)
- `-p, --project <name>` β filter by project
- `-g, --group <name>` β restrict to a group
- `--json` β machine-readable output
### Environment overrides
| Var | Default | What |
| ---------------------- | ----------------------- | ------------------------------------------ |
| `PAPARATS_SERVER_URL` | `http://localhost:9876` | MCP server base URL (used by CLI commands) |
| `PAPARATS_INDEXER_URL` | `http://localhost:9877` | Indexer base URL (`add`, `list`, `edit`) |
---
## Monitoring
Paparats exposes Prometheus metrics for operational visibility. Opt in by setting `PAPARATS_METRICS=true` in the server's environment:
```yaml
# In ~/.paparats/docker-compose.yml, under paparats service:
environment:
PAPARATS_METRICS: 'true'
```
### Metrics endpoint
```bash
curl http://localhost:9876/metrics
```
### Key metrics
| Metric | Type | Description |
| ----------------------------------- | --------- | ----------------------------------- |
| `paparats_search_total` | Counter | Search requests by group and method |
| `paparats_search_duration_seconds` | Histogram | Search latency |
| `paparats_index_files_total` | Counter | Files indexed |
| `paparats_index_chunks_total` | Counter | Chunks indexed |
| `paparats_query_cache_hit_rate` | Gauge | Query result cache hit rate |
| `paparats_embedding_cache_hit_rate` | Gauge | Embedding cache hit rate |
| `paparats_watcher_events_total` | Counter | File watcher events |
### Prometheus scrape config
```yaml
scrape_configs:
- job_name: paparats
scrape_interval: 15s
static_configs:
- targets: ['localhost:9876']
```
### Query cache
Search results are cached in-memory (LRU, default 1000 entries, 5-minute TTL). The cache is automatically invalidated when files change. Configure via environment variables:
- `QUERY_CACHE_MAX_ENTRIES` β max cached queries (default: 1000)
- `QUERY_CACHE_TTL_MS` β TTL in milliseconds (default: 300000)
Cache stats are included in `GET /api/stats` under the `queryCache` field.
---
## Analytics & Observability
Paparats ships with three observability layers that work together:
1. **Prometheus** (`PAPARATS_METRICS=true`, see above) β scrape `/metrics`.
2. **Local SQLite analytics store** at `~/.paparats/analytics.db` (default ON) β raw search/tool/indexing events. Six MCP tools query it directly: `token_savings_report`, `top_queries`, `cross_project_share`, `retry_rate`, `slowest_searches`, `failed_chunks`.
3. **OpenTelemetry** (`PAPARATS_OTEL_ENABLED=true` + `OTEL_EXPORTER_OTLP_ENDPOINT`) β spans for every search, MCP tool call, embedding, indexing run, chunking error. Works with Tempo, Jaeger, Honeycomb, Datadog, Grafana Cloud β anything that speaks OTLP/HTTP.
### Operator console (`/ui`)
Open `http://localhost:9876/ui` for a single-screen dashboard ([see screenshot at top of README](#paparats-mcp)) that visualises the analytics store above: ROI, top / slowest queries, cross-project usage, per-user activity, indexer status, embedding p95/p99, and recent failures. Polls every 5 s, no extra services to run.
- Protect it (optional): `PAPARATS_UI_BASIC_AUTH=user:pass` β applies to `/ui` and `/api/analytics` only; `/mcp` and `/api/search` stay open so agents keep working.
- Show the screenshot view to anyone without touching real data: `PAPARATS_UI_DEMO=true` (or append `?demo=1` to the URL once).
### Pre-built Grafana dashboard
The built-in `/ui` covers the current snapshot. For history (latency p99 over weeks, GC trends, CPU under indexing bursts) wire `/metrics` to Prometheus and import [`docs/grafana/paparats.json`](docs/grafana/paparats.json) β 15 panels across four rows: **Traffic & latency**, **Embeddings**, **Indexing**, **Process health**.
```bash
# 1. Enable Prometheus surface on the server.
PAPARATS_METRICS=true paparats up # or set in your docker-compose.yml
# 2. Point your Prometheus at http://<server>:9876/metrics.
# 3. In Grafana: Dashboards β Import β upload docs/grafana/paparats.json
# β pick your Prometheus datasource β Import.
```
The dashboard uses a `${DS_PROMETHEUS}` variable, so it works with any Prometheus instance (local, Grafana Cloud, Mimir, VictoriaMetrics).
### Sending traces to Elastic APM (or any OTLP backend)
Elastic APM Server accepts OpenTelemetry natively since 7.14 β no agent install, no SDK injection. Set four env vars on the paparats container and restart:
```bash
PAPARATS_OTEL_ENABLED=true
OTEL_EXPORTER_OTLP_ENDPOINT=https://your-apm-server:8200
OTEL_EXPORTER_OTLP_HEADERS=Authorization=Bearer <apm-secret-token>
OTEL_SERVICE_NAME=paparats-mcp
```
Within a minute a new service `paparats-mcp` appears in APM β Services. The same env vars work for Tempo, Jaeger, Honeycomb, Datadog, Grafana Cloud Traces β change the endpoint and auth header.
**What gets recorded** β one span per event, with paparats-specific attributes for filtering and grouping:
| Span name | Key attributes | When |
| --------------------------------- | ------------------------------------------------------------------------------------ | ------------------------------------- |
| `paparats.search` | `tool`, `group`, `query.hash`, `query.length`, `search.duration_ms`, `result_count`, `cache_hit` | every `search_code` / `find_usages` |
| `paparats.get_chunk` | `chunk_id`, `fetch.radius_lines`, `fetch.duration_ms`, `fetch.found` | every `get_chunk` call |
| `paparats.mcp.tool` | `tool`, `tool.duration_ms`, `tool.ok` | every MCP tool invocation |
| `paparats.embedding` | `kind`, `batch_size`, `cache_hits`, `cache_miss`, `duration_ms`, `timeout` | every embedding request |
| `paparats.indexing.run` | `group`, `project`, `trigger`, `status`, `files_total`, `chunks_total`, `errors_total` | every indexer cycle |
| `paparats.indexing.chunking_error`| `group`, `project`, `file`, `language`, `error_class` | per-file chunking failure |
Every span also carries `paparats.user`, `paparats.session`, `paparats.client`, `paparats.request_id`, and (when present) `paparats.anchor_project` from the identity headers above β so you can filter APM by user or correlate spans across a single MCP session.
**What this is good for in Elastic APM:**
- **Errors view** β chunking and embedding failures with stacktrace + file/language/error_class context, aggregated by error class.
- **Transactions** β `paparats.search` becomes a transaction type. Sort by p95/p99/error rate to find the slow workloads. Filter by `paparats.tool=search_code` or `paparats.group=β¦` to slice by repo.
- **Custom queries / metrics** β every paparats attribute is indexed. Build APM queries like `paparats.embedding.cache_miss:true AND duration_ms>500` to find slow cache-miss embeddings, or aggregate `paparats.search.result_count` per `paparats.group`.
- **Log correlation** β if you ship paparats stdout to Elastic via Filebeat, the `trace.id` field links a log line back to its span.
**What this is _not_ β honest caveats:**
- Spans are flat (one event = one span), not parented. Service Map will show `paparats-mcp` as an isolated node; you won't see a "search β embedding β Qdrant" waterfall. Use the per-span `duration_ms` attributes for stage timing instead.
- Outbound HTTP to Qdrant / the embed server is not auto-instrumented β to see those as separate dependencies in APM you'd need to enable `@opentelemetry/instrumentation-http` (planned, not shipped). For now, embedding and search latency live on the existing spans.
- Per-request token-savings, top queries, and cross-project usage stay in the local SQLite store β they're aggregations, not events. View them in the built-in `/ui` console, not in APM.
For pure metrics (CPU, GC, RSS, request rates) Elastic Metricbeat or our Prometheus exporter (above) is a better fit than APM.
### Identity attribution
Clients (IDE plugins, CLI) can set `X-Paparats-User`, `X-Paparats-Session`, `X-Paparats-Client`, `X-Paparats-Anchor-Project` headers. The header name for `user` is configurable via `PAPARATS_IDENTITY_HEADER` (default `X-Paparats-User`). Missing header β events are attributed to `anonymous`. There is no cryptographic verification β this is for attribution, not access control.
`GET /api/stats` echoes the resolved identity, useful for verifying header propagation:
```bash
curl -H 'X-Paparats-User: alice' http://localhost:9876/api/stats | jq .identity
```
### Token-savings estimators
Three levels, computed from raw events at query-time:
- **Naive baseline** β what a model would have read if it pulled the whole file for each result.
- **Search-only** β tokens actually returned by `search_code`.
- **Actually consumed** β tokens that the client subsequently fetched via `get_chunk`. The most honest signal, since it discounts noisy results that were never used.
Run `token_savings_report` from any MCP client connected to `/support/mcp`.
### Cross-project noise
When a client passes `X-Paparats-Anchor-Project` (or specifies a single project in the search call), the share of results from _other_ projects in the same group is recorded. Use `cross_project_share` to see how noisy your group's index is for each user.
### Indexer-pipeline visibility
`failed_chunks` aggregates AST parse failures, regex fallbacks, zero-chunk files, and binary skips. `slowest_searches` ranks individual searches by latency.
### Configuration matrix
| Env var | Default | Purpose |
| --------------------------------------- | -------------------------- | ----------------------------------------------------------- |
| `PAPARATS_METRICS` | `false` | Prometheus surface (existing, unchanged) |
| `PAPARATS_ANALYTICS_ENABLED` | `true` | Local SQLite analytics writes |
| `PAPARATS_ANALYTICS_DB_PATH` | `~/.paparats/analytics.db` | Analytics DB file |
| `PAPARATS_ANALYTICS_RETENTION_DAYS` | `90` | Daily prune cutoff |
| `PAPARATS_ANALYTICS_RETENTION_RUN_HOUR` | `3` | Hour-of-day for prune (local time) |
| `PAPARATS_IDENTITY_HEADER` | `X-Paparats-User` | Header name for user attribution |
| `PAPARATS_LOG_RESULT_FILES` | `true` | If `false`, store NULL for `search_results.file` |
| `PAPARATS_LOG_QUERY_TEXT` | `true` | If `false`, store NULL for `search_events.query_text` |
| `PAPARATS_REFORMULATION_WINDOW_MS` | `90000` | Reformulation detection window |
| `PAPARATS_TELEMETRY_SAMPLE_RATE` | `1.0` | Sampling rate (errors are always kept) |
| `PAPARATS_OTEL_ENABLED` | `false` | Enable OTel SDK + OTLP exporter |
| `OTEL_EXPORTER_OTLP_ENDPOINT` | unset | OTLP HTTP endpoint (e.g. `http://localhost:4318/v1/traces`) |
| `OTEL_EXPORTER_OTLP_HEADERS` | unset | OTLP auth headers (`key=value,key2=value2`) |
| `OTEL_SERVICE_NAME` | `paparats-mcp` | OTel resource attribute |
| `OTEL_RESOURCE_ATTRIBUTES` | unset | Extra resource attrs (`key=value,key2=value2`) |
### PII guidance
- File paths and query text are stored locally by default. For shared deployments where paths could leak sensitive info, set `PAPARATS_LOG_RESULT_FILES=false` and/or `PAPARATS_LOG_QUERY_TEXT=false`.
- OTel spans never carry full query text by default β only `paparats.query.hash` and length.
---
## Architecture
```
paparats-mcp/
βββ packages/
β βββ server/ # MCP server (Docker image: ibaz/paparats-server)
β β βββ src/
β β β βββ lib.ts # Public library exports (for programmatic use)
β β β βββ index.ts # HTTP server bootstrap + graceful shutdown
β β β βββ app.ts # Express app + HTTP API routes
β β β βββ indexer.ts # Group-aware indexing, single-parse chunkFile()
β β β βββ searcher.ts # Search with query expansion, cache, metrics
β β β βββ query-expansion.ts # Abbreviation, case, plural expansion
β β β βββ task-prefixes.ts # Jina task prefix detection
β β β βββ query-cache.ts # In-memory LRU search result cache
β β β βββ metrics.ts # Prometheus metrics (opt-in)
β β β βββ ast-chunker.ts # AST-based code chunking (tree-sitter, primary strategy)
β β β βββ chunker.ts # Regex-based code chunking (fallback for unsupported languages)
β β β βββ ast-symbol-extractor.ts # AST-based symbol extraction (module-level only, 11 languages)
β β β βββ ast-queries.ts # Tree-sitter S-expression queries per language
β β β βββ tree-sitter-parser.ts # WASM tree-sitter manager
β β β βββ symbol-graph.ts # Cross-chunk symbol edges (calls/called_by/refs)
β β β βββ embeddings.ts # llama-server provider + SQLite cache
β β β βββ config.ts # .paparats.yml reader + validation
β β β βββ metadata.ts # Tag resolution + auto-detection
β β β βββ metadata-db.ts # SQLite store for git commits + tickets + symbol edges
β β β βββ git-metadata.ts # Git history extraction + chunk mapping
β β β βββ ticket-extractor.ts # Jira/GitHub/custom ticket parsing
β β β βββ mcp-handler.ts # MCP protocol β dual-mode (coding /mcp + support /support/mcp)
β β β βββ watcher.ts # File watcher (chokidar)
β β β βββ arch/ # Architectural memory layer (components, decisions, lessons)
β β β β βββ types.ts # ArchComponent, ArchDecision, ArchLesson, ArchWriteResult
β β β β βββ collection.ts # Per-group Qdrant collection (`paparats_<group>_arch`) lifecycle
β β β β βββ text-embeddings.ts # qwen3-embedding-0.6b text embedder (1024d, last-pooled, llama-server)
β β β β βββ store.ts # CRUD + server-side similarity gate (cosine 0.85 / 0.70)
β β β β βββ context.ts # `arch_context` query β top-N across kinds with age stamps
β β β βββ types.ts # Shared types
β β βββ Dockerfile
β βββ indexer/ # Automated repo indexer (Docker image: ibaz/paparats-indexer)
β β βββ src/
β β β βββ index.ts # Entry: Express mini-server + cron scheduler
β β β βββ config-loader.ts # projects.yml parser + per-repo overrides
β β β βββ config-watcher.ts # chokidar watcher for hot-reloading the project list
β β β βββ repo-manager.ts # parseReposEnv(), cloneOrPull() using simple-git
β β β βββ scheduler.ts # node-cron wrapper
β β β βββ types.ts # IndexerConfig, RepoConfig, RepoOverrides, IndexerFileConfig
β β βββ Dockerfile
β βββ embed/ # llama.cpp llama-server + llama-swap, models pre-baked (Docker image: ibaz/paparats-embed)
β β βββ Dockerfile
β βββ cli/ # CLI tool (npm package: @paparats/cli)
β β βββ src/
β β βββ index.ts # Commander entry
β β βββ docker-compose-generator.ts # Programmatic YAML generation
β β βββ projects-yml.ts # projects.yml + install.json read/write
β β βββ commands/ # install, projects (add/remove/list), lifecycle, edit, etc.
β βββ shared/ # Shared utilities (npm package: @paparats/shared)
β βββ src/
β βββ path-validation.ts # Path validation
β βββ gitignore.ts # Gitignore parsing
β βββ exclude-patterns.ts # Glob exclude normalization
β βββ language-excludes.ts # Language-specific exclude defaults
βββ examples/
βββ paparats.yml.* # Config examples per language
```
---
## Stack
- **Qdrant** β vector database (1 collection per group with `paparats_` prefix for code, plus a separate `paparats_<group>_arch` collection per group for the architectural memory layer; cosine similarity, payload filtering)
- **Embed server** β llama.cpp `llama-server` + `llama-swap` serving local embeddings via `bge-code-v1` for code (1536d, task-specific prefixes) **and `qwen3-embedding-0.6b` for the architectural memory layer** (1024d, last-pooled). Both Apache-2.0. llama-swap routes by model name; models stay resident by default (`EMBED_TTL=0`), with opt-in idle unload via `EMBED_TTL`
- **SQLite** β embedding cache (`~/.paparats/cache/embeddings.db`) + git metadata + symbol edges store (`~/.paparats/metadata.db`)
- **MCP** β Model Context Protocol (SSE for Cursor, Streamable HTTP for Claude Code). Dual endpoints: `/mcp` (coding) and `/support/mcp` (support)
- **TypeScript** monorepo with Yarn workspaces
---
## Integration Examples
### Support Chatbot
Use paparats as the knowledge backend for a product support bot. Connect the bot to the **support endpoint** (`/support/mcp`) for access to `explain_feature`, `recent_changes`, `find_usages`, and other support-oriented tools:
```
User: "How do I configure rate limiting?"
Bot workflow (via /support/mcp):
1. explain_feature("rate limiting", group="my-app")
β returns code locations + recent changes + related modules
2. get_chunk_meta(<chunk_id>)
β returns who last modified it, when, linked tickets
3. Bot synthesizes response in plain language with ticket references
```
### CI/CD reindex on push
Indexing lives in the indexer container. To force a reindex of a project from CI,
trigger the indexer's HTTP endpoint:
```yaml
name: Reindex Paparats
on:
push:
branches: [main]
jobs:
reindex:
runs-on: ubuntu-latest
steps:
- run: |
curl -X POST http://your-paparats-host:9877/trigger \
-H 'Content-Type: application/json' \
-d '{"repos": ["your-org/your-repo"]}'
```
Pass `"force": true` in the body to drop existing chunks first (destructive β use after
schema/config changes). If the project isn't yet in `projects.yml`, add it once
during your initial setup and the indexer's cron + hot-reload will keep it in sync going
forward.
### Code-review assistant
Combine multiple tools to analyze the impact of a pull request:
```
1. explain_feature("the feature being changed")
β understand what the code does and how it connects
2. find_usages({chunk_id: "<changed chunk>", direction: "both"})
β blast radius via the symbol graph
3. search_changes("related area", since="2024-01-01")
β recent changes that might conflict or overlap
```
---
## Embedding Model Setup
Paparats supports three embedding backends. **Pick one** β the choice is
sticky per Qdrant collection (changing it requires reindexing; the server
refuses to mix providers in one collection and surfaces a clear error).
| Provider | Model | Dims | Privacy | Speed (1k chunks) | Cost |
| ------------ | --------------------------- | ----- | -------------- | ------------------------ | --------------------- |
| **llama** | `bge-code-v1` | 1536 | 100% local | ~2β4 min (CPU) | Free, ~1.5 GB on disk |
| **OpenAI** | `text-embedding-3-small` | 1536 | Sent to OpenAI | ~30 s | ~$0.02 / 1 M tokens |
| **Voyage** | `voyage-code-3` | 1024 | Sent to Voyage | ~30 s | ~$0.18 / 1 M tokens |
Valid `EMBEDDING_PROVIDER` values: `llama` | `openai` | `voyage`. Selection
precedence: explicit `EMBEDDING_PROVIDER` β `OPENAI_API_KEY` present β
`VOYAGE_API_KEY` present β `llama`. So setting just your API key in the
environment is enough to switch.
```bash
# OpenAI β cheapest cloud option
export OPENAI_API_KEY=sk-...
docker compose up -d
# Voyage AI β best quality on code per recent benchmarks
export VOYAGE_API_KEY=pa-...
docker compose up -d
# Force a provider explicitly (overrides auto-detect)
export EMBEDDING_PROVIDER=voyage
```
Overrides: `EMBEDDING_MODEL` (defaults: `text-embedding-3-small`,
`voyage-code-3`, `bge-code-v1`) and `EMBEDDING_DIMENSIONS` (1536 /
1024 / 1536). Voyage `voyage-code-3` supports 256/512/1024/2048 via
Matryoshka β set `EMBEDDING_DIMENSIONS` to opt into a non-default size.
### Local (llama) β defaults below
Default model: [BAAI/bge-code-v1](https://huggingface.co/BAAI/bge-code-v1) (served as a Q8_0 GGUF) β code-optimized, Apache-2.0, 1536 dims, 32k context. The embed server is llama.cpp `llama-server` + `llama-swap`: `llama-server` loads the GGUF directly and `llama-swap` routes by model name and lazy-loads it on first request. There is **no Modelfile and no model-registration step** β llama-swap discovers the model by name.
**Recommended:** `paparats install` automates this:
- **Native mode** (`--embed-mode native`, default on macOS): installs the embed server via `brew install llama.cpp mostlygeek/llama-swap/llama-swap` (Metal-accelerated) and downloads the GGUF to `~/.paparats/models/`
- **Docker mode** (`--embed-mode docker`, default on Linux): Uses the `ibaz/paparats-embed` image with `bge-code-v1` + `qwen3-embedding-0.6b` pre-baked β zero setup
**Manual setup (Docker):**
```bash
# Run the pre-baked embed server. llama-swap serves on 8080 inside the container
# (mapped to host 11434) and loads models by name β nothing else to configure.
docker run -d -p 11434:8080 ibaz/paparats-embed:latest
# Verify (llama-swap exposes an OpenAI-style API)
curl http://localhost:8080/health
curl http://localhost:8080/v1/models
```
| Spec | Value |
| ------------ | ----------------------------------- |
| Model | BAAI/bge-code-v1 (Apache-2.0) |
| Dimensions | 1536 |
| Context | 32,768 tokens (recommended β€ 8,192) |
| Quantization | Q8_0 (~1.5 GB) |
| Languages | 15+ programming languages |
Task-specific prefixes (nl2code, code2code, techqa) applied automatically.
### Model benchmarks & methodology
We benchmark embedding models on standard BEIR-format retrieval datasets before
adopting them, on real hardware rather than trusting vendor-reported numbers.
**Methodology.** Each model is served by `llama-server` (Q8_0 GGUF) and queried
through the OpenAI-style `/v1/embeddings` endpoint β the same path the server
uses in production. We embed the full corpus, run brute-force cosine search per
query, and report nDCG / Recall @K against the dataset's published `qrels`. Each
model uses its own recommended query instruction (documents are embedded
unprefixed); pooling matches the architecture (`cls` for BERT-family, `last` for
decoder-family). No reranking β the consumer is an agent that reads the whole
top-K, so recall matters more than exact ordering.
**Code retrieval** β CoIR CosQA (20,604 code docs, 500 queries):
| Model | License | Dim | Recall@5 | nDCG@5 | Recall@10 |
| ----- | ------- | --- | -------- | ------ | --------- |
| **BAAI/bge-code-v1** | Apache-2.0 | 1536 | **0.480** | **0.323** | **0.722** |
| jinaai/jina-code-embeddings-1.5b | CC-BY-NC | 1536 | 0.408 | 0.269 | 0.626 |
**Prose retrieval** β SciFact (5,183 docs, 300 queries):
| Model | License | Dim | Recall@5 | nDCG@5 | Recall@10 |
| ----- | ------- | --- | -------- | ------ | --------- |
| jinaai/jina-embeddings-v3 + retrieval-LoRA | CC-BY-NC | 1024 | 0.795 | 0.709 | 0.843 |
| **Qwen/Qwen3-Embedding-0.6B** | Apache-2.0 | 1024 | 0.775 | 0.681 | 0.831 |
| BAAI/bge-m3 | MIT | 1024 | 0.720 | 0.625 | 0.783 |
**Prose retrieval** β FiQA business/finance (16.7k-doc subsample, 648 queries):
| Model | License | Dim | Recall@5 | nDCG@5 | Recall@10 |
| ----- | ------- | --- | -------- | ------ | --------- |
| **Qwen/Qwen3-Embedding-0.6B** | Apache-2.0 | 1024 | **0.575** | **0.555** | **0.666** |
| BAAI/bge-m3 | MIT | 1024 | 0.499 | 0.496 | 0.596 |
**Takeaways.**
- For **code**, `bge-code-v1` beats `jina-code-embeddings` by ~7pt Recall@5, is
permissively licensed (Apache-2.0 vs CC-BY-NC), and keeps the same 1536 dims.
- For **prose/docs**, `Qwen3-Embedding-0.6B` beats `bge-m3` by 5β8pt Recall@5
(and by more on business prose), is Apache-2.0, and keeps 1024 dims β at the
cost of ~2Γ slower indexing on CPU. `jina-embeddings-v3 + LoRA` scores highest
but its CC-BY-NC license rules it out for commercial/self-hosted adoption.
- A cross-encoder reranker (`bge-reranker-v2-m3`) added +4β6pt nDCG but ~2.7s per
query on CPU and mostly reorders rather than improving recall β not worth it
when an agent consumes the top-K directly.
- **Late chunking** was evaluated and rejected: decoder-family models (Qwen3)
can't do it (causal attention), and it degrades multi-topic documents anyway.
Structural / per-topic chunking is the dominant quality lever, not the model.
> Numbers were measured on Apple-silicon (Metal) with Q8_0 quantization; absolute
> values shift with hardware, quantization, and corpus, but the relative ordering
> between models is what drives model selection here.
---
## Comparison with Alternatives
### Feature Matrix
#### Deployment
| Feature | Paparats | Vexify | SeaGOAT | Augment | Sourcegraph | Greptile | Bloop |
| :---------- | :------: | :----: | :-----: | :-----: | :---------: | :------: | :---: |
| Open source | β
MIT | β
MIT | β
MIT | β | β οΈ Partial | β | β οΈ 1 |
| Fully local | β
| β
| β
| β οΈ No 2 | β | β | β
|
#### Search Quality
| Feature | Paparats | Vexify | SeaGOAT | Augment | Sourcegraph | Greptile | Bloop |
| :-------------- | :-------: | :----: | :------: | :--------: | :---------: | :--------: | :----: |
| Code embeddings | β
Jina 3 | β οΈ 4 | β 5 | β οΈ Partial | β οΈ Partial | β οΈ Partial | β
|
| Vector database | β
Qdrant | SQLite | ChromaDB | Propri. | Propri. | pgvector | Qdrant |
| AST chunking | β
| β | β | β οΈ Partial | β οΈ Partial | β οΈ Partial | β
|
| Query expansion | β
6 | β | β | β οΈ Partial | β οΈ Partial | β οΈ Partial | β |
#### Developer Experience
| Feature | Paparats | Vexify | SeaGOAT | Augment | Sourcegraph | Greptile | Bloop |
| :----------------- | :-------: | :--------: | :--------: | :--------: | :---------: | :--------: | :--------: |
| Real-time watching | β
Auto | β | β | β οΈ CI/CD | β
| β οΈ Partial | β οΈ Partial |
| Embedding cache | β
SQLite | β οΈ Partial | β | β οΈ Partial | β οΈ Partial | β οΈ Partial | β |
| Multi-project | β
Groups | β
| β | β
| β
| β
| β
|
| One-cmd install | β
| β οΈ Partial | β οΈ Partial | β | β | β | β |
#### AI Integration
| Feature | Paparats | Vexify | SeaGOAT | Augment | Sourcegraph | Greptile | Bloop |
| :--------------------- | :------: | :----: | :-----: | :--------: | :---------: | :------: | :---: |
| MCP native | β
| β
| β | β
| β | β οΈ API | β |
| Symbol graph | β
| β | β | β | β οΈ Partial | β | β |
| Token metrics | β
| β | β | β οΈ Partial | β | β | β |
| Git history | β
| β | β | β | β οΈ Partial | β | β |
| Ticket extraction | β
| β | β | β | β | β | β |
| Architectural memory 7 | β
ADRs | β | β | β | β | β | β |
#### Pricing
| | Paparats | Vexify | SeaGOAT | Augment | Sourcegraph | Greptile | Bloop |
| :--- | :---------: | :---------: | :---------: | :-----: | :---------: | :------: | :---------: |
| Cost | β
**Free** | β
**Free** | β
**Free** | β Paid | β Paid | β Paid | β οΈ Archived |
<details>
<summary>Notes</summary>
1. Bloop archived January 2, 2025
2. Augment Context Engine indexes locally but stores vectors in cloud
3. Jina Code Embeddings 1.5B (1536 dims) with task-specific prefixes (nl2code, code2code, techqa)
4. Vexify supports Ollama models but limited to specific embeddings (jina-embeddings-2-base-code, nomic-embed-text)
5. SeaGOAT locked to all-MiniLM-L6-v2 (384 dims, general-purpose)
6. Abbreviations, case variants, plurals, filler word removal
7. Agent-maintained components / decisions (ADRs) / lessons in a second Qdrant collection per group; server-side similarity gate deduplicates writes, `supersedes` links replace stale decisions, every card carries an "updated N ago" stamp on read
</details>
---
## Token Savings Metrics
### What we measure (and what we don't)
Paparats provides **estimated** token savings to help you understand the order of magnitude of context reduction. These are heuristics, not precise measurements.
#### Per-search response
```json
{
"metrics": {
"tokensReturned": 150,
"estimatedFullFileTokens": 5000,
"tokensSaved": 4850,
"savingsPercent": 97
}
}
```
| Field | Calculation | Reality Check |
| ------------------------- | --------------------------- | -------------------------------------------------------------- |
| `tokensReturned` | `ceil(content.length / 4)` | Based on actual returned content; /4 is rough approximation |
| `estimatedFullFileTokens` | `ceil(endLine * 50 / 4)` | **Heuristic**: assumes 50 chars/line, never loads actual files |
| `tokensSaved` | `estimated - returned` | **Derived**: difference between two estimates |
| `savingsPercent` | `(saved / estimated) * 100` | **Relative**: percentage of heuristic estimate |
#### Cumulative stats
```bash
curl -s http://localhost:9876/api/stats | jq '.usage'
```
```json
{
"searchCount": 47,
"totalTokensSaved": 152340,
"avgTokensSavedPerSearch": 3241
}
```
These are **sums of estimates**, not measured token counts from a real tokenizer.
---
## License
MIT
---
## Releasing (maintainers)
Releases are driven by [Changesets](https://github.com/changesets/changesets). Versioning + CHANGELOG generation happen in CI; **publishing to npm and tagging happen locally** from a maintainer machine that's authenticated with npm. There are no npm credentials in CI.
### Authoring a changeset (per PR)
```bash
yarn changeset
# Pick affected packages, bump type (patch/minor/major), and write the user-facing summary.
git add .changeset/
git commit -m "chore: changeset"
```
All four packages (`@paparats/shared`, `@paparats/cli`, `@paparats/server`, `@paparats/indexer`) are kept on a **fixed version** β pick any one and the rest are bumped to match.
### How a release happens
**1. CI opens a release PR (automatic).** The [Release workflow](.github/workflows/release.yml) runs on every push to `main`. If pending `.changeset/*.md` files exist, it opens (or updates) a `chore: release` PR with: version bumps in every `package.json`, regenerated per-package `CHANGELOG.md` files, `server.json` synced via `scripts/sync-server-json.js`, and the consumed `.changeset/*.md` files deleted.
**2. Maintainer merges the release PR.** No further CI publish step runs.
**3. Maintainer publishes locally.** From a clean checkout of `main` after the merge:
```bash
git checkout main && git pull
yarn release:local # or `--dry-run` to preview
```
`yarn release:local` runs `scripts/release-local.sh`, which:
- refuses to run unless you're on `main`, the tree is clean, and you're in sync with `origin/main`;
- refuses if any pending `.changeset/*.md` are present (means the release PR wasn't merged);
- reads the new version from `packages/cli/package.json`;
- builds, runs `yarn changeset publish` (skips already-published versions), then tags `vX.Y.Z` and pushes the tag.
**4. Downstream workflows fire on the tag.** Pushing `vX.Y.Z` triggers [docker-publish.yml](.github/workflows/docker-publish.yml) and [publish-mcp.yml](.github/workflows/publish-mcp.yml) automatically.
### Required credentials
| Where | What | Purpose |
| ----- | ----------------------------------- | ------------------------------------------------- |
| CI | `GITHUB_TOKEN` (auto) | Open/update the `chore: release` PR |
| Local | `npm login` (or `NPM_TOKEN` in env) | `yarn changeset publish` to publish `@paparats/*` |
No npm token lives in GitHub secrets β publishing is intentionally a manual, authenticated step.
### Manual / fallback flows
`./scripts/release-docker.sh --push` still builds and pushes the Docker images by hand if needed (e.g. between official releases). It reads the version from `package.json`.
### Docker images
| Image | Source | Size |
| ----------------------- | ----------------------------- | ---------------------- |
| `ibaz/paparats-server` | `packages/server/Dockerfile` | ~200 MB |
| `ibaz/paparats-indexer` | `packages/indexer/Dockerfile` | ~200 MB |
| `ibaz/paparats-embed` | `packages/embed/Dockerfile` | ~2.3 GB (includes models) |
---
## Contributing
Contributions welcome! Areas of interest:
- Additional language support (PHP, Elixir, Scala, Kotlin, Swift)
- Alternative embedding providers (OpenAI, Cohere, local GGUF via llama.cpp)
- Performance optimizations (chunking strategies, cache eviction)
- Agent use cases (support bots, QA automation, code analytics)
Open an issue or pull request to get started.
---
## Links
- [BAAI/bge-code-v1](https://huggingface.co/BAAI/bge-code-v1) β code embedding model
- [Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) β arch/docs text embedding model
- [Qdrant](https://qdrant.tech) β vector database
- [llama.cpp](https://github.com/ggml-org/llama.cpp) β local embedding runtime (`llama-server`)
- [llama-swap](https://github.com/mostlygeek/llama-swap) β model router / lazy loader in front of llama-server
- [MCP](https://modelcontextprotocol.io) β Model Context Protocol
---
**Star the repo if Paparats helps you code faster!**
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues