confluence-rag
by tealtadpole
README.md
# confluence-rag
A local RAG index over **Confluence Data Center / Server**, served to Claude Code
(or any MCP client) as a local **stdio MCP server**.
```
Confluence REST API ──(your PAT, read-only GETs)──► sync ──► SQLite index (~/.local/share/confluence-rag)
├─ page text
├─ chunks + embeddings (local fastembed model)
└─ FTS5 keyword index
Claude Code ◄──stdio / MCP──► confluence-rag serve ──► hybrid search (vectors + BM25, rank fusion)
```
Everything except the Confluence API calls runs on your machine: embeddings come from a small
local ONNX model, and nothing is sent anywhere else. Text the MCP tools return does go to the
model provider as part of your Claude Code conversation.
## Requirements
| What | Details |
|---|---|
| **OS** | Linux or macOS. Windows isn't supported (the sync lock uses `fcntl`); use WSL instead. |
| **Python** | 3.11 or newer (uses the built-in `tomllib`). On Debian/Ubuntu also install `python3-venv`. |
| **Database** | **Nothing to install.** The index is a single SQLite file created automatically in `data_dir`. It uses SQLite's FTS5 full-text extension, which is built into standard Python builds. Check with: `python3 -c "import sqlite3; sqlite3.connect(':memory:').execute('create virtual table t using fts5(x)'); print('ok')"` |
| **Confluence** | **Data Center / Server 7.9 or newer** (the first version with Personal Access Tokens). Not Confluence Cloud (see Limitations). |
| **Access token** | A Personal Access Token: Confluence → your avatar → *Settings* → *Personal Access Tokens* → *Create token*. Read access is enough. Set an expiry date. |
| **Network** | HTTPS access to your Confluence (VPN if it's internal). On the **first run only**, access to Hugging Face to download the embedding model. On a locked-down network, see "Offline model" below. |
| **Disk** | ~280 MB for the Python environment, ~65 MB for the model, plus the index: about 100–150 MB per 10,000 pages of average length (text + embeddings + keyword index). |
| **Memory** | ~700 MB while syncing or serving (mostly the embedding model), plus ~1.5 KB per chunk for the in-memory vectors (100k chunks ≈ 150 MB). |
| **CPU / time** | Embedding runs on the CPU at roughly **5–10 chunks per second** on a 4-core laptop, and a typical page is 2–4 chunks. A 2,000-page space takes ~10–15 minutes; 100,000 pages takes roughly 10 hours. Only the first sync is slow: later syncs process only changed pages. |
| **Claude Code** | Optional. Any MCP client that can launch a stdio server works; the CLI also works on its own. |
Python packages (installed by `pip` below): `mcp` 2.x, `fastembed` (ONNX runtime, no PyTorch or GPU needed), `beautifulsoup4`, `requests`, `numpy`.
## Installation
```bash
git clone https://github.com/tealtadpole/local-confluence-RAG.git
cd local-confluence-RAG
python3 -m venv .venv
.venv/bin/pip install -e . # add ".[dev]" to also install pytest
cp config.example.toml config.toml # gitignored
chmod 600 config.toml
$EDITOR config.toml # at least base_url and spaces
export CONFLUENCE_PAT='...' # or set token = "..." in config.toml
```
Then:
```bash
.venv/bin/confluence-rag check # tests URL + token, counts pages per space
.venv/bin/confluence-rag sync # first sync downloads and indexes everything
.venv/bin/confluence-rag search "how do I request VPN access"
.venv/bin/confluence-rag status
```
Run `check` first: it shows how many pages each space has, so you can estimate the first sync
time from the table above before starting it. A sync can be stopped with Ctrl+C and resumed later.
**Offline model:** the model is downloaded once into `data_dir/models`. If Hugging Face is blocked,
download it on a machine with internet access and copy the folder over:
```bash
.venv/bin/python -c "from fastembed import TextEmbedding; TextEmbedding('BAAI/bge-small-en-v1.5', cache_dir='models')"
# then copy ./models to <data_dir>/models (default ~/.local/share/confluence-rag/models)
```
**Common problems**
| Symptom | Fix |
|---|---|
| `HTTP 401 ... expired or been revoked` | Create a new PAT and update `CONFLUENCE_PAT` / `token`. |
| `Expected JSON ... got 'text/html'` | `base_url` is wrong (missing context path such as `/confluence`?) or an SSO page or proxy is intercepting. |
| `SSL: CERTIFICATE_VERIFY_FAILED` | Corporate TLS inspection: set `verify_ssl = "/path/to/corp-ca.pem"`. |
| `No module named 'tomllib'` | Python is older than 3.11. |
| `ensurepip is not available` | `sudo apt install python3-venv` |
## Use it from Claude Code
Run this from the repository folder (`$PWD` becomes the absolute path):
```bash
claude mcp add -s user confluence \
-e CONFLUENCE_PAT="$CONFLUENCE_PAT" \
-- "$PWD/.venv/bin/confluence-rag" --config "$PWD/config.toml" serve
```
`-e` stores the token in Claude Code's MCP config (`~/.claude.json`). If you'd rather keep it
only in `config.toml`, leave out `-e`. Restart Claude Code, then check `/mcp`. Tools:
| Tool | What it does |
|---|---|
| `search_confluence(query, max_results?, space?)` | Hybrid search; returns passages with title, URL, space, date, page_id |
| `get_confluence_page(page_id)` | Full text of one indexed page |
| `confluence_index_status()` | Page counts, last sync, last errors |
| `refresh_confluence_index(full?)` | Starts a background sync |
To allow the read-only tools without permission prompts, add them to `permissions.allow` in
`~/.claude/settings.json`, e.g. `"mcp__confluence__search_confluence"`.
## How the index stays fresh
Each sync, per space:
1. **List** every page id and version. This is cheap: no bodies, 200 per request.
2. **Fetch** only pages that are new, or whose version or space changed. More than
`bulk_threshold` changes in a space downloads the space in bulk; fewer fetches them one by one.
3. **Delete** local pages that no longer exist anywhere: deleted, newly restricted from you,
or in a space you removed from the config. Deletion is **skipped** if any space failed to
list, so a Confluence outage never empties your index.
Triggers:
- **Automatically**, while the MCP server runs, when the index is older than
`sync.auto_refresh_minutes` (default 60; 0 disables).
- **On demand** with `confluence-rag sync` (e.g. from cron) or the `refresh_confluence_index` tool.
- Only one sync runs at a time (a file lock), so the CLI and server don't collide.
`sync --full` re-downloads and re-indexes everything. Changing `embedding_model`,
`chunk_size` or `chunk_overlap` clears the index and rebuilds it automatically on the next sync.
Interrupted syncs are safe: each page is committed individually, so the next run continues.
**Token expiry:** a PAT can't be renewed through the API. When it expires, Confluence returns
401. Syncs then stop with a clear message ("create a new Personal Access Token"), visible in
`status` and `confluence_index_status`. The existing index stays searchable. Replace the token
and sync again.
## Security notes
- **Your PAT sees what you see.** The index contains every page you can read. Don't share the
index or expose this server to other people: they would get your access.
- The index holds full page text. `data_dir` is created with mode 700. Treat it like the wiki itself.
- Keep the token out of files where you can (`token_env`). If it's in `config.toml`, the tool
warns unless the file is `chmod 600`. It's never written to logs.
- Search results are labelled as reference data, not instructions. Wiki pages are editable by
many people and could contain prompt-injection text.
- Check your company's policy before indexing many spaces. `spaces = ["*"]` indexes every
space your token can see.
- Behind a corporate TLS proxy, set `verify_ssl = "/path/to/corp-ca.pem"`.
## Limitations
- Confluence **Data Center / Server** only (Bearer PAT, REST v1). Cloud uses email + API
token and the v2 API; it would need a different client in `confluence.py`.
- Pages only: no blog posts, comments or attachments.
- Vector search is brute force in memory (numpy). Fine for tens of thousands of pages. For
millions, move vectors to pgvector, Qdrant or sqlite-vec.
- Macros that render content dynamically (Jira issues, child page lists, includes) are stored
without their rendered output.
## Layout
```
src/confluence_rag/
config.py load + validate config.toml, resolve the token
confluence.py read-only REST client: pagination, retries, 429 handling, auth errors
textproc.py storage-format XHTML -> text; heading-aware chunking
embed.py fastembed wrapper (local ONNX model)
store.py SQLite: pages, chunks + embeddings, FTS5
search.py hybrid search with reciprocal rank fusion
sync.py list -> diff -> fetch -> index -> prune
server.py MCP tools + background refresher
cli.py check / sync / search / status / serve
tests/ runs against a fake Confluence HTTP server (`.venv/bin/pytest`)
```
## License
[MIT](LICENSE). Not affiliated with or endorsed by Atlassian. "Confluence" is a trademark of Atlassian.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues