Skip to main content
Glama

confluence-rag

A local RAG index over Confluence Data Center / Server, served to Claude Code (or any MCP client) as a local stdio MCP server.

Confluence REST API ──(your PAT, read-only GETs)──► sync ──► SQLite index (~/.local/share/confluence-rag)
                                                               ├─ page text
                                                               ├─ chunks + embeddings (local fastembed model)
                                                               └─ FTS5 keyword index
Claude Code ◄──stdio / MCP──► confluence-rag serve ──► hybrid search (vectors + BM25, rank fusion)

Everything except the Confluence API calls runs on your machine: embeddings come from a small local ONNX model, and nothing is sent anywhere else. Text the MCP tools return does go to the model provider as part of your Claude Code conversation.

Requirements

What

Details

OS

Linux or macOS. Windows isn't supported (the sync lock uses fcntl); use WSL instead.

Python

3.11 or newer (uses the built-in tomllib). On Debian/Ubuntu also install python3-venv.

Database

Nothing to install. The index is a single SQLite file created automatically in data_dir. It uses SQLite's FTS5 full-text extension, which is built into standard Python builds. Check with: python3 -c "import sqlite3; sqlite3.connect(':memory:').execute('create virtual table t using fts5(x)'); print('ok')"

Confluence

Data Center / Server 7.9 or newer (the first version with Personal Access Tokens). Not Confluence Cloud (see Limitations).

Access token

A Personal Access Token: Confluence → your avatar → Settings → Personal Access Tokens → Create token. Read access is enough. Set an expiry date.

Network

HTTPS access to your Confluence (VPN if it's internal). On the first run only, access to Hugging Face to download the embedding model. On a locked-down network, see "Offline model" below.

Disk

~280 MB for the Python environment, ~65 MB for the model, plus the index: about 100–150 MB per 10,000 pages of average length (text + embeddings + keyword index).

Memory

~700 MB while syncing or serving (mostly the embedding model), plus ~1.5 KB per chunk for the in-memory vectors (100k chunks ≈ 150 MB).

CPU / time

Embedding runs on the CPU at roughly 5–10 chunks per second on a 4-core laptop, and a typical page is 2–4 chunks. A 2,000-page space takes ~10–15 minutes; 100,000 pages takes roughly 10 hours. Only the first sync is slow: later syncs process only changed pages.

Claude Code

Optional. Any MCP client that can launch a stdio server works; the CLI also works on its own.

Python packages (installed by pip below): mcp 2.x, fastembed (ONNX runtime, no PyTorch or GPU needed), beautifulsoup4, requests, numpy.

Related MCP server: Confluence MCP Server

Installation

git clone https://github.com/tealtadpole/local-confluence-RAG.git
cd local-confluence-RAG
python3 -m venv .venv
.venv/bin/pip install -e .              # add ".[dev]" to also install pytest

cp config.example.toml config.toml      # gitignored
chmod 600 config.toml
$EDITOR config.toml                     # at least base_url and spaces
export CONFLUENCE_PAT='...'             # or set token = "..." in config.toml

Then:

.venv/bin/confluence-rag check          # tests URL + token, counts pages per space
.venv/bin/confluence-rag sync           # first sync downloads and indexes everything
.venv/bin/confluence-rag search "how do I request VPN access"
.venv/bin/confluence-rag status

Run check first: it shows how many pages each space has, so you can estimate the first sync time from the table above before starting it. A sync can be stopped with Ctrl+C and resumed later.

Offline model: the model is downloaded once into data_dir/models. If Hugging Face is blocked, download it on a machine with internet access and copy the folder over:

.venv/bin/python -c "from fastembed import TextEmbedding; TextEmbedding('BAAI/bge-small-en-v1.5', cache_dir='models')"
# then copy ./models to <data_dir>/models (default ~/.local/share/confluence-rag/models)

Common problems

Symptom

Fix

HTTP 401 ... expired or been revoked

Create a new PAT and update CONFLUENCE_PAT / token.

Expected JSON ... got 'text/html'

base_url is wrong (missing context path such as /confluence?) or an SSO page or proxy is intercepting.

SSL: CERTIFICATE_VERIFY_FAILED

Corporate TLS inspection: set verify_ssl = "/path/to/corp-ca.pem".

No module named 'tomllib'

Python is older than 3.11.

ensurepip is not available

sudo apt install python3-venv

Use it from Claude Code

Run this from the repository folder ($PWD becomes the absolute path):

claude mcp add -s user confluence \
  -e CONFLUENCE_PAT="$CONFLUENCE_PAT" \
  -- "$PWD/.venv/bin/confluence-rag" --config "$PWD/config.toml" serve

-e stores the token in Claude Code's MCP config (~/.claude.json). If you'd rather keep it only in config.toml, leave out -e. Restart Claude Code, then check /mcp. Tools:

Tool

What it does

search_confluence(query, max_results?, space?)

Hybrid search; returns passages with title, URL, space, date, page_id

get_confluence_page(page_id)

Full text of one indexed page

confluence_index_status()

Page counts, last sync, last errors

refresh_confluence_index(full?)

Starts a background sync

To allow the read-only tools without permission prompts, add them to permissions.allow in ~/.claude/settings.json, e.g. "mcp__confluence__search_confluence".

How the index stays fresh

Each sync, per space:

  1. List every page id and version. This is cheap: no bodies, 200 per request.

  2. Fetch only pages that are new, or whose version or space changed. More than bulk_threshold changes in a space downloads the space in bulk; fewer fetches them one by one.

  3. Delete local pages that no longer exist anywhere: deleted, newly restricted from you, or in a space you removed from the config. Deletion is skipped if any space failed to list, so a Confluence outage never empties your index.

Triggers:

  • Automatically, while the MCP server runs, when the index is older than sync.auto_refresh_minutes (default 60; 0 disables).

  • On demand with confluence-rag sync (e.g. from cron) or the refresh_confluence_index tool.

  • Only one sync runs at a time (a file lock), so the CLI and server don't collide.

sync --full re-downloads and re-indexes everything. Changing embedding_model, chunk_size or chunk_overlap clears the index and rebuilds it automatically on the next sync.

Interrupted syncs are safe: each page is committed individually, so the next run continues.

Token expiry: a PAT can't be renewed through the API. When it expires, Confluence returns 401. Syncs then stop with a clear message ("create a new Personal Access Token"), visible in status and confluence_index_status. The existing index stays searchable. Replace the token and sync again.

Security notes

  • Your PAT sees what you see. The index contains every page you can read. Don't share the index or expose this server to other people: they would get your access.

  • The index holds full page text. data_dir is created with mode 700. Treat it like the wiki itself.

  • Keep the token out of files where you can (token_env). If it's in config.toml, the tool warns unless the file is chmod 600. It's never written to logs.

  • Search results are labelled as reference data, not instructions. Wiki pages are editable by many people and could contain prompt-injection text.

  • Check your company's policy before indexing many spaces. spaces = ["*"] indexes every space your token can see.

  • Behind a corporate TLS proxy, set verify_ssl = "/path/to/corp-ca.pem".

Limitations

  • Confluence Data Center / Server only (Bearer PAT, REST v1). Cloud uses email + API token and the v2 API; it would need a different client in confluence.py.

  • Pages only: no blog posts, comments or attachments.

  • Vector search is brute force in memory (numpy). Fine for tens of thousands of pages. For millions, move vectors to pgvector, Qdrant or sqlite-vec.

  • Macros that render content dynamically (Jira issues, child page lists, includes) are stored without their rendered output.

Layout

src/confluence_rag/
  config.py      load + validate config.toml, resolve the token
  confluence.py  read-only REST client: pagination, retries, 429 handling, auth errors
  textproc.py    storage-format XHTML -> text; heading-aware chunking
  embed.py       fastembed wrapper (local ONNX model)
  store.py       SQLite: pages, chunks + embeddings, FTS5
  search.py      hybrid search with reciprocal rank fusion
  sync.py        list -> diff -> fetch -> index -> prune
  server.py      MCP tools + background refresher
  cli.py         check / sync / search / status / serve
tests/           runs against a fake Confluence HTTP server (`.venv/bin/pytest`)

License

MIT. Not affiliated with or endorsed by Atlassian. "Confluence" is a trademark of Atlassian.

Related MCP Connectors

Related MCP Servers