Skip to main content
Glama
msg43

imessage-index

by msg43
README.md
# imessage-index

A local-first iMessage retrieval system. Extracts your full message
history, resolves identities, segments conversations topically, embeds
and indexes them, and serves hybrid full-text + vector search over an
MCP server — so an AI assistant can actually answer questions about
your own correspondence.

Designed to run unattended on a headless Mac, with all derived state on
an encrypted volume and nothing leaving the machine unless you
explicitly allowlist it.

---

## Is this for you?

**What it does.** It reads your Mac's Messages history, figures out who
each contact really is across their various numbers and emails, groups
messages into topical conversations, and builds a local search index —
full-text plus vector — that an AI assistant (Claude Desktop or Claude
Code, over MCP) can query. Everything stays on your machine unless you
explicitly export a filtered subset.

**Who it's for right now.** Someone comfortable running commands in a
terminal, installing developer tools (Python, Rust, PostgreSQL), and
troubleshooting when a step doesn't go cleanly — not a turnkey app for a
non-technical user yet. The step-by-step guide in
[`docs/install-macos.md`](docs/install-macos.md) walks through the whole
setup and the [`scripts/doctor.py`](scripts/doctor.py) diagnostic checks
your machine before you start, but real command-line comfort still
helps.

**Hardware.** An Apple Silicon Mac (M1 or later) — the pipeline depends
on macOS-only APIs (Full Disk Access to Messages, the Contacts
framework, Apple Vision) and on Apple's MLX framework for local model
inference, which does not run on Intel Macs. RAM: the README's own
requirement below is 32 GB unified memory for the full 8B-class local
models; smaller models work with less, but no specific lower figure has
been measured [UNVERIFIED].

**Time and disk.** Not measured end to end for a first-time setup
[UNVERIFIED] — expect at least an hour between installing dependencies,
downloading pinned models, and running the reranker conversion step (see
[Requirements](#requirements) below). Disk usage depends heavily on the
size of your own Messages history and attachments; no general figure has
been measured [UNVERIFIED].

**Before you read further**, the section immediately below is important:
this project has never been run end to end against real models, so
treat it as a rigorous, well-tested foundation rather than a finished
product.

---

## ⚠️ Read this before you clone

**No model has run on the pipeline end to end yet.** With a plain
`uv sync --extra dev` and no database, the test suite is 1,803 passed
and 503 skipped (most skips need a scratch PostgreSQL; the rest need the
`models` and/or `export` extras, which are not required for a plain
install). With every extra installed (`--extra dev --extra models
--extra export`) and still no database, it's 1,847 passed and 477
skipped, almost all of those needing PostgreSQL — measured 2026-09-24.
The CLI works,
migrations apply against real PostgreSQL + pgvector, and the snapshot →
extract → identity stages have run against a real corpus. Segmentation,
embedding, enrichment and retrieval have only ever run on the fake
providers or on tiny stand-in models (see
[what is still unverified](#replacing-the-model-providers)).

**It ships deterministic *fake* model providers alongside the real
wiring.** The embedding, reranking, captioning, OCR, transcription, and
segmentation providers each have a correctly-dimensioned stub that lets
the pipeline run end to end in tests. `imsg.providers.factory` selects
between the real implementations and those stubs with one config field,
`models.backend` (`real` is the default; `fake` is explicit opt-in), and
every command that builds providers prints `models: backend=<real|fake>`
first. **Run the pipeline with `fake` and every stage will report success
while your search results are meaningless.** The real implementations
and their pins are described in
[Replacing the model providers](#replacing-the-model-providers).

Treat this as a **thoroughly tested skeleton with a complete design
behind it**, not a working product. If you want something that works
today, this isn't it. If you want a rigorous starting point that has
already made and documented the non-obvious decisions, read on.

### What is genuinely done

| Area | State |
|---|---|
| Schema + migrations | Complete; applied against live PostgreSQL 17 + pgvector |
| Pipeline stages (snapshot → extract → identity → segment → enrich → embed → sync → export) | Implemented; unit + integration tested |
| Hybrid retrieval (BM25 + text vector + multimodal vector, RRF-fused, reranked) | Implemented |
| Local MCP surface (stdio) | Implemented |
| Public MCP surface (StreamableHTTP + OAuth) | Implemented; never exposed |
| Export gate (default-deny, plan/approve/push/purge) | Implemented and wired to the CLI; **the GCS + Discovery Engine transport has never run against a live API** |
| Eval harness (nDCG@k, recall@k, MRR) | Implemented; metrics verified against hand-computed fixtures |
| Real model providers (MLX text/reranker/boundary LLM, Apple Vision OCR, mlx-whisper, mlx-vlm captions, PE-Core) | Implemented behind `models.backend: real` (the default); every pinned model smoke-run once on a synthetic input (2026-09-14/15, results in the lock); **never run on the pipeline end to end** |
| Nightly backup (`imsg backup`) | Implemented and wired to the CLI; verified `pg_dump` + FTS sidecar copy, 14 retained. **Scope is deliberately narrow — see below** |
| AT-1 auth probe (`imsg mcp public --probe`) | Wired to the CLI and fully tested up to the point where live OAuth tokens are required; **never run against real tokens or a live gate** |
| Runs against real data | Snapshot → extract → identity: yes. Segment / embed / enrich with the pinned models: **never** |

---

## Architecture in one pass

```
chat.db ──.backup──▶ snapshot ──▶ extract ──▶ identity resolution
                                                     │
                                              segmentation
                                                     │
                                  ┌──────────────────┴─────────────┐
                             text bodies                     attachments
                                                                   │
                                          ┌────────────────────────┼───────────┐
                                    PDF text layer              images    audio/video
                                          └────────────────────────┼───────────┘
                                                                   │
                                      embed ──▶ pgvector + FTS5 + multimodal vector
                                                                   │
                                          ┌────────────────────────┴──────────┐
                                    local MCP                         filtered export
                                   (full corpus)                       (allowlisted)
```

Four choices that shaped everything else:

- **Segments, not messages.** Retrieval returns topically coherent
  conversation segments. Message-granular results flood a model's
  context with "sounds good" and prevent reconstructing what was
  actually discussed.
- **Everything becomes text first.** OCR output, captions, transcripts,
  and PDF text layers all embed into one text space, so failures stay
  debuggable — when something doesn't surface you can read the
  extracted text and see why. A second multimodal vector runs alongside
  for visual similarity.
- **Identity resolution precedes segmentation.** Nothing downstream keys
  on a raw handle. One person legitimately has many numbers, emails and
  aliases across a decade.
- **Default deny on export.** Nothing reaches an external index unless
  explicitly allowlisted, and a group thread requires *every*
  participant allowlisted.

---

## The export gate

The one path by which message content can leave the machine, so every
command on it is shaped to refuse. Default deny: a thread exports only
if *every* participant — and every message and tapback sender, the owner
included — is explicitly allowlisted. Attachments are gated separately
from text bodies.

```bash
uv run imsg export plan                  # eligibility → staged bytes → review report
uv run imsg export approve <run-id>      # pins the exact bytes you reviewed
uv run imsg export push <run-id>         # re-verifies, then promotes
uv run imsg export purge-person <who>    # revocation; exempt from the approval gate
uv run imsg export unclassified-report   # weekly: whose threads are still unclassified
```


`plan` writes the review report the whole design rests on — per thread:
participants, message count, date range, sample lines. `approve` pins
the manifest hash and every staged file hash. `push` re-checks all of
that **and re-derives eligibility from the live database**, because the
hashes prove the bytes did not change, not that the world did not: a
participant added to a group between approval and push leaves every hash
green while changing who is in the export. Any drift aborts the push and
requires a new plan.

Revocation is deliberately faster than export — it only ever narrows
scope — so a purge needs no approval, though every drift check still
applies and the run is recorded in full.

`plan`, `push`, `purge-person` and `unclassified-report` all take
`--dry-run`. `push --dry-run` runs every verification the real push runs
and builds no transport at all, so rehearsing the gate cannot reach the
network.

**Nothing can reach Google without a credential you named.**
`export.gcp_credentials` is a `keychain:` / `env:` secret reference with
**no default**; with it unset, `push` refuses before it opens the
database or imports a Google client library. Leave it out until you mean
it.

**Honest limit.** A purge reaches the Discovery Engine index and the GCS
bucket. Copies already swept into organizational retention, backups, or
another person's hands are beyond it. The gate at export time is the
actual protection, which is why it denies by default.

---

## Operations

Seven `com.imsgindex.*` LaunchAgents are rendered by `imsg
install-agents`, and every command they invoke exists — installing them
schedules no job that fails nightly.

```bash
uv run imsg backup                       # daily 04:00: verified pg_dump + FTS copy, 14 kept
uv run imsg backup --dry-run             # preconditions + the retention plan, writes nothing
uv run imsg status                       # mount, Postgres, disk, posture, unclassified threads
```

**What `imsg backup` covers, and what it does not.** In scope: the
Postgres dump (everything the pipeline derived exists only there, and
it is the one component that can suffer *logical* corruption) and the
FTS5 sidecar (rebuildable, copied for recovery speed, under SPEC's
checkpoint + integrity-check conditions). Out of scope, on purpose:

- **`attachments/` (~147 GB)** — 14 nightly copies of it would be over
  two terabytes on the same volume as the original, it is a
  content-addressed cache that `imsg backfill-attachments` rebuilds,
  and a blob store does not suffer the logical corruption these copies
  defend against. **But:** anything iCloud has already purged exists
  nowhere else, and that genuinely needs an independently encrypted
  off-box archive, which this command is not and does not pretend to be.
- **`models/`** — public weights already pinned by repo + immutable
  revision in `models/manifest.lock.yaml`. Recovery is `uv sync --extra
  models && imsg models verify`.
- **`ops/`** — small and irreplaceable, but outside what the spec scopes
  to this job; flagged in the command's own output rather than silently
  added.

These copies share the physical device with the data they copy: they
protect against logical corruption, **not** theft or disk failure. The
command prints that on every run.

**Retention deletes only what it can prove.** A directory under
`backups/` is removed only if it is a real (non-symlink) direct child,
its name matches the exact backup-set pattern, it holds a `MANIFEST.json`
that parses and says `"complete": true` (written last, so its presence
*is* the completeness evidence), and it is not among the newest 14.
Anything else — a half-written set from an interrupted run, a staging
directory, an operator's own file — is counted, reported, and left
alone.

**Attachments whose file is not on this host** can be fetched from
wherever else a copy exists: another Mac's Messages folder, its
attached drives, a NAS share (migration 0008, `attachment_location`).

```bash
# on the index host: find candidate copies (counts only are printed)
uv run imsg locate-attachments --listing other-mac=listing.tsv \
    --catalog catalogs/ --seed-db other-mac=other-mac-chat.db --dry-run
# on the index host: this host's folder, the cache, attachments.pull shares
uv run imsg backfill-attachments
# on the other Mac, which the index host cannot reach: copy what it has
uv run imsg push-attachments --ssh-host index-host \
    --remote-imsg /path/to/imsg --remote-config /path/to/config.yaml \
    --root "other-mac=$HOME/Library/Messages/Attachments" --root "D-XXXXX=/Volumes/Drive"
# on the index host again: verify and materialize what was pushed
uv run imsg backfill-attachments
```

A copy is accepted at a path a chat.db recorded for the attachment, in a
folder named after its GUID under its name, or, flagged, under its name
at its byte size anywhere else; never on a name alone. Each copy is
checked against the size and hash its location reported before it enters
the cache. Every source is only read: both copies are plain rsync runs
whose source side is rsync's sender, and the flags that would make it
delete or remove anything are refused. No command prints a path.

**Attachment text (OCR, PDF and document text, transcripts, captions)**
comes from the S5b queue. `imsg backfill-attachments` queues each
attachment's kinds as it materializes it; `imsg enrich --plan` queues
everything already materialized (migration 0009 first). Routing reads the
file's content, never its name or chat.db's MIME claim.

```bash
uv run imsg migrate                                    # 0009 adds the doc_text kind
uv run imsg enrich --plan --dry-run                    # per-kind counts, unroutable types; writes nothing
uv run imsg enrich --plan                              # fill the queue (insert-only, safe to repeat)
uv run imsg enrich --kinds doc_text,pdf_text,transcript,ocr,frame_ocr --limit 200000
uv run imsg enrich --dry-run                           # what is still claimable, by kind
```

A worker claims one task at a time, cheap kinds first and captions last
unless `--kinds` names kinds (then that list is the order), newest
attachment first within a kind, and stands aside between tasks while a
search is running. Captions are the slow part (about 16 s each on an
M4 Pro, estimated); the nightly `…enrich` agent works through them.

**Model-heavy commands run one at a time on a host.** Each one loads
tens of GiB of model weights, and two together have exhausted a 64 GiB
host's memory. `sync`, `segment`, `embed`, `enrich` (the worker, not
`--plan`), `eval run` and `eval pool` take an exclusive lock on
`<data_root>/run/heavy-models.lock` before their models load, and a
second one waits for the first, logging `heavy_lock.waiting` with the
holder's pid and command. `sync` snapshots and extracts first and takes
the lock only at segmentation. The kernel releases the lock when its
holder exits, including when it is killed, so there is nothing to clean
up. The MCP servers never take it. A dry run that loads no model
(`embed --dry-run`, `enrich --dry-run`, `enrich --plan`) takes no lock;
`segment --dry-run` does, because it still runs the boundary model.

```bash
uv run imsg status | grep heavy_models_lock   # held or not, and by which pid/command
uv run imsg embed --no-wait                   # exit 1 naming the holder, instead of waiting
```

**Nothing loads a model the host has no memory for.** Before any model
set loads — an MCP server at startup or reloading after an idle unload,
`segment`, `embed`, the `enrich` worker, `sync`'s segmentation and
embedding, `eval run`/`eval pool` — it asks `imsg.memory_admission`: is
the kernel's memory pressure normal, and does available memory (`vm_stat`
free + inactive + speculative, less what other processes were admitted
for and have not loaded yet) cover the role's expected footprint
(`memory.footprints`, measured) plus `memory.reserve_bytes` (8 GiB) for
the rest of the host? A host that cannot be measured is a no. Background
commands wait for a yes, re-checking every 30 s for up to 600 s, then
exit **75** (`deferred: memory`) with the heavy lock released. An MCP
server loads nothing and answers retrieval calls with the retryable
`WARMING_UP` code and a "host memory busy" message; the local server tries
again on a later call, the public one by itself every
`memory.admission_retry_seconds` (15 s). Two processes admitted in the
same second cannot both claim the same free memory: admission runs under
a short host-wide lock and leaves a reservation in
`<data_root>/run/memory-reservations/`.

**MCP servers load before background work.** An MCP server whose load is
refused posts a notice, `<data_root>/run/live-servers-waiting/<pid>.json`,
held under a lock for as long as it waits (the kernel drops the lock if
the server dies). While one is posted, no background command starts a
model load, and a running one stops after its current unit of work (a
chat, a batch, a task) and exits **75**, dropping its models and its
reservation, so the server loads at its next try. What background jobs
were admitted for and have not loaded yet counts against the server only
for `memory.background_yield_seconds` (60 s); after that, only the memory
the host really has, the reserve and the pressure level count, so a job
that cannot stop soon cannot hold the server off. Other MCP servers'
reservations always count.

```bash
uv run imsg status | grep -A3 -E 'live_servers_waiting|mcp_public_waiting_on'   # who waits, on what
uv run imsg background status                                                  # what background work gives way to
```

**Heavy background work can be paused.** `imsg background pause`
stops segment, embed, the enrich worker, backfill-attachments and `sync`'s
segmentation and embedding: a command started while paused exits **76**
(`deferred: paused`) without loading anything, and a running one stops
after the unit it is on (a task, a batch, a chat, a file). `sync` keeps
doing its light work — snapshot, extract, identity — so no message is
lost; it says it skipped segmentation and embedding and exits 76. The MCP
servers are never paused. Another project on the host can pause without
this project's config or encrypted volume by creating the host pause file
(`background.host_pause_file`, default
`~/.config/imessage-index/pause-background`):

```bash
uv run imsg background pause --reason "photo import" --until 6h   # or an ISO time; --until is optional
uv run imsg background resume
uv run imsg background status

# from another project on the same host (same user), no imsg needed:
mkdir -p ~/.config/imessage-index
printf 'reason=photo import\npid=%d\n' $$ > ~/.config/imessage-index/pause-background   # paused while this shell runs
rm -f ~/.config/imessage-index/pause-background                                              # resume
```

A `pid=` line makes the pause end by itself when that process does, so a
crashed import cannot leave the index paused; an `until=<ISO time>` line
ends it at that time. Without either, it lasts until the file is removed.

**Running work stops when memory runs short.** Between units of work every
heavy background command also reads the kernel's memory-pressure level:
at critical it stops at once, at warn it stops if warn is still there 10 s
later (`memory.background_stop_at`, `memory.warn_confirm_seconds`), exiting
75. Every MCP server's watchdog reads the same level every 5 s: at critical
a server with its models loaded unloads them at once, without waiting for
its idle timer, as soon as no call is in flight
(`memory.local_server_release_at`; `memory.public_server_release_at` can be
`never`). The public server loads again by itself after
`memory.public_rewarm_cooldown_seconds` (300 s), if admitted. Each role
also runs under its own MLX memory limit (`memory.mlx_memory_limits`, set
with `mx.set_memory_limit`): a guideline MLX uses to release cached
buffers and pace evaluation, never to refuse an allocation, set at 1.25x
each role's measured MLX peak so it changes neither results nor latency.
`imsg status` shows available memory and the pressure level, the pause
state and its reason, and every running model process with its footprint
(read without root) and what it was admitted for.

**Heavy work waits while a search is running.** The enrich worker waits
between tasks, segmentation before each boundary-model call and embedding
before each batch (in `sync` and on their own) while any search is in
flight: a public or local MCP search, or the search page's embedding or
rerank. The servers hold a Postgres advisory lock while they answer, and a
background step checks it, which costs two short statements when nobody
is searching. A wait lasts at most `enrichment.yield_max_pause_seconds`
(300 s), then the unit runs anyway; for segmentation and embedding, a
pause or a live server waiting for memory ends it at once.
`enrichment.yield_to_queries: false` turns it off everywhere.
`imsg status` shows `query_in_flight` and which step is waiting
(`enrichment_yielding_now`, `segment_yielding_now`, `embed_yielding_now`),
and each command's last lines say how often and how long it waited. It
does not preempt: a search that starts during a batch shares the GPU with
that batch until the batch ends.

**Before exposing the public surface**, AT-1 must pass:

```bash
security add-generic-password -a "$USER" -s imsgindex-at1-owner -w      # prompts; no shell history
security add-generic-password -a "$USER" -s imsgindex-at1-nonowner -w
uv run imsg mcp public --probe \
    --owner-token-ref keychain:imsgindex-at1-owner \
    --foreign-token-ref keychain:imsgindex-at1-nonowner
```

Tokens are passed as `keychain:` / `env:` **references**, never values —
a token in `argv` is readable by every process on the host via `ps -ww`
and is recorded in shell history. Exit codes are the verdict: `0` pass,
`1` fail (a breach), `2` invalid (*proved nothing* — which is not a
pass), `78` the probe never ran because a precondition was missing. An
invalid result is treated exactly like a failure: scope stays
`allowlist`.

---

## Requirements

- **macOS on Apple Silicon.** The pipeline depends on macOS-only APIs:
  Full Disk Access to Messages, the Contacts framework, Apple Vision.
- **Python 3.12+** and [`uv`](https://docs.astral.sh/uv/).
- **Rust toolchain**, to build the extraction shim.
- **≥ 32 GB unified memory** for 8B-class local models; smaller models
  work with reduced quality.
- **An encrypted volume to hold all derived state.** Either your Mac's
  FileVault-encrypted startup disk, or a separate encrypted APFS volume
  — either is accepted as `data_root`, and the mount gate refuses to run
  against an unencrypted or unmounted location
  (`src/imsg/mount/guard.py`). Whichever you use, it must contain
  a sentinel file named `.imsgindex-volume` at the root of `data_root` so
  the gate can confirm it is the intended volume and not, say, an
  unmounted mount point silently resolving to the boot disk underneath
  it (`src/imsg/mount/guard.py`).
- **PostgreSQL 17 + pgvector, as a dedicated instance on port 5433**,
  with its data directory under `$DATA_ROOT/pg17` (`src/imsg/config/
  schema.py`, `src/imsg/db/fingerprint.py`). A generic
  Postgres install on the default port will not work — config validation
  rejects any other port, and a two-sided fingerprint check refuses to
  treat any other data directory as this project's own instance. The
  `pgvector` and `pg_prewarm` extensions must both be installed into that
  instance; `pg_prewarm` fills PostgreSQL's shared buffer cache with the
  search index at startup so queries don't wait on disk (migration
  `0004_pg_prewarm.sql`). [`scripts/bootstrap_local_postgres.sh`](scripts/bootstrap_local_postgres.sh)
  sets up a cluster meeting all of this for you.
- **A one-time reranker conversion step.** The pinned reranker model is
  not downloaded ready-to-use — it's converted locally from a Hugging
  Face checkpoint using the exact command recorded in
  [`models/manifest.lock.yaml`](models/manifest.lock.yaml) (see the
  `qwen3-reranker-0.6b-bf16` entry). This requires the `models` extra
  installed first. [`docs/install-macos.md`](docs/install-macos.md) has
  this command spelled out step by step.

## Getting started

```bash
git clone https://github.com/msg43/imessage_mcp.git imessage-index
cd imessage-index && uv sync --extra models --extra dev
```
The `models` extra installs the real model runtimes (mlx, mlx-lm,
mlx-whisper, mlx-vlm, torch, open_clip, pyobjc's Vision bridge,
pillow-heif); `uv sync --extra dev` alone is enough for the `fake`
backend and the test suite — tests that need `models` or the `export`
extra (below) skip cleanly without them.
```bash
cargo build --release --manifest-path tools/imsg-dump/Cargo.toml
```
```bash
uv run pytest        # 1803 passed, 503 skipped with `--extra dev` alone and no database;
                     # 1847 passed, 477 skipped with every extra installed and still no database — 2026-09-24
```

Copy `config.example.yaml`, fill it in, and point the CLI at it:

```bash
export IMSG_CONFIG=/path/to/your/config.yaml
```
```bash
uv run imsg check-permissions && uv run imsg migrate
```

Secrets are never stored in config — they resolve from the macOS
Keychain (`keychain:<item>`), the environment (`env:<VAR>`), or a file
only you can read (`file:/absolute/path`, mode 0600; the usual choice on
a headless host, where the Keychain is unreadable over SSH). Config
validation rejects anything that looks like a literal secret.

> **Two macOS gotchas that will each cost you an hour.**
> PostgreSQL needs `export LC_ALL=C` or the postmaster dies at startup
> with *"postmaster became multithreaded during startup"* — which reads
> like a corrupt installation and is not. And **Full Disk Access cannot
> be granted over SSH**: TCC prompts require a GUI session, and the
> grant goes to the binary that *launches* the job, not to Messages.

Then `uv run imsg --help`. Every stage supports `--dry-run`.

## Connect to Claude

Once the pipeline is set up and indexed, point an AI assistant at the
local MCP server so it can search your messages: `uv run imsg mcp local`
speaks the MCP protocol over stdio. Full step-by-step instructions for
registering it with both **Claude Desktop** and **Claude Code** —
including example config — are in
[`docs/install-macos.md`](docs/install-macos.md) and
[`examples/`](examples/).

## Replacing the model providers

This is the work between "tests pass" and "it does something." The
real implementations exist; what is missing is any run of them against
the pinned weights. `imsg.providers.factory` is the only place providers
are constructed. `models.backend` in `config.yaml` selects `real` (the
default — MLX text embedding, reranker and boundary LLM; Apple Vision
OCR; mlx-whisper transcription; mlx-vlm captioning; PE-Core multimodal
embedding) or `fake` (the deterministic stand-ins, explicit opt-in);
every command that builds providers prints `models: backend=<...>`
first. The factory imports the real classes lazily, by dotted path, and
builds them from the repo ids and immutable revisions in `config.yaml`,
whose defaults mirror `models/manifest.lock.yaml` — repo, commit sha,
license, expected dimension, quantization, runtime floors and a
smoke-test record per model. `uv run imsg models verify` (also
`scripts/verify_model_manifest.py`) re-resolves each repo against the
Hugging Face API and checks the installed packages against the floors;
it reports drift and never rewrites the lock unless given `--write`.
The runtime packages live behind the `models` extra; a missing package
fails as one clear `imsg: ...` line, not a traceback. The two fixed
prompts (`prompts/segment_boundaries.txt`, `prompts/caption.txt`) ship
in the repo and are used unless `paths.data_root` holds a copy at the
same relative path; each run prints which file it used, because those
bytes are hashed into `seg_config_hash` and the caption provenance.

**What has been run, and what is still unverified (2026-09-15).**
`uv run python scripts/smoke_test_models.py` (also
`imsg.providers.model_smoke`) downloads every pinned model at its sha,
checksums it (`artifact_sha256`, defined in the lock header), builds the
real provider through the factory and runs one fictional input per role,
recording load time, inference time and peak memory per model into the
lock with `--write`. On an M2 Ultra with 128 GB every entry passes. Still
unverified: `segment` / `embed` / `enrich` on a real chat, batched
throughput and memory (the recorded peaks are single-input), the
deployment host's memory, and retrieval quality with the pinned weights —
establishing that is the first real task.

**The reranker is a local conversion, not a Hub download.** The only
8-bit MLX conversion of Qwen3-Reranker-8B on the Hub ships no `lm_head`
tensor, so the model card's yes/no logits cannot be computed from it (the
2026-09-14 smoke run found this; the evidence is kept in the lock entry's
notes). The lock therefore pins the reranker as `source: local_conversion`:
the upstream repo and commit sha, the license, the converter
(`mlx-lm==0.31.3`), the exact `command` that produces it, its
`output_dir` relative to `paths.data_root`, and the `artifact_sha256` of
that directory. The pinned reranker is **Qwen3-Reranker-0.6B** since
2026-09-17, for the latency budget below, and since 2026-09-25 its
unquantized 16-bit (bf16) conversion (1.12 GiB on disk), which runs faster
on the GPU than the 8-bit mxfp8 conversion before it. Both earlier
conversions, the 0.6B's mxfp8 build and the 8B, stay in the lock as
`status: retained` — pinned and verified exactly like an active entry,
claiming no role; setting `retrieval.reranker_model` to the mxfp8 build's
directory switches back to it. `imsg models verify` prints which entry is
active for every role and lists the retained ones separately. To
reproduce a conversion, run the recorded command with `$DATA_ROOT` set to
your data root and the `models` extra installed (a few seconds for the
0.6B, about 16 for the 8B), then
`scripts/smoke_test_models.py --only qwen3-reranker-0.6b-bf16 --data-root
$DATA_ROOT` to confirm the digest — two runs of the bf16 recipe on
2026-09-25 produced the same artifact byte for byte. In `config.yaml`,
`retrieval.reranker_model` names that directory (data-root-relative) and
`retrieval.reranker_revision` the *upstream* commit; the factory reads
the value as a local directory when it exists under the data root and as
a Hugging Face repo id otherwise, and the provider records `model_id` as
`<dir>@<upstream sha>`. `uv run imsg models verify --data-root
$DATA_ROOT` re-resolves the upstream repo for drift, recomputes the
directory's digest against the lock, and checks the runtimes; `--write`
never advances an upstream pin (the directory was converted from the
pinned commit — re-convert and re-pin by hand to move it).

**Search latency: the reranker's size, and whether the index is in
memory.** Measured end to end through the real MCP surface on an M2 Ultra
on 2026-09-17 — `imsg mcp local` under a stdio client, 20 fictional
queries, three cold starts, with another job using the same disk array at
216-2,177 MB/s throughout — a whole `search_messages` round trip took
**p50 0.87 s, p95 1.14 s, max 1.28 s**. Where it goes, p95 per stage:
reranking 0.67 s, the full-text channel 0.24 s, the two vector channels
0.05 and 0.03 s, the query embedding 0.06 s, the multimodal text tower
0.02 s, summary fetch 0.002 s, the audit row 0.07 s.

Three settings decide most of that, and all three were chosen by
measurement rather than taste:

- **Which reranker.** Qwen3-Reranker-0.6B, at `retrieval.rerank_top` 20
  and `retrieval.rerank_doc_max_tokens` 256. The 8B conversion pinned
  before it reads roughly 650-780 tokens a second here: scoring 50
  uncapped candidates meant ~20,000 tokens and a p95 of 36 s, and its
  fastest setting that fit a budget (10 candidates, 64 tokens, p95 1.97 s
  of reranking) agreed with its own full ranking no better than doing no
  reranking at all. How many candidates are reranked matters more than
  the reranker's size: with 10, the returned top 10 *is* the fused top 10
  in another order, so reranking cannot lift anything from further down.
  Since 2026-09-25 the 0.6B runs as its bf16 build, with up to 8,192
  padded tokens per forward pass (`retrieval.rerank_max_batch_tokens`)
  and the prompt text every candidate shares read once per query
  (`retrieval.rerank_reuse_prefix`). On the same machine that took
  reranking 20 candidates of 40-700 tokens, cut to 256, from p95 0.674 s
  to 0.397 s (`scripts/bench_query_stages.py`, 2026-09-25); the figures
  above predate it.
- **How much of the HNSW index each search reads.**
  `retrieval.hnsw_ef_search`, applied with `SET LOCAL` in each channel's
  own transaction, defaults to 1000 — pgvector's maximum. On the live
  index, recall@100 against exact search rises 0.944 -> 1.000 (text
  channel) and 0.792 -> 0.998 (image channel) from pgvector's default of
  40 to 1000, and the worst single query rises from 0.79 and 0.38 to 1.00
  and 0.98, for about 50 ms more per query.
- **Whether the pages are cached.** See "Things that will bite you" below:
  `shared_buffers` holds the search working set, `pg_prewarm` fills it at
  startup, and `imsg status` says so.

`scripts/bench_retrieval_latency.py` sweeps the reranker settings against
the real index and reports a quality *proxy* — agreement with scoring 50
uncapped candidates, not an evaluation. Re-run it before moving either
setting, and let the eval harness settle what latency costs.

**Startup.** `imsg mcp local` answers the MCP handshake in about 1 s
(Claude Code gives up on a server that has not answered within
`MCP_TIMEOUT`, 30 s by default) and, by default, loads no model until the
first retrieval call (`mcp.local.warm_at_start: false`): every client
session starts its own server, most never search, and a server that
loaded at start held a full model set regardless. The warm-up then runs in
the background, logging each step's time on stderr: text embedder
4.8-12.5 s, PE-Core text tower 33.8-50.2 s, reranker 2.7-3.2 s, database
buffer pool 0.1-14.3 s — 41.8-65.4 s in all, against 121 s before the 0.6B
was pinned. A retrieval tool call that arrives during warm-up (the first
one starts it) waits up to 90 s for it, then returns `WARMING_UP` with an
estimate of the seconds remaining; `mcp.local.warm_at_start: true` loads
at start instead, as before. `imsg mcp public` always loads at start; a model
that fails to load makes every such call return `WARM_UP_FAILED` with the
cause. `check_permissions` is the exception: it is diagnostics, so it
answers straight away and carries the warm-up's own state (which step is
loading, how many are done, the estimate, the failure) — the way to find
out what a server that is not answering searches is doing. The first
query after warm-up still costs a little more than the rest (0.94-4.08 s
against a 0.87 s median), and the extra is the query embedder's first
forward pass at a real shape: 0.13-3.07 s against 0.055 s afterwards.

**Idle unload.** Every client session runs its own `imsg mcp local`, and
each process holds its own copy of the models (text embedder, PE-Core text
tower, reranker, plus MLX's buffer cache) once it has loaded them. Idle
sessions used to keep that memory until they exited; five of them
exhausted a 64 GiB index host. Now a server loads only on its first
retrieval call (above), and drops its models after
`mcp.local.idle_unload_seconds` (default 600; 0 = never) with no retrieval
tool call. It clears MLX's and torch's caches
so the memory goes back to the system, and never unloads while a call is
in flight. The next retrieval call reloads the models through the same
warm-up and the same 90 s wait. A reload from a warm OS file cache took
1.2 s for the 8B embedder and 0.6B reranker on an M2 Ultra (9.4 GB of
process footprint down to 0.35 GB after unload, 2026-09-24). A reload from
disk costs about what a cold start does, and past 90 s the call answers
`WARMING_UP` for the client to retry. `check_permissions` reports the state
as `unloaded`. `mcp.public.idle_unload_seconds` defaults to 0: the public
surface is a single process, and its 20 s wait is shorter than a reload.
A server whose stdin closes, which is what happens when the SSH client
disconnects, exits at once, even mid-warm-up. A client that vanishes
without closing the connection (a machine that sleeps) is noticed only
when sshd's `ClientAliveInterval` gives up on it.

| Interface | What it needs |
|---|---|
| `TextEmbeddingProvider` | `embed_documents()` (bare) and `embed_query()` (instruction-prefixed); **2048-dim**, L2-normalized |
| `MultimodalEmbeddingProvider` | `embed_images()` and `embed_text()` (paired towers); **1280-dim** |
| `BoundaryProvider` | Topical boundary indices for a window of messages |
| `OcrProvider` / `CaptionProvider` / `TranscriptionProvider` | One method each |

The reference design uses local MLX-hosted models throughout, on the
reasoning that sending a decade of personal messages to a hosted API is
a categorically different decision from indexing them on your own
machine — and that as of mid-2026 the leading open text-embedding models
top the benchmarks anyway, so there is little quality left to trade for
it. Nothing in the code requires that choice; the interfaces are
provider-agnostic.

⚠️ **The dimensions are load-bearing** — see the first item below.

---

## Things that will bite you

Learned the expensive way; written down so you don't have to.

- **pgvector's index caps are lower than its type limits.** The
  `vector` and `halfvec` types accept up to 16,000 dimensions, but
  **HNSW/IVFFlat indexes cap at 2,000 (`vector`) and 4,000
  (`halfvec`)**. A column can be perfectly legal DDL whose index can
  never be created — the error surfaces at `CREATE INDEX`, and ignoring
  it means silently falling back to sequential scan.
  `scripts/lint_ddl.py` exists solely to catch this. An earlier
  revision of this project specified an unbuildable `halfvec(4096)` for
  exactly this reason.
- **Audience validation and subject validation are not redundant.** On
  the public surface the subject check answers *"is this the owner?"*
  and the audience check answers *"was this token minted for this
  system?"* A user's OAuth subject is identical across every app they
  sign into, so subject-checking alone does not stop a token minted for
  another application being replayed here. Both, or neither works.
- **Applied migrations are immutable**, enforced by hash. Correct a
  mistake in a *later* migration; never edit a shipped one.
- **`updated_at` is enforced by trigger, not convention.**
  Re-segmentation keys off it, so a writer that forgot to bump it would
  strand chats out of reprocessing — no error, stale results the only
  symptom.
- **Filters must overfetch, not post-filter.** Post-filtering a fixed
  top-K silently starves results when filters are selective: you get
  few or zero hits and it looks like "nothing matched."
- **Ingest-time and query-time text normalization must match exactly.**
  If they drift, exact-phrase search silently stops working.
- **Vector search is only fast while its index is in memory.** An HNSW
  query touches a few thousand pages of an index far larger than the
  default 128 MB `shared_buffers`. Measured on the external encrypted
  volume (2026-09-16): queries whose pages the OS had not cached took p95
  302 ms (text vectors) and 609 ms (image vectors), against 37 ms and
  70 ms once the index files were in the page cache — and with another
  job's heavy I/O on the same disk, several seconds each. The fix is in
  three parts: `shared_buffers` sized to hold what search reads (measured
  with `pg_statio_*` deltas over the benchmark queries — 2,296 MiB here,
  of which 1,193 MiB is the three HNSW indexes and 857 MiB the embedding
  tables' TOAST), `pg_prewarm` (migration 0004) filling it — at every
  server start via `RetrievalService.warm_up()`, and after a reboot via
  `pg_prewarm.autoprewarm` — and `imsg status`, which prints
  `shared_buffers` against the total HNSW index size and warns when the
  pool is the smaller of the two.
- **Two providers pinned to the same checkpoint still load it twice
  unless something makes them share.** Captioning goes through `mlx-vlm`
  and topical boundary detection through `mlx-lm`; both name the same
  35B repo and revision, and each used to build its own model object.
  Measured: MLX active memory 18.99 → 37.15 GiB as the second one
  loaded, in one process, for one set of weights. The fix is
  `imsg.shared_vlm_runtime.SharedVlmRuntime`, an object a caller passes
  to both providers — boundary detection then runs text-only through the
  already-loaded vision-language model, which is the *same* computation
  (the rendered chat prompt is byte-identical and the last-position logit
  row is bit-identical across all 248,320 float32 values). It is
  explicit rather than a module-level cache precisely so the sharing is
  visible at the call site; `models.share_boundary_and_caption_weights`
  turns it off.
- **A CLIP-style model loads both towers whether or not you use both.**
  PE-Core is 9.01 GiB at fp32, of which the vision tower is 7.01 and the
  text tower 2.00 — and the query server only ever calls `embed_text`
  while the embedding pipeline only ever calls `embed_images`. The
  provider now decides on first use and releases the other tower;
  `scripts/verify_pe_core_tower_selection.py` proves the vectors are
  identical bit for bit before and after, because "it cannot change the
  answer" is an argument, not a measurement. A process that does use both
  rebuilds once and keeps both.
- **MLX's buffer cache defaults to its memory limit, which is not a
  bound.** It pools freed GPU buffers rather than returning them, and on
  a 64 GB host the default limit probes at 60.8 GiB; one enrichment
  process was seen holding 37.25 GiB of freed buffers, which is
  indistinguishable from a leak and hides real regressions. Every MLX
  provider now calls `imsg.mlx_runtime.bound_buffer_cache` at load
  (`models.query_cache_limit_bytes` / `models.enrichment_cache_limit_bytes`).
  The call is process-wide, so one provider bounding it covers the rest
  — including `mlx_whisper`, which has no load hook of its own.
- **Two GPU-heavy jobs on one machine need an arbiter, and it has to
  survive a crash.** The nightly enrichment window overlaps the
  always-on MCP server; with the duplicate weights gone and no swap at
  all, query p95 still trebled purely from GPU contention. Enrichment
  now stands aside for in-flight queries
  (`imsg.db.enrichment_yield_locks`, `enrichment.yield_to_queries`), and
  the signal is a **Postgres session-level advisory lock** specifically
  because the server releases it when the session ends, however it ends
  — a killed MCP server cannot leave enrichment paused, with no timeout
  to tune and no stale marker to reap. It is checked between units of
  work, never during one, so no claimed task is ever abandoned;
  `imsg status` reports whether it is yielding right now. **What it
  cannot do** is help when the unit of work is long relative to the
  query rate: measured with it on and off, search p95 during the window
  was 3.00 s versus 3.06 s — one caption takes 14 s on that host and a
  search arrives every 2 s, so a worker that resumes between tasks
  resumes straight into another 14-second caption. The memory fixes
  above are what closed the gap; this is kept because it costs one round
  trip when nobody is searching and it will matter wherever the batch is
  finer-grained than the traffic — not because it earned its place on
  this workload. Segmentation and embedding use the same gate since
  2026-09-26 (`imsg.search_yield`).
- **The planner can abandon an HNSW index at some `ef_search` values, and
  that is not a monotonic effect.** pgvector's own cost estimate bounds
  layer-0 tuples by `ef_search` while its selectivity term carries
  `log(ef_search)` in a denominator, so the estimated cost climbs and
  then drops back: on this index it was 3,252 at `ef_search` 40, 13,287
  at 280 and 3,719 at 285, against a flat 8,827 for the sequential
  alternative — which the planner undercosts anyway, because (pgvector's
  own FAQ) it "doesn't consider out-of-line storage in cost estimates"
  and the vectors it would sort are 642 MiB of TOAST. Between 170 and 284
  the same query took 245 ms instead of 9 ms, exactly and only because of
  the plan. Every vector channel therefore sets `enable_seqscan = off`
  for its own transaction, pgvector's documented remedy.

---

## Layout

```
src/imsg/
  config/      config surface + validation (enforces the safety rules)
  db/          connection, migrations, cluster fingerprint
  stages/      snapshot, extract, identity, sync
  segment/  backfill/  enrich/  embed/    indexing pipeline
  retrieval/   hybrid query flow, RRF fusion, reranking
  mcp/         auth boundary, local + public surfaces, tools
  export/      default-deny eligibility, plan/approve/push
  backup/      nightly pg_dump + FTS copy, verification, retention
  eval/        metrics, runner, diff
  verify/      seed completeness, attachment reconciliation
migrations/         schema, applied in order by a hash-checked runner
tools/imsg-dump/    GPL-3.0 Rust extraction shim (subprocess only)
```

## Development

```bash
uv run ruff check . && uv run mypy . && uv run pytest
```

Integration tests run against a live PostgreSQL when one is reachable
and skip cleanly when it isn't; the unit suite never needs a database.

## Licensing

The core is **MIT** — see [`LICENSE`](LICENSE). Component licensing and
the reasoning behind the split are in [`NOTICE`](NOTICE).

In short: `tools/imsg-dump/` is **GPL-3.0** and carries its own
`LICENSE`. It links the GPL `imessage-database` crate to parse the
`attributedBody` typedstream format, which is not optional — since Big
Sur much of a message's text is not in the `text` column at all, and
readers that only query that column silently return empty strings for
large portions of modern history.

**It is invoked strictly across a process boundary** — spawned as a
subprocess, never linked into the Python code. That boundary is
deliberate and load-bearing for the licensing split. Vendor it
differently and that is yours to reason about.

## What isn't here

Instance configuration, by design: real config values, contact seed
data, allowlists and eval queries live in a separate private overlay
you supply and point at with `IMSG_CONFIG`, and `config.example.yaml`
ships placeholders only. **This repo's git history was rewritten on
2026-09-24 to remove real contact data that had been committed as test
fixtures.** If you find a real name, number, address, or other personal
identifier anywhere in this repo — in code, history, or an issue — do
not open a public issue about it; report it privately as described in
[`SECURITY.md`](SECURITY.md).

The design record — architecture rationale, full build spec, and the
decision log explaining *why* each choice above was made — is kept
private, since it's written against a specific deployment.