Skip to main content
Glama
msg43

imessage-index

by msg43

imessage-index

A local-first iMessage retrieval system. Extracts your full message history, resolves identities, segments conversations topically, embeds and indexes them, and serves hybrid full-text + vector search over an MCP server — so an AI assistant can actually answer questions about your own correspondence.

Designed to run unattended on a headless Mac, with all derived state on an encrypted volume and nothing leaving the machine unless you explicitly allowlist it.


Is this for you?

What it does. It reads your Mac's Messages history, figures out who each contact really is across their various numbers and emails, groups messages into topical conversations, and builds a local search index — full-text plus vector — that an AI assistant (Claude Desktop or Claude Code, over MCP) can query. Everything stays on your machine unless you explicitly export a filtered subset.

Who it's for right now. Someone comfortable running commands in a terminal, installing developer tools (Python, Rust, PostgreSQL), and troubleshooting when a step doesn't go cleanly — not a turnkey app for a non-technical user yet. The step-by-step guide in docs/install-macos.md walks through the whole setup and the scripts/doctor.py diagnostic checks your machine before you start, but real command-line comfort still helps.

Hardware. An Apple Silicon Mac (M1 or later) — the pipeline depends on macOS-only APIs (Full Disk Access to Messages, the Contacts framework, Apple Vision) and on Apple's MLX framework for local model inference, which does not run on Intel Macs. RAM: the README's own requirement below is 32 GB unified memory for the full 8B-class local models; smaller models work with less, but no specific lower figure has been measured [UNVERIFIED].

Time and disk. Not measured end to end for a first-time setup [UNVERIFIED] — expect at least an hour between installing dependencies, downloading pinned models, and running the reranker conversion step (see Requirements below). Disk usage depends heavily on the size of your own Messages history and attachments; no general figure has been measured [UNVERIFIED].

Before you read further, the section immediately below is important: this project has never been run end to end against real models, so treat it as a rigorous, well-tested foundation rather than a finished product.


Related MCP server: imessage-rich-search

⚠️ Read this before you clone

No model has run on the pipeline end to end yet. With a plain uv sync --extra dev and no database, the test suite is 1,803 passed and 503 skipped (most skips need a scratch PostgreSQL; the rest need the models and/or export extras, which are not required for a plain install). With every extra installed (--extra dev --extra models --extra export) and still no database, it's 1,847 passed and 477 skipped, almost all of those needing PostgreSQL — measured 2026-09-24. The CLI works, migrations apply against real PostgreSQL + pgvector, and the snapshot → extract → identity stages have run against a real corpus. Segmentation, embedding, enrichment and retrieval have only ever run on the fake providers or on tiny stand-in models (see what is still unverified).

It ships deterministic fake model providers alongside the real wiring. The embedding, reranking, captioning, OCR, transcription, and segmentation providers each have a correctly-dimensioned stub that lets the pipeline run end to end in tests. imsg.providers.factory selects between the real implementations and those stubs with one config field, models.backend (real is the default; fake is explicit opt-in), and every command that builds providers prints models: backend=<real|fake> first. Run the pipeline with fake and every stage will report success while your search results are meaningless. The real implementations and their pins are described in Replacing the model providers.

Treat this as a thoroughly tested skeleton with a complete design behind it, not a working product. If you want something that works today, this isn't it. If you want a rigorous starting point that has already made and documented the non-obvious decisions, read on.

What is genuinely done

Area

State

Schema + migrations

Complete; applied against live PostgreSQL 17 + pgvector

Pipeline stages (snapshot → extract → identity → segment → enrich → embed → sync → export)

Implemented; unit + integration tested

Hybrid retrieval (BM25 + text vector + multimodal vector, RRF-fused, reranked)

Implemented

Local MCP surface (stdio)

Implemented

Public MCP surface (StreamableHTTP + OAuth)

Implemented; never exposed

Export gate (default-deny, plan/approve/push/purge)

Implemented and wired to the CLI; the GCS + Discovery Engine transport has never run against a live API

Eval harness (nDCG@k, recall@k, MRR)

Implemented; metrics verified against hand-computed fixtures

Real model providers (MLX text/reranker/boundary LLM, Apple Vision OCR, mlx-whisper, mlx-vlm captions, PE-Core)

Implemented behind models.backend: real (the default); every pinned model smoke-run once on a synthetic input (2026-09-14/15, results in the lock); never run on the pipeline end to end

Nightly backup (imsg backup)

Implemented and wired to the CLI; verified pg_dump + FTS sidecar copy, 14 retained. Scope is deliberately narrow — see below

AT-1 auth probe (imsg mcp public --probe)

Wired to the CLI and fully tested up to the point where live OAuth tokens are required; never run against real tokens or a live gate

Runs against real data

Snapshot → extract → identity: yes. Segment / embed / enrich with the pinned models: never


Architecture in one pass

chat.db ──.backup──▶ snapshot ──▶ extract ──▶ identity resolution
                                                     │
                                              segmentation
                                                     │
                                  ┌──────────────────┴─────────────┐
                             text bodies                     attachments
                                                                   │
                                          ┌────────────────────────┼───────────┐
                                    PDF text layer              images    audio/video
                                          └────────────────────────┼───────────┘
                                                                   │
                                      embed ──▶ pgvector + FTS5 + multimodal vector
                                                                   │
                                          ┌────────────────────────┴──────────┐
                                    local MCP                         filtered export
                                   (full corpus)                       (allowlisted)

Four choices that shaped everything else:

  • Segments, not messages. Retrieval returns topically coherent conversation segments. Message-granular results flood a model's context with "sounds good" and prevent reconstructing what was actually discussed.

  • Everything becomes text first. OCR output, captions, transcripts, and PDF text layers all embed into one text space, so failures stay debuggable — when something doesn't surface you can read the extracted text and see why. A second multimodal vector runs alongside for visual similarity.

  • Identity resolution precedes segmentation. Nothing downstream keys on a raw handle. One person legitimately has many numbers, emails and aliases across a decade.

  • Default deny on export. Nothing reaches an external index unless explicitly allowlisted, and a group thread requires every participant allowlisted.


The export gate

The one path by which message content can leave the machine, so every command on it is shaped to refuse. Default deny: a thread exports only if every participant — and every message and tapback sender, the owner included — is explicitly allowlisted. Attachments are gated separately from text bodies.

uv run imsg export plan                  # eligibility → staged bytes → review report
uv run imsg export approve <run-id>      # pins the exact bytes you reviewed
uv run imsg export push <run-id>         # re-verifies, then promotes
uv run imsg export purge-person <who>    # revocation; exempt from the approval gate
uv run imsg export unclassified-report   # weekly: whose threads are still unclassified

plan writes the review report the whole design rests on — per thread: participants, message count, date range, sample lines. approve pins the manifest hash and every staged file hash. push re-checks all of that and re-derives eligibility from the live database, because the hashes prove the bytes did not change, not that the world did not: a participant added to a group between approval and push leaves every hash green while changing who is in the export. Any drift aborts the push and requires a new plan.

Revocation is deliberately faster than export — it only ever narrows scope — so a purge needs no approval, though every drift check still applies and the run is recorded in full.

plan, push, purge-person and unclassified-report all take --dry-run. push --dry-run runs every verification the real push runs and builds no transport at all, so rehearsing the gate cannot reach the network.

Nothing can reach Google without a credential you named. export.gcp_credentials is a keychain: / env: secret reference with no default; with it unset, push refuses before it opens the database or imports a Google client library. Leave it out until you mean it.

Honest limit. A purge reaches the Discovery Engine index and the GCS bucket. Copies already swept into organizational retention, backups, or another person's hands are beyond it. The gate at export time is the actual protection, which is why it denies by default.


Operations

Seven com.imsgindex.* LaunchAgents are rendered by imsg install-agents, and every command they invoke exists — installing them schedules no job that fails nightly.

uv run imsg backup                       # daily 04:00: verified pg_dump + FTS copy, 14 kept
uv run imsg backup --dry-run             # preconditions + the retention plan, writes nothing
uv run imsg status                       # mount, Postgres, disk, posture, unclassified threads

What imsg backup covers, and what it does not. In scope: the Postgres dump (everything the pipeline derived exists only there, and it is the one component that can suffer logical corruption) and the FTS5 sidecar (rebuildable, copied for recovery speed, under SPEC's checkpoint + integrity-check conditions). Out of scope, on purpose:

  • attachments/ (~147 GB) — 14 nightly copies of it would be over two terabytes on the same volume as the original, it is a content-addressed cache that imsg backfill-attachments rebuilds, and a blob store does not suffer the logical corruption these copies defend against. But: anything iCloud has already purged exists nowhere else, and that genuinely needs an independently encrypted off-box archive, which this command is not and does not pretend to be.

  • models/ — public weights already pinned by repo + immutable revision in models/manifest.lock.yaml. Recovery is uv sync --extra models && imsg models verify.

  • ops/ — small and irreplaceable, but outside what the spec scopes to this job; flagged in the command's own output rather than silently added.

These copies share the physical device with the data they copy: they protect against logical corruption, not theft or disk failure. The command prints that on every run.

Retention deletes only what it can prove. A directory under backups/ is removed only if it is a real (non-symlink) direct child, its name matches the exact backup-set pattern, it holds a MANIFEST.json that parses and says "complete": true (written last, so its presence is the completeness evidence), and it is not among the newest 14. Anything else — a half-written set from an interrupted run, a staging directory, an operator's own file — is counted, reported, and left alone.

Attachments whose file is not on this host can be fetched from wherever else a copy exists: another Mac's Messages folder, its attached drives, a NAS share (migration 0008, attachment_location).

# on the index host: find candidate copies (counts only are printed)
uv run imsg locate-attachments --listing other-mac=listing.tsv \
    --catalog catalogs/ --seed-db other-mac=other-mac-chat.db --dry-run
# on the index host: this host's folder, the cache, attachments.pull shares
uv run imsg backfill-attachments
# on the other Mac, which the index host cannot reach: copy what it has
uv run imsg push-attachments --ssh-host index-host \
    --remote-imsg /path/to/imsg --remote-config /path/to/config.yaml \
    --root "other-mac=$HOME/Library/Messages/Attachments" --root "D-XXXXX=/Volumes/Drive"
# on the index host again: verify and materialize what was pushed
uv run imsg backfill-attachments

A copy is accepted at a path a chat.db recorded for the attachment, in a folder named after its GUID under its name, or, flagged, under its name at its byte size anywhere else; never on a name alone. Each copy is checked against the size and hash its location reported before it enters the cache. Every source is only read: both copies are plain rsync runs whose source side is rsync's sender, and the flags that would make it delete or remove anything are refused. No command prints a path.

Attachment text (OCR, PDF and document text, transcripts, captions) comes from the S5b queue. imsg backfill-attachments queues each attachment's kinds as it materializes it; imsg enrich --plan queues everything already materialized (migration 0009 first). Routing reads the file's content, never its name or chat.db's MIME claim.

uv run imsg migrate                                    # 0009 adds the doc_text kind
uv run imsg enrich --plan --dry-run                    # per-kind counts, unroutable types; writes nothing
uv run imsg enrich --plan                              # fill the queue (insert-only, safe to repeat)
uv run imsg enrich --kinds doc_text,pdf_text,transcript,ocr,frame_ocr --limit 200000
uv run imsg enrich --dry-run                           # what is still claimable, by kind

A worker claims one task at a time, cheap kinds first and captions last unless --kinds names kinds (then that list is the order), newest attachment first within a kind, and stands aside between tasks while a search is running. Captions are the slow part (about 16 s each on an M4 Pro, estimated); the nightly …enrich agent works through them.

Model-heavy commands run one at a time on a host. Each one loads tens of GiB of model weights, and two together have exhausted a 64 GiB host's memory. sync, segment, embed, enrich (the worker, not --plan), eval run and eval pool take an exclusive lock on <data_root>/run/heavy-models.lock before their models load, and a second one waits for the first, logging heavy_lock.waiting with the holder's pid and command. sync snapshots and extracts first and takes the lock only at segmentation. The kernel releases the lock when its holder exits, including when it is killed, so there is nothing to clean up. The MCP servers never take it. A dry run that loads no model (embed --dry-run, enrich --dry-run, enrich --plan) takes no lock; segment --dry-run does, because it still runs the boundary model.

uv run imsg status | grep heavy_models_lock   # held or not, and by which pid/command
uv run imsg embed --no-wait                   # exit 1 naming the holder, instead of waiting

Nothing loads a model the host has no memory for. Before any model set loads — an MCP server at startup or reloading after an idle unload, segment, embed, the enrich worker, sync's segmentation and embedding, eval run/eval pool — it asks imsg.memory_admission: is the kernel's memory pressure normal, and does available memory (vm_stat free + inactive + speculative, less what other processes were admitted for and have not loaded yet) cover the role's expected footprint (memory.footprints, measured) plus memory.reserve_bytes (8 GiB) for the rest of the host? A host that cannot be measured is a no. Background commands wait for a yes, re-checking every 30 s for up to 600 s, then exit 75 (deferred: memory) with the heavy lock released. An MCP server loads nothing and answers retrieval calls with the retryable WARMING_UP code and a "host memory busy" message; the local server tries again on a later call, the public one by itself every memory.admission_retry_seconds (15 s). Two processes admitted in the same second cannot both claim the same free memory: admission runs under a short host-wide lock and leaves a reservation in <data_root>/run/memory-reservations/.

MCP servers load before background work. An MCP server whose load is refused posts a notice, <data_root>/run/live-servers-waiting/<pid>.json, held under a lock for as long as it waits (the kernel drops the lock if the server dies). While one is posted, no background command starts a model load, and a running one stops after its current unit of work (a chat, a batch, a task) and exits 75, dropping its models and its reservation, so the server loads at its next try. What background jobs were admitted for and have not loaded yet counts against the server only for memory.background_yield_seconds (60 s); after that, only the memory the host really has, the reserve and the pressure level count, so a job that cannot stop soon cannot hold the server off. Other MCP servers' reservations always count.

uv run imsg status | grep -A3 -E 'live_servers_waiting|mcp_public_waiting_on'   # who waits, on what
uv run imsg background status                                                  # what background work gives way to

Heavy background work can be paused. imsg background pause stops segment, embed, the enrich worker, backfill-attachments and sync's segmentation and embedding: a command started while paused exits 76 (deferred: paused) without loading anything, and a running one stops after the unit it is on (a task, a batch, a chat, a file). sync keeps doing its light work — snapshot, extract, identity — so no message is lost; it says it skipped segmentation and embedding and exits 76. The MCP servers are never paused. Another project on the host can pause without this project's config or encrypted volume by creating the host pause file (background.host_pause_file, default ~/.config/imessage-index/pause-background):

uv run imsg background pause --reason "photo import" --until 6h   # or an ISO time; --until is optional
uv run imsg background resume
uv run imsg background status

# from another project on the same host (same user), no imsg needed:
mkdir -p ~/.config/imessage-index
printf 'reason=photo import\npid=%d\n' $$ > ~/.config/imessage-index/pause-background   # paused while this shell runs
rm -f ~/.config/imessage-index/pause-background                                              # resume

A pid= line makes the pause end by itself when that process does, so a crashed import cannot leave the index paused; an until=<ISO time> line ends it at that time. Without either, it lasts until the file is removed.

Running work stops when memory runs short. Between units of work every heavy background command also reads the kernel's memory-pressure level: at critical it stops at once, at warn it stops if warn is still there 10 s later (memory.background_stop_at, memory.warn_confirm_seconds), exiting 75. Every MCP server's watchdog reads the same level every 5 s: at critical a server with its models loaded unloads them at once, without waiting for its idle timer, as soon as no call is in flight (memory.local_server_release_at; memory.public_server_release_at can be never). The public server loads again by itself after memory.public_rewarm_cooldown_seconds (300 s), if admitted. Each role also runs under its own MLX memory limit (memory.mlx_memory_limits, set with mx.set_memory_limit): a guideline MLX uses to release cached buffers and pace evaluation, never to refuse an allocation, set at 1.25x each role's measured MLX peak so it changes neither results nor latency. imsg status shows available memory and the pressure level, the pause state and its reason, and every running model process with its footprint (read without root) and what it was admitted for.

Heavy work waits while a search is running. The enrich worker waits between tasks, segmentation before each boundary-model call and embedding before each batch (in sync and on their own) while any search is in flight: a public or local MCP search, or the search page's embedding or rerank. The servers hold a Postgres advisory lock while they answer, and a background step checks it, which costs two short statements when nobody is searching. A wait lasts at most enrichment.yield_max_pause_seconds (300 s), then the unit runs anyway; for segmentation and embedding, a pause or a live server waiting for memory ends it at once. enrichment.yield_to_queries: false turns it off everywhere. imsg status shows query_in_flight and which step is waiting (enrichment_yielding_now, segment_yielding_now, embed_yielding_now), and each command's last lines say how often and how long it waited. It does not preempt: a search that starts during a batch shares the GPU with that batch until the batch ends.

Before exposing the public surface, AT-1 must pass:

security add-generic-password -a "$USER" -s imsgindex-at1-owner -w      # prompts; no shell history
security add-generic-password -a "$USER" -s imsgindex-at1-nonowner -w
uv run imsg mcp public --probe \
    --owner-token-ref keychain:imsgindex-at1-owner \
    --foreign-token-ref keychain:imsgindex-at1-nonowner

Tokens are passed as keychain: / env: references, never values — a token in argv is readable by every process on the host via ps -ww and is recorded in shell history. Exit codes are the verdict: 0 pass, 1 fail (a breach), 2 invalid (proved nothing — which is not a pass), 78 the probe never ran because a precondition was missing. An invalid result is treated exactly like a failure: scope stays allowlist.


Requirements

  • macOS on Apple Silicon. The pipeline depends on macOS-only APIs: Full Disk Access to Messages, the Contacts framework, Apple Vision.

  • Python 3.12+ and uv.

  • Rust toolchain, to build the extraction shim.

  • ≥ 32 GB unified memory for 8B-class local models; smaller models work with reduced quality.

  • An encrypted volume to hold all derived state. Either your Mac's FileVault-encrypted startup disk, or a separate encrypted APFS volume — either is accepted as data_root, and the mount gate refuses to run against an unencrypted or unmounted location (src/imsg/mount/guard.py). Whichever you use, it must contain a sentinel file named .imsgindex-volume at the root of data_root so the gate can confirm it is the intended volume and not, say, an unmounted mount point silently resolving to the boot disk underneath it (src/imsg/mount/guard.py).

  • PostgreSQL 17 + pgvector, as a dedicated instance on port 5433, with its data directory under $DATA_ROOT/pg17 (src/imsg/config/ schema.py, src/imsg/db/fingerprint.py). A generic Postgres install on the default port will not work — config validation rejects any other port, and a two-sided fingerprint check refuses to treat any other data directory as this project's own instance. The pgvector and pg_prewarm extensions must both be installed into that instance; pg_prewarm fills PostgreSQL's shared buffer cache with the search index at startup so queries don't wait on disk (migration 0004_pg_prewarm.sql). scripts/bootstrap_local_postgres.sh sets up a cluster meeting all of this for you.

  • A one-time reranker conversion step. The pinned reranker model is not downloaded ready-to-use — it's converted locally from a Hugging Face checkpoint using the exact command recorded in models/manifest.lock.yaml (see the qwen3-reranker-0.6b-bf16 entry). This requires the models extra installed first. docs/install-macos.md has this command spelled out step by step.

Getting started

git clone https://github.com/msg43/imessage_mcp.git imessage-index
cd imessage-index && uv sync --extra models --extra dev

The models extra installs the real model runtimes (mlx, mlx-lm, mlx-whisper, mlx-vlm, torch, open_clip, pyobjc's Vision bridge, pillow-heif); uv sync --extra dev alone is enough for the fake backend and the test suite — tests that need models or the export extra (below) skip cleanly without them.

cargo build --release --manifest-path tools/imsg-dump/Cargo.toml
uv run pytest        # 1803 passed, 503 skipped with `--extra dev` alone and no database;
                     # 1847 passed, 477 skipped with every extra installed and still no database — 2026-09-24

Copy config.example.yaml, fill it in, and point the CLI at it:

export IMSG_CONFIG=/path/to/your/config.yaml
uv run imsg check-permissions && uv run imsg migrate

Secrets are never stored in config — they resolve from the macOS Keychain (keychain:<item>), the environment (env:<VAR>), or a file only you can read (file:/absolute/path, mode 0600; the usual choice on a headless host, where the Keychain is unreadable over SSH). Config validation rejects anything that looks like a literal secret.

Two macOS gotchas that will each cost you an hour. PostgreSQL needs export LC_ALL=C or the postmaster dies at startup with "postmaster became multithreaded during startup" — which reads like a corrupt installation and is not. And Full Disk Access cannot be granted over SSH: TCC prompts require a GUI session, and the grant goes to the binary that launches the job, not to Messages.

Then uv run imsg --help. Every stage supports --dry-run.

Connect to Claude

Once the pipeline is set up and indexed, point an AI assistant at the local MCP server so it can search your messages: uv run imsg mcp local speaks the MCP protocol over stdio. Full step-by-step instructions for registering it with both Claude Desktop and Claude Code — including example config — are in docs/install-macos.md and examples/.

Replacing the model providers

This is the work between "tests pass" and "it does something." The real implementations exist; what is missing is any run of them against the pinned weights. imsg.providers.factory is the only place providers are constructed. models.backend in config.yaml selects real (the default — MLX text embedding, reranker and boundary LLM; Apple Vision OCR; mlx-whisper transcription; mlx-vlm captioning; PE-Core multimodal embedding) or fake (the deterministic stand-ins, explicit opt-in); every command that builds providers prints models: backend=<...> first. The factory imports the real classes lazily, by dotted path, and builds them from the repo ids and immutable revisions in config.yaml, whose defaults mirror models/manifest.lock.yaml — repo, commit sha, license, expected dimension, quantization, runtime floors and a smoke-test record per model. uv run imsg models verify (also scripts/verify_model_manifest.py) re-resolves each repo against the Hugging Face API and checks the installed packages against the floors; it reports drift and never rewrites the lock unless given --write. The runtime packages live behind the models extra; a missing package fails as one clear imsg: ... line, not a traceback. The two fixed prompts (prompts/segment_boundaries.txt, prompts/caption.txt) ship in the repo and are used unless paths.data_root holds a copy at the same relative path; each run prints which file it used, because those bytes are hashed into seg_config_hash and the caption provenance.

What has been run, and what is still unverified (2026-09-15). uv run python scripts/smoke_test_models.py (also imsg.providers.model_smoke) downloads every pinned model at its sha, checksums it (artifact_sha256, defined in the lock header), builds the real provider through the factory and runs one fictional input per role, recording load time, inference time and peak memory per model into the lock with --write. On an M2 Ultra with 128 GB every entry passes. Still unverified: segment / embed / enrich on a real chat, batched throughput and memory (the recorded peaks are single-input), the deployment host's memory, and retrieval quality with the pinned weights — establishing that is the first real task.

The reranker is a local conversion, not a Hub download. The only 8-bit MLX conversion of Qwen3-Reranker-8B on the Hub ships no lm_head tensor, so the model card's yes/no logits cannot be computed from it (the 2026-09-14 smoke run found this; the evidence is kept in the lock entry's notes). The lock therefore pins the reranker as source: local_conversion: the upstream repo and commit sha, the license, the converter (mlx-lm==0.31.3), the exact command that produces it, its output_dir relative to paths.data_root, and the artifact_sha256 of that directory. The pinned reranker is Qwen3-Reranker-0.6B since 2026-09-17, for the latency budget below, and since 2026-09-25 its unquantized 16-bit (bf16) conversion (1.12 GiB on disk), which runs faster on the GPU than the 8-bit mxfp8 conversion before it. Both earlier conversions, the 0.6B's mxfp8 build and the 8B, stay in the lock as status: retained — pinned and verified exactly like an active entry, claiming no role; setting retrieval.reranker_model to the mxfp8 build's directory switches back to it. imsg models verify prints which entry is active for every role and lists the retained ones separately. To reproduce a conversion, run the recorded command with $DATA_ROOT set to your data root and the models extra installed (a few seconds for the 0.6B, about 16 for the 8B), then scripts/smoke_test_models.py --only qwen3-reranker-0.6b-bf16 --data-root $DATA_ROOT to confirm the digest — two runs of the bf16 recipe on 2026-09-25 produced the same artifact byte for byte. In config.yaml, retrieval.reranker_model names that directory (data-root-relative) and retrieval.reranker_revision the upstream commit; the factory reads the value as a local directory when it exists under the data root and as a Hugging Face repo id otherwise, and the provider records model_id as <dir>@<upstream sha>. uv run imsg models verify --data-root $DATA_ROOT re-resolves the upstream repo for drift, recomputes the directory's digest against the lock, and checks the runtimes; --write never advances an upstream pin (the directory was converted from the pinned commit — re-convert and re-pin by hand to move it).

Search latency: the reranker's size, and whether the index is in memory. Measured end to end through the real MCP surface on an M2 Ultra on 2026-09-17 — imsg mcp local under a stdio client, 20 fictional queries, three cold starts, with another job using the same disk array at 216-2,177 MB/s throughout — a whole search_messages round trip took p50 0.87 s, p95 1.14 s, max 1.28 s. Where it goes, p95 per stage: reranking 0.67 s, the full-text channel 0.24 s, the two vector channels 0.05 and 0.03 s, the query embedding 0.06 s, the multimodal text tower 0.02 s, summary fetch 0.002 s, the audit row 0.07 s.

Three settings decide most of that, and all three were chosen by measurement rather than taste:

  • Which reranker. Qwen3-Reranker-0.6B, at retrieval.rerank_top 20 and retrieval.rerank_doc_max_tokens 256. The 8B conversion pinned before it reads roughly 650-780 tokens a second here: scoring 50 uncapped candidates meant ~20,000 tokens and a p95 of 36 s, and its fastest setting that fit a budget (10 candidates, 64 tokens, p95 1.97 s of reranking) agreed with its own full ranking no better than doing no reranking at all. How many candidates are reranked matters more than the reranker's size: with 10, the returned top 10 is the fused top 10 in another order, so reranking cannot lift anything from further down. Since 2026-09-25 the 0.6B runs as its bf16 build, with up to 8,192 padded tokens per forward pass (retrieval.rerank_max_batch_tokens) and the prompt text every candidate shares read once per query (retrieval.rerank_reuse_prefix). On the same machine that took reranking 20 candidates of 40-700 tokens, cut to 256, from p95 0.674 s to 0.397 s (scripts/bench_query_stages.py, 2026-09-25); the figures above predate it.

  • How much of the HNSW index each search reads. retrieval.hnsw_ef_search, applied with SET LOCAL in each channel's own transaction, defaults to 1000 — pgvector's maximum. On the live index, recall@100 against exact search rises 0.944 -> 1.000 (text channel) and 0.792 -> 0.998 (image channel) from pgvector's default of 40 to 1000, and the worst single query rises from 0.79 and 0.38 to 1.00 and 0.98, for about 50 ms more per query.

  • Whether the pages are cached. See "Things that will bite you" below: shared_buffers holds the search working set, pg_prewarm fills it at startup, and imsg status says so.

scripts/bench_retrieval_latency.py sweeps the reranker settings against the real index and reports a quality proxy — agreement with scoring 50 uncapped candidates, not an evaluation. Re-run it before moving either setting, and let the eval harness settle what latency costs.

Startup. imsg mcp local answers the MCP handshake in about 1 s (Claude Code gives up on a server that has not answered within MCP_TIMEOUT, 30 s by default) and, by default, loads no model until the first retrieval call (mcp.local.warm_at_start: false): every client session starts its own server, most never search, and a server that loaded at start held a full model set regardless. The warm-up then runs in the background, logging each step's time on stderr: text embedder 4.8-12.5 s, PE-Core text tower 33.8-50.2 s, reranker 2.7-3.2 s, database buffer pool 0.1-14.3 s — 41.8-65.4 s in all, against 121 s before the 0.6B was pinned. A retrieval tool call that arrives during warm-up (the first one starts it) waits up to 90 s for it, then returns WARMING_UP with an estimate of the seconds remaining; mcp.local.warm_at_start: true loads at start instead, as before. imsg mcp public always loads at start; a model that fails to load makes every such call return WARM_UP_FAILED with the cause. check_permissions is the exception: it is diagnostics, so it answers straight away and carries the warm-up's own state (which step is loading, how many are done, the estimate, the failure) — the way to find out what a server that is not answering searches is doing. The first query after warm-up still costs a little more than the rest (0.94-4.08 s against a 0.87 s median), and the extra is the query embedder's first forward pass at a real shape: 0.13-3.07 s against 0.055 s afterwards.

Idle unload. Every client session runs its own imsg mcp local, and each process holds its own copy of the models (text embedder, PE-Core text tower, reranker, plus MLX's buffer cache) once it has loaded them. Idle sessions used to keep that memory until they exited; five of them exhausted a 64 GiB index host. Now a server loads only on its first retrieval call (above), and drops its models after mcp.local.idle_unload_seconds (default 600; 0 = never) with no retrieval tool call. It clears MLX's and torch's caches so the memory goes back to the system, and never unloads while a call is in flight. The next retrieval call reloads the models through the same warm-up and the same 90 s wait. A reload from a warm OS file cache took 1.2 s for the 8B embedder and 0.6B reranker on an M2 Ultra (9.4 GB of process footprint down to 0.35 GB after unload, 2026-09-24). A reload from disk costs about what a cold start does, and past 90 s the call answers WARMING_UP for the client to retry. check_permissions reports the state as unloaded. mcp.public.idle_unload_seconds defaults to 0: the public surface is a single process, and its 20 s wait is shorter than a reload. A server whose stdin closes, which is what happens when the SSH client disconnects, exits at once, even mid-warm-up. A client that vanishes without closing the connection (a machine that sleeps) is noticed only when sshd's ClientAliveInterval gives up on it.

Interface

What it needs

TextEmbeddingProvider

embed_documents() (bare) and embed_query() (instruction-prefixed); 2048-dim, L2-normalized

MultimodalEmbeddingProvider

embed_images() and embed_text() (paired towers); 1280-dim

BoundaryProvider

Topical boundary indices for a window of messages

OcrProvider / CaptionProvider / TranscriptionProvider

One method each

The reference design uses local MLX-hosted models throughout, on the reasoning that sending a decade of personal messages to a hosted API is a categorically different decision from indexing them on your own machine — and that as of mid-2026 the leading open text-embedding models top the benchmarks anyway, so there is little quality left to trade for it. Nothing in the code requires that choice; the interfaces are provider-agnostic.

⚠️ The dimensions are load-bearing — see the first item below.


Things that will bite you

Learned the expensive way; written down so you don't have to.

  • pgvector's index caps are lower than its type limits. The vector and halfvec types accept up to 16,000 dimensions, but HNSW/IVFFlat indexes cap at 2,000 (vector) and 4,000 (halfvec). A column can be perfectly legal DDL whose index can never be created — the error surfaces at CREATE INDEX, and ignoring it means silently falling back to sequential scan. scripts/lint_ddl.py exists solely to catch this. An earlier revision of this project specified an unbuildable halfvec(4096) for exactly this reason.

  • Audience validation and subject validation are not redundant. On the public surface the subject check answers "is this the owner?" and the audience check answers "was this token minted for this system?" A user's OAuth subject is identical across every app they sign into, so subject-checking alone does not stop a token minted for another application being replayed here. Both, or neither works.

  • Applied migrations are immutable, enforced by hash. Correct a mistake in a later migration; never edit a shipped one.

  • updated_at is enforced by trigger, not convention. Re-segmentation keys off it, so a writer that forgot to bump it would strand chats out of reprocessing — no error, stale results the only symptom.

  • Filters must overfetch, not post-filter. Post-filtering a fixed top-K silently starves results when filters are selective: you get few or zero hits and it looks like "nothing matched."

  • Ingest-time and query-time text normalization must match exactly. If they drift, exact-phrase search silently stops working.

  • Vector search is only fast while its index is in memory. An HNSW query touches a few thousand pages of an index far larger than the default 128 MB shared_buffers. Measured on the external encrypted volume (2026-09-16): queries whose pages the OS had not cached took p95 302 ms (text vectors) and 609 ms (image vectors), against 37 ms and 70 ms once the index files were in the page cache — and with another job's heavy I/O on the same disk, several seconds each. The fix is in three parts: shared_buffers sized to hold what search reads (measured with pg_statio_* deltas over the benchmark queries — 2,296 MiB here, of which 1,193 MiB is the three HNSW indexes and 857 MiB the embedding tables' TOAST), pg_prewarm (migration 0004) filling it — at every server start via RetrievalService.warm_up(), and after a reboot via pg_prewarm.autoprewarm — and imsg status, which prints shared_buffers against the total HNSW index size and warns when the pool is the smaller of the two.

  • Two providers pinned to the same checkpoint still load it twice unless something makes them share. Captioning goes through mlx-vlm and topical boundary detection through mlx-lm; both name the same 35B repo and revision, and each used to build its own model object. Measured: MLX active memory 18.99 → 37.15 GiB as the second one loaded, in one process, for one set of weights. The fix is imsg.shared_vlm_runtime.SharedVlmRuntime, an object a caller passes to both providers — boundary detection then runs text-only through the already-loaded vision-language model, which is the same computation (the rendered chat prompt is byte-identical and the last-position logit row is bit-identical across all 248,320 float32 values). It is explicit rather than a module-level cache precisely so the sharing is visible at the call site; models.share_boundary_and_caption_weights turns it off.

  • A CLIP-style model loads both towers whether or not you use both. PE-Core is 9.01 GiB at fp32, of which the vision tower is 7.01 and the text tower 2.00 — and the query server only ever calls embed_text while the embedding pipeline only ever calls embed_images. The provider now decides on first use and releases the other tower; scripts/verify_pe_core_tower_selection.py proves the vectors are identical bit for bit before and after, because "it cannot change the answer" is an argument, not a measurement. A process that does use both rebuilds once and keeps both.

  • MLX's buffer cache defaults to its memory limit, which is not a bound. It pools freed GPU buffers rather than returning them, and on a 64 GB host the default limit probes at 60.8 GiB; one enrichment process was seen holding 37.25 GiB of freed buffers, which is indistinguishable from a leak and hides real regressions. Every MLX provider now calls imsg.mlx_runtime.bound_buffer_cache at load (models.query_cache_limit_bytes / models.enrichment_cache_limit_bytes). The call is process-wide, so one provider bounding it covers the rest — including mlx_whisper, which has no load hook of its own.

  • Two GPU-heavy jobs on one machine need an arbiter, and it has to survive a crash. The nightly enrichment window overlaps the always-on MCP server; with the duplicate weights gone and no swap at all, query p95 still trebled purely from GPU contention. Enrichment now stands aside for in-flight queries (imsg.db.enrichment_yield_locks, enrichment.yield_to_queries), and the signal is a Postgres session-level advisory lock specifically because the server releases it when the session ends, however it ends — a killed MCP server cannot leave enrichment paused, with no timeout to tune and no stale marker to reap. It is checked between units of work, never during one, so no claimed task is ever abandoned; imsg status reports whether it is yielding right now. What it cannot do is help when the unit of work is long relative to the query rate: measured with it on and off, search p95 during the window was 3.00 s versus 3.06 s — one caption takes 14 s on that host and a search arrives every 2 s, so a worker that resumes between tasks resumes straight into another 14-second caption. The memory fixes above are what closed the gap; this is kept because it costs one round trip when nobody is searching and it will matter wherever the batch is finer-grained than the traffic — not because it earned its place on this workload. Segmentation and embedding use the same gate since 2026-09-26 (imsg.search_yield).

  • The planner can abandon an HNSW index at some ef_search values, and that is not a monotonic effect. pgvector's own cost estimate bounds layer-0 tuples by ef_search while its selectivity term carries log(ef_search) in a denominator, so the estimated cost climbs and then drops back: on this index it was 3,252 at ef_search 40, 13,287 at 280 and 3,719 at 285, against a flat 8,827 for the sequential alternative — which the planner undercosts anyway, because (pgvector's own FAQ) it "doesn't consider out-of-line storage in cost estimates" and the vectors it would sort are 642 MiB of TOAST. Between 170 and 284 the same query took 245 ms instead of 9 ms, exactly and only because of the plan. Every vector channel therefore sets enable_seqscan = off for its own transaction, pgvector's documented remedy.


Layout

src/imsg/
  config/      config surface + validation (enforces the safety rules)
  db/          connection, migrations, cluster fingerprint
  stages/      snapshot, extract, identity, sync
  segment/  backfill/  enrich/  embed/    indexing pipeline
  retrieval/   hybrid query flow, RRF fusion, reranking
  mcp/         auth boundary, local + public surfaces, tools
  export/      default-deny eligibility, plan/approve/push
  backup/      nightly pg_dump + FTS copy, verification, retention
  eval/        metrics, runner, diff
  verify/      seed completeness, attachment reconciliation
migrations/         schema, applied in order by a hash-checked runner
tools/imsg-dump/    GPL-3.0 Rust extraction shim (subprocess only)

Development

uv run ruff check . && uv run mypy . && uv run pytest

Integration tests run against a live PostgreSQL when one is reachable and skip cleanly when it isn't; the unit suite never needs a database.

Licensing

The core is MIT — see LICENSE. Component licensing and the reasoning behind the split are in NOTICE.

In short: tools/imsg-dump/ is GPL-3.0 and carries its own LICENSE. It links the GPL imessage-database crate to parse the attributedBody typedstream format, which is not optional — since Big Sur much of a message's text is not in the text column at all, and readers that only query that column silently return empty strings for large portions of modern history.

It is invoked strictly across a process boundary — spawned as a subprocess, never linked into the Python code. That boundary is deliberate and load-bearing for the licensing split. Vendor it differently and that is yours to reason about.

What isn't here

Instance configuration, by design: real config values, contact seed data, allowlists and eval queries live in a separate private overlay you supply and point at with IMSG_CONFIG, and config.example.yaml ships placeholders only. This repo's git history was rewritten on 2026-09-24 to remove real contact data that had been committed as test fixtures. If you find a real name, number, address, or other personal identifier anywhere in this repo — in code, history, or an issue — do not open a public issue about it; report it privately as described in SECURITY.md.

The design record — architecture rationale, full build spec, and the decision log explaining why each choice above was made — is kept private, since it's written against a specific deployment.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    A read-only MCP server for macOS that enables users to search through iMessage history and analyze conversation patterns using AI-powered tools. It provides detailed statistics on messaging habits, streaks, and contact analytics while keeping all data private and local.
    381 npm
    1
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables full-text search of macOS iMessages including link preview metadata. Works as an MCP server for Claude Desktop to search your messages locally.
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Local-only MCP server that exposes Beeper Texts data (messages, chats, contacts) from the macOS Beeper Desktop SQLite database for AI assistants and automation tools, supporting search and media retrieval across multiple platforms.
    6
    4
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A local MCP server that provides a durable, searchable archive of your WhatsApp history using hybrid retrieval to navigate conversations.
    4
    MIT