imessage-index
Extracts iMessage history, resolves identities, segments conversations topically, embeds and indexes them, and serves hybrid full-text + vector search over an MCP server.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@imessage-indexsearch for messages about the weekend trip"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
imessage-index
A local-first iMessage retrieval system. Extracts your full message history, resolves identities, segments conversations topically, embeds and indexes them, and serves hybrid full-text + vector search over an MCP server — so an AI assistant can actually answer questions about your own correspondence.
Designed to run unattended on a headless Mac, with all derived state on an encrypted volume and nothing leaving the machine unless you explicitly allowlist it.
Is this for you?
What it does. It reads your Mac's Messages history, figures out who each contact really is across their various numbers and emails, groups messages into topical conversations, and builds a local search index — full-text plus vector — that an AI assistant (Claude Desktop or Claude Code, over MCP) can query. Everything stays on your machine unless you explicitly export a filtered subset.
Who it's for right now. Someone comfortable running commands in a
terminal, installing developer tools (Python, Rust, PostgreSQL), and
troubleshooting when a step doesn't go cleanly — not a turnkey app for a
non-technical user yet. The step-by-step guide in
docs/install-macos.md walks through the whole
setup and the scripts/doctor.py diagnostic checks
your machine before you start, but real command-line comfort still
helps.
Hardware. An Apple Silicon Mac (M1 or later) — the pipeline depends on macOS-only APIs (Full Disk Access to Messages, the Contacts framework, Apple Vision) and on Apple's MLX framework for local model inference, which does not run on Intel Macs. RAM: the README's own requirement below is 32 GB unified memory for the full 8B-class local models; smaller models work with less, but no specific lower figure has been measured [UNVERIFIED].
Time and disk. Not measured end to end for a first-time setup [UNVERIFIED] — expect at least an hour between installing dependencies, downloading pinned models, and running the reranker conversion step (see Requirements below). Disk usage depends heavily on the size of your own Messages history and attachments; no general figure has been measured [UNVERIFIED].
Before you read further, the section immediately below is important: this project has never been run end to end against real models, so treat it as a rigorous, well-tested foundation rather than a finished product.
Related MCP server: imessage-rich-search
⚠️ Read this before you clone
No model has run on the pipeline end to end yet. With a plain
uv sync --extra dev and no database, the test suite is 1,803 passed
and 503 skipped (most skips need a scratch PostgreSQL; the rest need the
models and/or export extras, which are not required for a plain
install). With every extra installed (--extra dev --extra models --extra export) and still no database, it's 1,847 passed and 477
skipped, almost all of those needing PostgreSQL — measured 2026-09-24.
The CLI works,
migrations apply against real PostgreSQL + pgvector, and the snapshot →
extract → identity stages have run against a real corpus. Segmentation,
embedding, enrichment and retrieval have only ever run on the fake
providers or on tiny stand-in models (see
what is still unverified).
It ships deterministic fake model providers alongside the real
wiring. The embedding, reranking, captioning, OCR, transcription, and
segmentation providers each have a correctly-dimensioned stub that lets
the pipeline run end to end in tests. imsg.providers.factory selects
between the real implementations and those stubs with one config field,
models.backend (real is the default; fake is explicit opt-in), and
every command that builds providers prints models: backend=<real|fake>
first. Run the pipeline with fake and every stage will report success
while your search results are meaningless. The real implementations
and their pins are described in
Replacing the model providers.
Treat this as a thoroughly tested skeleton with a complete design behind it, not a working product. If you want something that works today, this isn't it. If you want a rigorous starting point that has already made and documented the non-obvious decisions, read on.
What is genuinely done
Area | State |
Schema + migrations | Complete; applied against live PostgreSQL 17 + pgvector |
Pipeline stages (snapshot → extract → identity → segment → enrich → embed → sync → export) | Implemented; unit + integration tested |
Hybrid retrieval (BM25 + text vector + multimodal vector, RRF-fused, reranked) | Implemented |
Local MCP surface (stdio) | Implemented |
Public MCP surface (StreamableHTTP + OAuth) | Implemented; never exposed |
Export gate (default-deny, plan/approve/push/purge) | Implemented and wired to the CLI; the GCS + Discovery Engine transport has never run against a live API |
Eval harness (nDCG@k, recall@k, MRR) | Implemented; metrics verified against hand-computed fixtures |
Real model providers (MLX text/reranker/boundary LLM, Apple Vision OCR, mlx-whisper, mlx-vlm captions, PE-Core) | Implemented behind |
Nightly backup ( | Implemented and wired to the CLI; verified |
AT-1 auth probe ( | Wired to the CLI and fully tested up to the point where live OAuth tokens are required; never run against real tokens or a live gate |
Runs against real data | Snapshot → extract → identity: yes. Segment / embed / enrich with the pinned models: never |
Architecture in one pass
chat.db ──.backup──▶ snapshot ──▶ extract ──▶ identity resolution
│
segmentation
│
┌──────────────────┴─────────────┐
text bodies attachments
│
┌────────────────────────┼───────────┐
PDF text layer images audio/video
└────────────────────────┼───────────┘
│
embed ──▶ pgvector + FTS5 + multimodal vector
│
┌────────────────────────┴──────────┐
local MCP filtered export
(full corpus) (allowlisted)Four choices that shaped everything else:
Segments, not messages. Retrieval returns topically coherent conversation segments. Message-granular results flood a model's context with "sounds good" and prevent reconstructing what was actually discussed.
Everything becomes text first. OCR output, captions, transcripts, and PDF text layers all embed into one text space, so failures stay debuggable — when something doesn't surface you can read the extracted text and see why. A second multimodal vector runs alongside for visual similarity.
Identity resolution precedes segmentation. Nothing downstream keys on a raw handle. One person legitimately has many numbers, emails and aliases across a decade.
Default deny on export. Nothing reaches an external index unless explicitly allowlisted, and a group thread requires every participant allowlisted.
The export gate
The one path by which message content can leave the machine, so every command on it is shaped to refuse. Default deny: a thread exports only if every participant — and every message and tapback sender, the owner included — is explicitly allowlisted. Attachments are gated separately from text bodies.
uv run imsg export plan # eligibility → staged bytes → review report
uv run imsg export approve <run-id> # pins the exact bytes you reviewed
uv run imsg export push <run-id> # re-verifies, then promotes
uv run imsg export purge-person <who> # revocation; exempt from the approval gate
uv run imsg export unclassified-report # weekly: whose threads are still unclassifiedplan writes the review report the whole design rests on — per thread:
participants, message count, date range, sample lines. approve pins
the manifest hash and every staged file hash. push re-checks all of
that and re-derives eligibility from the live database, because the
hashes prove the bytes did not change, not that the world did not: a
participant added to a group between approval and push leaves every hash
green while changing who is in the export. Any drift aborts the push and
requires a new plan.
Revocation is deliberately faster than export — it only ever narrows scope — so a purge needs no approval, though every drift check still applies and the run is recorded in full.
plan, push, purge-person and unclassified-report all take
--dry-run. push --dry-run runs every verification the real push runs
and builds no transport at all, so rehearsing the gate cannot reach the
network.
Nothing can reach Google without a credential you named.
export.gcp_credentials is a keychain: / env: secret reference with
no default; with it unset, push refuses before it opens the
database or imports a Google client library. Leave it out until you mean
it.
Honest limit. A purge reaches the Discovery Engine index and the GCS bucket. Copies already swept into organizational retention, backups, or another person's hands are beyond it. The gate at export time is the actual protection, which is why it denies by default.
Operations
Seven com.imsgindex.* LaunchAgents are rendered by imsg install-agents, and every command they invoke exists — installing them
schedules no job that fails nightly.
uv run imsg backup # daily 04:00: verified pg_dump + FTS copy, 14 kept
uv run imsg backup --dry-run # preconditions + the retention plan, writes nothing
uv run imsg status # mount, Postgres, disk, posture, unclassified threadsWhat imsg backup covers, and what it does not. In scope: the
Postgres dump (everything the pipeline derived exists only there, and
it is the one component that can suffer logical corruption) and the
FTS5 sidecar (rebuildable, copied for recovery speed, under SPEC's
checkpoint + integrity-check conditions). Out of scope, on purpose:
attachments/(~147 GB) — 14 nightly copies of it would be over two terabytes on the same volume as the original, it is a content-addressed cache thatimsg backfill-attachmentsrebuilds, and a blob store does not suffer the logical corruption these copies defend against. But: anything iCloud has already purged exists nowhere else, and that genuinely needs an independently encrypted off-box archive, which this command is not and does not pretend to be.models/— public weights already pinned by repo + immutable revision inmodels/manifest.lock.yaml. Recovery isuv sync --extra models && imsg models verify.ops/— small and irreplaceable, but outside what the spec scopes to this job; flagged in the command's own output rather than silently added.
These copies share the physical device with the data they copy: they protect against logical corruption, not theft or disk failure. The command prints that on every run.
Retention deletes only what it can prove. A directory under
backups/ is removed only if it is a real (non-symlink) direct child,
its name matches the exact backup-set pattern, it holds a MANIFEST.json
that parses and says "complete": true (written last, so its presence
is the completeness evidence), and it is not among the newest 14.
Anything else — a half-written set from an interrupted run, a staging
directory, an operator's own file — is counted, reported, and left
alone.
Attachments whose file is not on this host can be fetched from
wherever else a copy exists: another Mac's Messages folder, its
attached drives, a NAS share (migration 0008, attachment_location).
# on the index host: find candidate copies (counts only are printed)
uv run imsg locate-attachments --listing other-mac=listing.tsv \
--catalog catalogs/ --seed-db other-mac=other-mac-chat.db --dry-run
# on the index host: this host's folder, the cache, attachments.pull shares
uv run imsg backfill-attachments
# on the other Mac, which the index host cannot reach: copy what it has
uv run imsg push-attachments --ssh-host index-host \
--remote-imsg /path/to/imsg --remote-config /path/to/config.yaml \
--root "other-mac=$HOME/Library/Messages/Attachments" --root "D-XXXXX=/Volumes/Drive"
# on the index host again: verify and materialize what was pushed
uv run imsg backfill-attachmentsA copy is accepted at a path a chat.db recorded for the attachment, in a folder named after its GUID under its name, or, flagged, under its name at its byte size anywhere else; never on a name alone. Each copy is checked against the size and hash its location reported before it enters the cache. Every source is only read: both copies are plain rsync runs whose source side is rsync's sender, and the flags that would make it delete or remove anything are refused. No command prints a path.
Attachment text (OCR, PDF and document text, transcripts, captions)
comes from the S5b queue. imsg backfill-attachments queues each
attachment's kinds as it materializes it; imsg enrich --plan queues
everything already materialized (migration 0009 first). Routing reads the
file's content, never its name or chat.db's MIME claim.
uv run imsg migrate # 0009 adds the doc_text kind
uv run imsg enrich --plan --dry-run # per-kind counts, unroutable types; writes nothing
uv run imsg enrich --plan # fill the queue (insert-only, safe to repeat)
uv run imsg enrich --kinds doc_text,pdf_text,transcript,ocr,frame_ocr --limit 200000
uv run imsg enrich --dry-run # what is still claimable, by kindA worker claims one task at a time, cheap kinds first and captions last
unless --kinds names kinds (then that list is the order), newest
attachment first within a kind, and stands aside between tasks while a
search is running. Captions are the slow part (about 16 s each on an
M4 Pro, estimated); the nightly …enrich agent works through them.
Model-heavy commands run one at a time on a host. Each one loads
tens of GiB of model weights, and two together have exhausted a 64 GiB
host's memory. sync, segment, embed, enrich (the worker, not
--plan), eval run and eval pool take an exclusive lock on
<data_root>/run/heavy-models.lock before their models load, and a
second one waits for the first, logging heavy_lock.waiting with the
holder's pid and command. sync snapshots and extracts first and takes
the lock only at segmentation. The kernel releases the lock when its
holder exits, including when it is killed, so there is nothing to clean
up. The MCP servers never take it. A dry run that loads no model
(embed --dry-run, enrich --dry-run, enrich --plan) takes no lock;
segment --dry-run does, because it still runs the boundary model.
uv run imsg status | grep heavy_models_lock # held or not, and by which pid/command
uv run imsg embed --no-wait # exit 1 naming the holder, instead of waitingNothing loads a model the host has no memory for. Before any model
set loads — an MCP server at startup or reloading after an idle unload,
segment, embed, the enrich worker, sync's segmentation and
embedding, eval run/eval pool — it asks imsg.memory_admission: is
the kernel's memory pressure normal, and does available memory (vm_stat
free + inactive + speculative, less what other processes were admitted
for and have not loaded yet) cover the role's expected footprint
(memory.footprints, measured) plus memory.reserve_bytes (8 GiB) for
the rest of the host? A host that cannot be measured is a no. Background
commands wait for a yes, re-checking every 30 s for up to 600 s, then
exit 75 (deferred: memory) with the heavy lock released. An MCP
server loads nothing and answers retrieval calls with the retryable
WARMING_UP code and a "host memory busy" message; the local server tries
again on a later call, the public one by itself every
memory.admission_retry_seconds (15 s). Two processes admitted in the
same second cannot both claim the same free memory: admission runs under
a short host-wide lock and leaves a reservation in
<data_root>/run/memory-reservations/.
MCP servers load before background work. An MCP server whose load is
refused posts a notice, <data_root>/run/live-servers-waiting/<pid>.json,
held under a lock for as long as it waits (the kernel drops the lock if
the server dies). While one is posted, no background command starts a
model load, and a running one stops after its current unit of work (a
chat, a batch, a task) and exits 75, dropping its models and its
reservation, so the server loads at its next try. What background jobs
were admitted for and have not loaded yet counts against the server only
for memory.background_yield_seconds (60 s); after that, only the memory
the host really has, the reserve and the pressure level count, so a job
that cannot stop soon cannot hold the server off. Other MCP servers'
reservations always count.
uv run imsg status | grep -A3 -E 'live_servers_waiting|mcp_public_waiting_on' # who waits, on what
uv run imsg background status # what background work gives way toHeavy background work can be paused. imsg background pause
stops segment, embed, the enrich worker, backfill-attachments and sync's
segmentation and embedding: a command started while paused exits 76
(deferred: paused) without loading anything, and a running one stops
after the unit it is on (a task, a batch, a chat, a file). sync keeps
doing its light work — snapshot, extract, identity — so no message is
lost; it says it skipped segmentation and embedding and exits 76. The MCP
servers are never paused. Another project on the host can pause without
this project's config or encrypted volume by creating the host pause file
(background.host_pause_file, default
~/.config/imessage-index/pause-background):
uv run imsg background pause --reason "photo import" --until 6h # or an ISO time; --until is optional
uv run imsg background resume
uv run imsg background status
# from another project on the same host (same user), no imsg needed:
mkdir -p ~/.config/imessage-index
printf 'reason=photo import\npid=%d\n' $$ > ~/.config/imessage-index/pause-background # paused while this shell runs
rm -f ~/.config/imessage-index/pause-background # resumeA pid= line makes the pause end by itself when that process does, so a
crashed import cannot leave the index paused; an until=<ISO time> line
ends it at that time. Without either, it lasts until the file is removed.
Running work stops when memory runs short. Between units of work every
heavy background command also reads the kernel's memory-pressure level:
at critical it stops at once, at warn it stops if warn is still there 10 s
later (memory.background_stop_at, memory.warn_confirm_seconds), exiting
75. Every MCP server's watchdog reads the same level every 5 s: at critical
a server with its models loaded unloads them at once, without waiting for
its idle timer, as soon as no call is in flight
(memory.local_server_release_at; memory.public_server_release_at can be
never). The public server loads again by itself after
memory.public_rewarm_cooldown_seconds (300 s), if admitted. Each role
also runs under its own MLX memory limit (memory.mlx_memory_limits, set
with mx.set_memory_limit): a guideline MLX uses to release cached
buffers and pace evaluation, never to refuse an allocation, set at 1.25x
each role's measured MLX peak so it changes neither results nor latency.
imsg status shows available memory and the pressure level, the pause
state and its reason, and every running model process with its footprint
(read without root) and what it was admitted for.
Heavy work waits while a search is running. The enrich worker waits
between tasks, segmentation before each boundary-model call and embedding
before each batch (in sync and on their own) while any search is in
flight: a public or local MCP search, or the search page's embedding or
rerank. The servers hold a Postgres advisory lock while they answer, and a
background step checks it, which costs two short statements when nobody
is searching. A wait lasts at most enrichment.yield_max_pause_seconds
(300 s), then the unit runs anyway; for segmentation and embedding, a
pause or a live server waiting for memory ends it at once.
enrichment.yield_to_queries: false turns it off everywhere.
imsg status shows query_in_flight and which step is waiting
(enrichment_yielding_now, segment_yielding_now, embed_yielding_now),
and each command's last lines say how often and how long it waited. It
does not preempt: a search that starts during a batch shares the GPU with
that batch until the batch ends.
Before exposing the public surface, AT-1 must pass:
security add-generic-password -a "$USER" -s imsgindex-at1-owner -w # prompts; no shell history
security add-generic-password -a "$USER" -s imsgindex-at1-nonowner -w
uv run imsg mcp public --probe \
--owner-token-ref keychain:imsgindex-at1-owner \
--foreign-token-ref keychain:imsgindex-at1-nonownerTokens are passed as keychain: / env: references, never values —
a token in argv is readable by every process on the host via ps -ww
and is recorded in shell history. Exit codes are the verdict: 0 pass,
1 fail (a breach), 2 invalid (proved nothing — which is not a
pass), 78 the probe never ran because a precondition was missing. An
invalid result is treated exactly like a failure: scope stays
allowlist.
Requirements
macOS on Apple Silicon. The pipeline depends on macOS-only APIs: Full Disk Access to Messages, the Contacts framework, Apple Vision.
Python 3.12+ and
uv.Rust toolchain, to build the extraction shim.
≥ 32 GB unified memory for 8B-class local models; smaller models work with reduced quality.
An encrypted volume to hold all derived state. Either your Mac's FileVault-encrypted startup disk, or a separate encrypted APFS volume — either is accepted as
data_root, and the mount gate refuses to run against an unencrypted or unmounted location (src/imsg/mount/guard.py). Whichever you use, it must contain a sentinel file named.imsgindex-volumeat the root ofdata_rootso the gate can confirm it is the intended volume and not, say, an unmounted mount point silently resolving to the boot disk underneath it (src/imsg/mount/guard.py).PostgreSQL 17 + pgvector, as a dedicated instance on port 5433, with its data directory under
$DATA_ROOT/pg17(src/imsg/config/ schema.py,src/imsg/db/fingerprint.py). A generic Postgres install on the default port will not work — config validation rejects any other port, and a two-sided fingerprint check refuses to treat any other data directory as this project's own instance. Thepgvectorandpg_prewarmextensions must both be installed into that instance;pg_prewarmfills PostgreSQL's shared buffer cache with the search index at startup so queries don't wait on disk (migration0004_pg_prewarm.sql).scripts/bootstrap_local_postgres.shsets up a cluster meeting all of this for you.A one-time reranker conversion step. The pinned reranker model is not downloaded ready-to-use — it's converted locally from a Hugging Face checkpoint using the exact command recorded in
models/manifest.lock.yaml(see theqwen3-reranker-0.6b-bf16entry). This requires themodelsextra installed first.docs/install-macos.mdhas this command spelled out step by step.
Getting started
git clone https://github.com/msg43/imessage_mcp.git imessage-index
cd imessage-index && uv sync --extra models --extra devThe models extra installs the real model runtimes (mlx, mlx-lm,
mlx-whisper, mlx-vlm, torch, open_clip, pyobjc's Vision bridge,
pillow-heif); uv sync --extra dev alone is enough for the fake
backend and the test suite — tests that need models or the export
extra (below) skip cleanly without them.
cargo build --release --manifest-path tools/imsg-dump/Cargo.tomluv run pytest # 1803 passed, 503 skipped with `--extra dev` alone and no database;
# 1847 passed, 477 skipped with every extra installed and still no database — 2026-09-24Copy config.example.yaml, fill it in, and point the CLI at it:
export IMSG_CONFIG=/path/to/your/config.yamluv run imsg check-permissions && uv run imsg migrateSecrets are never stored in config — they resolve from the macOS
Keychain (keychain:<item>), the environment (env:<VAR>), or a file
only you can read (file:/absolute/path, mode 0600; the usual choice on
a headless host, where the Keychain is unreadable over SSH). Config
validation rejects anything that looks like a literal secret.
Two macOS gotchas that will each cost you an hour. PostgreSQL needs
export LC_ALL=Cor the postmaster dies at startup with "postmaster became multithreaded during startup" — which reads like a corrupt installation and is not. And Full Disk Access cannot be granted over SSH: TCC prompts require a GUI session, and the grant goes to the binary that launches the job, not to Messages.
Then uv run imsg --help. Every stage supports --dry-run.
Connect to Claude
Once the pipeline is set up and indexed, point an AI assistant at the
local MCP server so it can search your messages: uv run imsg mcp local
speaks the MCP protocol over stdio. Full step-by-step instructions for
registering it with both Claude Desktop and Claude Code —
including example config — are in
docs/install-macos.md and
examples/.
Replacing the model providers
This is the work between "tests pass" and "it does something." The
real implementations exist; what is missing is any run of them against
the pinned weights. imsg.providers.factory is the only place providers
are constructed. models.backend in config.yaml selects real (the
default — MLX text embedding, reranker and boundary LLM; Apple Vision
OCR; mlx-whisper transcription; mlx-vlm captioning; PE-Core multimodal
embedding) or fake (the deterministic stand-ins, explicit opt-in);
every command that builds providers prints models: backend=<...>
first. The factory imports the real classes lazily, by dotted path, and
builds them from the repo ids and immutable revisions in config.yaml,
whose defaults mirror models/manifest.lock.yaml — repo, commit sha,
license, expected dimension, quantization, runtime floors and a
smoke-test record per model. uv run imsg models verify (also
scripts/verify_model_manifest.py) re-resolves each repo against the
Hugging Face API and checks the installed packages against the floors;
it reports drift and never rewrites the lock unless given --write.
The runtime packages live behind the models extra; a missing package
fails as one clear imsg: ... line, not a traceback. The two fixed
prompts (prompts/segment_boundaries.txt, prompts/caption.txt) ship
in the repo and are used unless paths.data_root holds a copy at the
same relative path; each run prints which file it used, because those
bytes are hashed into seg_config_hash and the caption provenance.
What has been run, and what is still unverified (2026-09-15).
uv run python scripts/smoke_test_models.py (also
imsg.providers.model_smoke) downloads every pinned model at its sha,
checksums it (artifact_sha256, defined in the lock header), builds the
real provider through the factory and runs one fictional input per role,
recording load time, inference time and peak memory per model into the
lock with --write. On an M2 Ultra with 128 GB every entry passes. Still
unverified: segment / embed / enrich on a real chat, batched
throughput and memory (the recorded peaks are single-input), the
deployment host's memory, and retrieval quality with the pinned weights —
establishing that is the first real task.
The reranker is a local conversion, not a Hub download. The only
8-bit MLX conversion of Qwen3-Reranker-8B on the Hub ships no lm_head
tensor, so the model card's yes/no logits cannot be computed from it (the
2026-09-14 smoke run found this; the evidence is kept in the lock entry's
notes). The lock therefore pins the reranker as source: local_conversion:
the upstream repo and commit sha, the license, the converter
(mlx-lm==0.31.3), the exact command that produces it, its
output_dir relative to paths.data_root, and the artifact_sha256 of
that directory. The pinned reranker is Qwen3-Reranker-0.6B since
2026-09-17, for the latency budget below, and since 2026-09-25 its
unquantized 16-bit (bf16) conversion (1.12 GiB on disk), which runs faster
on the GPU than the 8-bit mxfp8 conversion before it. Both earlier
conversions, the 0.6B's mxfp8 build and the 8B, stay in the lock as
status: retained — pinned and verified exactly like an active entry,
claiming no role; setting retrieval.reranker_model to the mxfp8 build's
directory switches back to it. imsg models verify prints which entry is
active for every role and lists the retained ones separately. To
reproduce a conversion, run the recorded command with $DATA_ROOT set to
your data root and the models extra installed (a few seconds for the
0.6B, about 16 for the 8B), then
scripts/smoke_test_models.py --only qwen3-reranker-0.6b-bf16 --data-root $DATA_ROOT to confirm the digest — two runs of the bf16 recipe on
2026-09-25 produced the same artifact byte for byte. In config.yaml,
retrieval.reranker_model names that directory (data-root-relative) and
retrieval.reranker_revision the upstream commit; the factory reads
the value as a local directory when it exists under the data root and as
a Hugging Face repo id otherwise, and the provider records model_id as
<dir>@<upstream sha>. uv run imsg models verify --data-root $DATA_ROOT re-resolves the upstream repo for drift, recomputes the
directory's digest against the lock, and checks the runtimes; --write
never advances an upstream pin (the directory was converted from the
pinned commit — re-convert and re-pin by hand to move it).
Search latency: the reranker's size, and whether the index is in
memory. Measured end to end through the real MCP surface on an M2 Ultra
on 2026-09-17 — imsg mcp local under a stdio client, 20 fictional
queries, three cold starts, with another job using the same disk array at
216-2,177 MB/s throughout — a whole search_messages round trip took
p50 0.87 s, p95 1.14 s, max 1.28 s. Where it goes, p95 per stage:
reranking 0.67 s, the full-text channel 0.24 s, the two vector channels
0.05 and 0.03 s, the query embedding 0.06 s, the multimodal text tower
0.02 s, summary fetch 0.002 s, the audit row 0.07 s.
Three settings decide most of that, and all three were chosen by measurement rather than taste:
Which reranker. Qwen3-Reranker-0.6B, at
retrieval.rerank_top20 andretrieval.rerank_doc_max_tokens256. The 8B conversion pinned before it reads roughly 650-780 tokens a second here: scoring 50 uncapped candidates meant ~20,000 tokens and a p95 of 36 s, and its fastest setting that fit a budget (10 candidates, 64 tokens, p95 1.97 s of reranking) agreed with its own full ranking no better than doing no reranking at all. How many candidates are reranked matters more than the reranker's size: with 10, the returned top 10 is the fused top 10 in another order, so reranking cannot lift anything from further down. Since 2026-09-25 the 0.6B runs as its bf16 build, with up to 8,192 padded tokens per forward pass (retrieval.rerank_max_batch_tokens) and the prompt text every candidate shares read once per query (retrieval.rerank_reuse_prefix). On the same machine that took reranking 20 candidates of 40-700 tokens, cut to 256, from p95 0.674 s to 0.397 s (scripts/bench_query_stages.py, 2026-09-25); the figures above predate it.How much of the HNSW index each search reads.
retrieval.hnsw_ef_search, applied withSET LOCALin each channel's own transaction, defaults to 1000 — pgvector's maximum. On the live index, recall@100 against exact search rises 0.944 -> 1.000 (text channel) and 0.792 -> 0.998 (image channel) from pgvector's default of 40 to 1000, and the worst single query rises from 0.79 and 0.38 to 1.00 and 0.98, for about 50 ms more per query.Whether the pages are cached. See "Things that will bite you" below:
shared_buffersholds the search working set,pg_prewarmfills it at startup, andimsg statussays so.
scripts/bench_retrieval_latency.py sweeps the reranker settings against
the real index and reports a quality proxy — agreement with scoring 50
uncapped candidates, not an evaluation. Re-run it before moving either
setting, and let the eval harness settle what latency costs.
Startup. imsg mcp local answers the MCP handshake in about 1 s
(Claude Code gives up on a server that has not answered within
MCP_TIMEOUT, 30 s by default) and, by default, loads no model until the
first retrieval call (mcp.local.warm_at_start: false): every client
session starts its own server, most never search, and a server that
loaded at start held a full model set regardless. The warm-up then runs in
the background, logging each step's time on stderr: text embedder
4.8-12.5 s, PE-Core text tower 33.8-50.2 s, reranker 2.7-3.2 s, database
buffer pool 0.1-14.3 s — 41.8-65.4 s in all, against 121 s before the 0.6B
was pinned. A retrieval tool call that arrives during warm-up (the first
one starts it) waits up to 90 s for it, then returns WARMING_UP with an
estimate of the seconds remaining; mcp.local.warm_at_start: true loads
at start instead, as before. imsg mcp public always loads at start; a model
that fails to load makes every such call return WARM_UP_FAILED with the
cause. check_permissions is the exception: it is diagnostics, so it
answers straight away and carries the warm-up's own state (which step is
loading, how many are done, the estimate, the failure) — the way to find
out what a server that is not answering searches is doing. The first
query after warm-up still costs a little more than the rest (0.94-4.08 s
against a 0.87 s median), and the extra is the query embedder's first
forward pass at a real shape: 0.13-3.07 s against 0.055 s afterwards.
Idle unload. Every client session runs its own imsg mcp local, and
each process holds its own copy of the models (text embedder, PE-Core text
tower, reranker, plus MLX's buffer cache) once it has loaded them. Idle
sessions used to keep that memory until they exited; five of them
exhausted a 64 GiB index host. Now a server loads only on its first
retrieval call (above), and drops its models after
mcp.local.idle_unload_seconds (default 600; 0 = never) with no retrieval
tool call. It clears MLX's and torch's caches
so the memory goes back to the system, and never unloads while a call is
in flight. The next retrieval call reloads the models through the same
warm-up and the same 90 s wait. A reload from a warm OS file cache took
1.2 s for the 8B embedder and 0.6B reranker on an M2 Ultra (9.4 GB of
process footprint down to 0.35 GB after unload, 2026-09-24). A reload from
disk costs about what a cold start does, and past 90 s the call answers
WARMING_UP for the client to retry. check_permissions reports the state
as unloaded. mcp.public.idle_unload_seconds defaults to 0: the public
surface is a single process, and its 20 s wait is shorter than a reload.
A server whose stdin closes, which is what happens when the SSH client
disconnects, exits at once, even mid-warm-up. A client that vanishes
without closing the connection (a machine that sleeps) is noticed only
when sshd's ClientAliveInterval gives up on it.
Interface | What it needs |
|
|
|
|
| Topical boundary indices for a window of messages |
| One method each |
The reference design uses local MLX-hosted models throughout, on the reasoning that sending a decade of personal messages to a hosted API is a categorically different decision from indexing them on your own machine — and that as of mid-2026 the leading open text-embedding models top the benchmarks anyway, so there is little quality left to trade for it. Nothing in the code requires that choice; the interfaces are provider-agnostic.
⚠️ The dimensions are load-bearing — see the first item below.
Things that will bite you
Learned the expensive way; written down so you don't have to.
pgvector's index caps are lower than its type limits. The
vectorandhalfvectypes accept up to 16,000 dimensions, but HNSW/IVFFlat indexes cap at 2,000 (vector) and 4,000 (halfvec). A column can be perfectly legal DDL whose index can never be created — the error surfaces atCREATE INDEX, and ignoring it means silently falling back to sequential scan.scripts/lint_ddl.pyexists solely to catch this. An earlier revision of this project specified an unbuildablehalfvec(4096)for exactly this reason.Audience validation and subject validation are not redundant. On the public surface the subject check answers "is this the owner?" and the audience check answers "was this token minted for this system?" A user's OAuth subject is identical across every app they sign into, so subject-checking alone does not stop a token minted for another application being replayed here. Both, or neither works.
Applied migrations are immutable, enforced by hash. Correct a mistake in a later migration; never edit a shipped one.
updated_atis enforced by trigger, not convention. Re-segmentation keys off it, so a writer that forgot to bump it would strand chats out of reprocessing — no error, stale results the only symptom.Filters must overfetch, not post-filter. Post-filtering a fixed top-K silently starves results when filters are selective: you get few or zero hits and it looks like "nothing matched."
Ingest-time and query-time text normalization must match exactly. If they drift, exact-phrase search silently stops working.
Vector search is only fast while its index is in memory. An HNSW query touches a few thousand pages of an index far larger than the default 128 MB
shared_buffers. Measured on the external encrypted volume (2026-09-16): queries whose pages the OS had not cached took p95 302 ms (text vectors) and 609 ms (image vectors), against 37 ms and 70 ms once the index files were in the page cache — and with another job's heavy I/O on the same disk, several seconds each. The fix is in three parts:shared_bufferssized to hold what search reads (measured withpg_statio_*deltas over the benchmark queries — 2,296 MiB here, of which 1,193 MiB is the three HNSW indexes and 857 MiB the embedding tables' TOAST),pg_prewarm(migration 0004) filling it — at every server start viaRetrievalService.warm_up(), and after a reboot viapg_prewarm.autoprewarm— andimsg status, which printsshared_buffersagainst the total HNSW index size and warns when the pool is the smaller of the two.Two providers pinned to the same checkpoint still load it twice unless something makes them share. Captioning goes through
mlx-vlmand topical boundary detection throughmlx-lm; both name the same 35B repo and revision, and each used to build its own model object. Measured: MLX active memory 18.99 → 37.15 GiB as the second one loaded, in one process, for one set of weights. The fix isimsg.shared_vlm_runtime.SharedVlmRuntime, an object a caller passes to both providers — boundary detection then runs text-only through the already-loaded vision-language model, which is the same computation (the rendered chat prompt is byte-identical and the last-position logit row is bit-identical across all 248,320 float32 values). It is explicit rather than a module-level cache precisely so the sharing is visible at the call site;models.share_boundary_and_caption_weightsturns it off.A CLIP-style model loads both towers whether or not you use both. PE-Core is 9.01 GiB at fp32, of which the vision tower is 7.01 and the text tower 2.00 — and the query server only ever calls
embed_textwhile the embedding pipeline only ever callsembed_images. The provider now decides on first use and releases the other tower;scripts/verify_pe_core_tower_selection.pyproves the vectors are identical bit for bit before and after, because "it cannot change the answer" is an argument, not a measurement. A process that does use both rebuilds once and keeps both.MLX's buffer cache defaults to its memory limit, which is not a bound. It pools freed GPU buffers rather than returning them, and on a 64 GB host the default limit probes at 60.8 GiB; one enrichment process was seen holding 37.25 GiB of freed buffers, which is indistinguishable from a leak and hides real regressions. Every MLX provider now calls
imsg.mlx_runtime.bound_buffer_cacheat load (models.query_cache_limit_bytes/models.enrichment_cache_limit_bytes). The call is process-wide, so one provider bounding it covers the rest — includingmlx_whisper, which has no load hook of its own.Two GPU-heavy jobs on one machine need an arbiter, and it has to survive a crash. The nightly enrichment window overlaps the always-on MCP server; with the duplicate weights gone and no swap at all, query p95 still trebled purely from GPU contention. Enrichment now stands aside for in-flight queries (
imsg.db.enrichment_yield_locks,enrichment.yield_to_queries), and the signal is a Postgres session-level advisory lock specifically because the server releases it when the session ends, however it ends — a killed MCP server cannot leave enrichment paused, with no timeout to tune and no stale marker to reap. It is checked between units of work, never during one, so no claimed task is ever abandoned;imsg statusreports whether it is yielding right now. What it cannot do is help when the unit of work is long relative to the query rate: measured with it on and off, search p95 during the window was 3.00 s versus 3.06 s — one caption takes 14 s on that host and a search arrives every 2 s, so a worker that resumes between tasks resumes straight into another 14-second caption. The memory fixes above are what closed the gap; this is kept because it costs one round trip when nobody is searching and it will matter wherever the batch is finer-grained than the traffic — not because it earned its place on this workload. Segmentation and embedding use the same gate since 2026-09-26 (imsg.search_yield).The planner can abandon an HNSW index at some
ef_searchvalues, and that is not a monotonic effect. pgvector's own cost estimate bounds layer-0 tuples byef_searchwhile its selectivity term carrieslog(ef_search)in a denominator, so the estimated cost climbs and then drops back: on this index it was 3,252 atef_search40, 13,287 at 280 and 3,719 at 285, against a flat 8,827 for the sequential alternative — which the planner undercosts anyway, because (pgvector's own FAQ) it "doesn't consider out-of-line storage in cost estimates" and the vectors it would sort are 642 MiB of TOAST. Between 170 and 284 the same query took 245 ms instead of 9 ms, exactly and only because of the plan. Every vector channel therefore setsenable_seqscan = offfor its own transaction, pgvector's documented remedy.
Layout
src/imsg/
config/ config surface + validation (enforces the safety rules)
db/ connection, migrations, cluster fingerprint
stages/ snapshot, extract, identity, sync
segment/ backfill/ enrich/ embed/ indexing pipeline
retrieval/ hybrid query flow, RRF fusion, reranking
mcp/ auth boundary, local + public surfaces, tools
export/ default-deny eligibility, plan/approve/push
backup/ nightly pg_dump + FTS copy, verification, retention
eval/ metrics, runner, diff
verify/ seed completeness, attachment reconciliation
migrations/ schema, applied in order by a hash-checked runner
tools/imsg-dump/ GPL-3.0 Rust extraction shim (subprocess only)Development
uv run ruff check . && uv run mypy . && uv run pytestIntegration tests run against a live PostgreSQL when one is reachable and skip cleanly when it isn't; the unit suite never needs a database.
Licensing
The core is MIT — see LICENSE. Component licensing and
the reasoning behind the split are in NOTICE.
In short: tools/imsg-dump/ is GPL-3.0 and carries its own
LICENSE. It links the GPL imessage-database crate to parse the
attributedBody typedstream format, which is not optional — since Big
Sur much of a message's text is not in the text column at all, and
readers that only query that column silently return empty strings for
large portions of modern history.
It is invoked strictly across a process boundary — spawned as a subprocess, never linked into the Python code. That boundary is deliberate and load-bearing for the licensing split. Vendor it differently and that is yours to reason about.
What isn't here
Instance configuration, by design: real config values, contact seed
data, allowlists and eval queries live in a separate private overlay
you supply and point at with IMSG_CONFIG, and config.example.yaml
ships placeholders only. This repo's git history was rewritten on
2026-09-24 to remove real contact data that had been committed as test
fixtures. If you find a real name, number, address, or other personal
identifier anywhere in this repo — in code, history, or an issue — do
not open a public issue about it; report it privately as described in
SECURITY.md.
The design record — architecture rationale, full build spec, and the decision log explaining why each choice above was made — is kept private, since it's written against a specific deployment.
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP connector for iMessage & Contacts via a local Mac agent + Vercel relay
- AmberOAuthcom.ambermem
Long-term memory for AI assistants. Hybrid retrieval, query expansion, auto-topics.
Search your AI chat history (ChatGPT, Claude, Codex) from any MCP client. Remote, private, read-only
Search, read, and write your Apple Notes from ChatGPT/Claude via a local Mac agent + MCP relay.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceA read-only MCP server for macOS that enables users to search through iMessage history and analyze conversation patterns using AI-powered tools. It provides detailed statistics on messaging habits, streaks, and contact analytics while keeping all data private and local.381 npm1-
- AlicenseAqualityCmaintenanceEnables full-text search of macOS iMessages including link preview metadata. Works as an MCP server for Claude Desktop to search your messages locally.1MIT
- AlicenseAqualityCmaintenanceLocal-only MCP server that exposes Beeper Texts data (messages, chats, contacts) from the macOS Beeper Desktop SQLite database for AI assistants and automation tools, supporting search and media retrieval across multiple platforms.64MIT
- AlicenseNot gradedqualityBmaintenanceA local MCP server that provides a durable, searchable archive of your WhatsApp history using hybrid retrieval to navigate conversations.4MIT