Skip to main content
Glama
fledgeling-co

sift-apple-mail-mcp

README.md
<div align="center">

<img src="assets/banner.png" alt="Sift: search your Apple Mail at the speed of thought" width="820">

<br>

[![npm](https://img.shields.io/npm/v/sift-apple-mail-mcp?color=F6AC5C&labelColor=22272F)](https://www.npmjs.com/package/sift-apple-mail-mcp)
[![node](https://img.shields.io/badge/node-%E2%89%A522.23-22272F?labelColor=22272F&color=F6AC5C)](https://nodejs.org)
[![MCP](https://img.shields.io/badge/MCP-2025--06--18-22272F?labelColor=22272F&color=F6AC5C)](https://modelcontextprotocol.io)
[![tests](https://img.shields.io/badge/tests-411%20hermetic-22272F?labelColor=22272F&color=F6AC5C)](docs/PERFORMANCE.md)
[![license](https://img.shields.io/badge/license-MIT-22272F?labelColor=22272F&color=F6AC5C)](LICENSE)

**Let Claude actually search your mail.**<br>
All of it, including the bodies and the PDFs, without sending a single message anywhere.

</div>

---

## What it does

You have years of email sitting on your Mac. Somewhere in it is the invoice, the
thread where you agreed the deadline, the address someone sent you in 2023.

Sift lets an AI assistant find it. Ask in your own words, get the actual thread
back, with the quoted replies stripped out so you read what people said rather
than nineteen copies of the first message.

Nothing leaves your machine. No account, no API key, no bill.

## Why speed is the whole feature

A fast search isn't about the seconds you wait. It's about what an assistant can
*afford to try*.

When a search costs 28 milliseconds, an assistant asks once, takes what comes
back, and moves on. When it costs 4 milliseconds, it can ask twelve different
ways and compare.

That's the difference between:

> *"I found three emails mentioning the invoice."*

and

> *"I searched for the invoice number, the supplier's name, the amount and the
> project code, then cross-checked the threads that matched more than one. The
> agreement is in the March thread; the two later ones are a different invoice
> with a similar number."*

The second answer isn't a cleverer model. It's the same model given room to be
thorough. Ten searches at 4 ms is still under a twentieth of a second.

Where that shows up:

- **Vague questions become answerable.** "That thing Amy sent about the rate
  change" needs several attempts with different words. Cheap attempts mean the
  assistant can make them.
- **Following a trail is viable.** Find a thread, pull its participants, search
  what each of them sent that month, check the attachments. A dozen calls, and
  at these speeds it costs less than one used to.
- **Cross-checking becomes routine.** An assistant that can search four ways
  will notice when three of them disagree, instead of confidently reporting the
  first hit.
- **Big mailboxes stop being special.** 188,000 messages behave like 18,000. The
  index does the work once, at build time.

## The numbers

Measured against [`imdinu/apple-mail-mcp`](https://github.com/imdinu/apple-mail-mcp),
using its own published methodology: five warmups discarded, ten measured calls,
one long-lived process.

| Operation | Baseline | Sift | |
|---|---:|---:|---|
| List accounts | ~1 ms | **0.061 ms** | |
| List 50 emails | ~5 ms | **0.279 ms** | |
| Fetch one email | ~3 ms | **0.010 ms** | warm resolver, not a disk read |
| Search subjects | ~10 ms | **2.71 ms** | |
| Search bodies | ~28 ms | **3.76 ms** | full coverage |

On a **2.6x larger mailbox**: 188,382 messages against their ~73,500.

> [!NOTE]
> Their run was on an M4 Max; this one wasn't, and nobody has run both stacks on
> one machine. Cross-machine timings are indicative rather than controlled. The
> `0.010 ms` fetch is a warm in-memory hit, not a cold read, so the search rows
> are the ones to believe.

**[How these numbers happen →](docs/PERFORMANCE.md)**: the daemon architecture,
the 9.8x warm-path measurement that was wrong twice before it was right, why
Apple's own catalogue needed no optimising at all, and what's deliberately not
claimed.

## Getting started

```bash
npm i -g sift-apple-mail-mcp
```

That's it. The first question you ask starts the index build on its own, answers
straight away, and tells you a build has begun with a rough file count and a
rough duration. Nothing waits on it. You can still run `sift-index build` by
hand if you'd rather watch it, but you don't have to.

Then point your MCP client at the installed binary:

```bash
readlink -f "$(which sift)"
```

Use that absolute path. macOS grants Full Disk Access per binary, so an `npx` or
`nvm` path breaks the permission the moment anything updates.

Sift needs **Full Disk Access** to read your mail: System Settings → Privacy &
Security → Full Disk Access. It reads Apple's files directly and never asks
Mail.app to do anything.

## What you can ask for

Thirteen tools, but you never call them by name; the assistant does.

| | |
|---|---|
| **Find** | Search bodies, subjects, senders, mailboxes, date ranges |
| **Read** | A message, its links, its attachments' text |
| **Follow** | A whole thread, collapsed, with the quoting removed |
| **Browse** | Accounts, mailboxes, recent, unread, flagged |
| **Check** | What fraction of your mail is actually searchable right now |
| **Change** | Create a mailbox, move messages, set a flag. Off by default |

That last one matters more than it sounds. Every answer carries its coverage and
how fresh the index is, so "I found nothing" and "I couldn't look" never get
confused for each other.

**"My inbox" means your inbox, on every account.** Mailboxes are matched on what
they are, not what they are called, which sounds like a detail and is not. A
Gmail account keeps no messages in `INBOX` at all: everything lives in
`[Gmail]/All Mail` and the inbox is a label. Ask for the inbox by name and you
get a folder that exists, reports a count, and is empty. Sift resolves the role
instead, and reads Gmail's label map, so a two-account search covers both. On
the mailbox this was found in that is the difference between 15,164 messages and
28,813.

## Changing mail, if you want that

Sift reads. It changes nothing unless you turn writes on, and it never deletes
anything at all — there is no delete tool, and that is a decision rather than a
gap.

Turn it on by setting `SIFT_ALLOW_WRITES=1` in the environment of the server
process and restarting it. Then the assistant can create a mailbox, move
messages into one, and set Mail's coloured flags. macOS will ask for Automation
permission for Mail the first time one of those runs; that is a separate grant
from Full Disk Access and neither implies the other.

Three things are worth knowing before you enable it.

**Every write asks first, if you want it to.** `dryRun` resolves everything and
reports exactly what would change, touching nothing.

**A result is a ledger, not a yes.** Moving 900 messages where 3 fail returns
897 applied and 3 reasons. Nothing is inferred from an exit code, and a message
the bridge never mentioned is reported as unknown rather than counted as moved.

**A destination can never come from an email.** The mailbox comes from the tool
call. There is no code path from a message body to a folder, so an email asking
to be filed somewhere is text, not an instruction. Every value handed to Mail
goes as an argument, never as script text, so a mailbox named
`" & (do shell script "…") & "` is a mailbox with a silly name.

## The parts worth knowing about

**The index builds itself and stays current.** Ask a question with no index and
Sift starts one in the background, answers from what it has, and reports the
fraction it could search. After that it watches for new mail with a two-stage
check against Mail's own catalogue: a free counter tells it something committed,
and a second query, `max(ROWID)` and a row count, tells it whether that was
actual mail arriving or just you marking something read. The first one alone
fires every time you open an email, which is measured and is why it only gates
rather than decides. A full reconcile runs every fifteen minutes anyway, because
one message arriving and another leaving between two checks looks like nothing
happening.

The build runs as its own detached process, not as a child of the server. Your
MCP client starts and kills that server constantly; a build takes about seven
minutes, so a child would die partway through every single time. Two builds at
once can't happen: the builder takes a lock the kernel holds, which is released
when the process dies however it dies.

**Threads come back collapsed.** The largest thread in the test mailbox is 202
messages. Sift returns 50 authored contributions in 612 ms, every quoted reply
removed. A search matching four messages in one thread says so, rather than
presenting four hits you'd read as four sources.

**PDFs are searchable.** 342 of 400 in the live mailbox return their text.
Scanned ones report `no-text-layer` rather than pretending to be empty. The
parser runs in a sandboxed process with a timeout, a memory ceiling and no
network, because attachments are bytes a stranger sent you.

**Message ids survive a rebuild.** Apple reassigns its internal row ids when Mail
rebuilds its database, so a stored one still resolves afterwards, to a different
message. Sift hands out keys derived from the RFC `Message-ID` instead, which
hold for 99.90% of the mailbox.

**Coverage is honest.** 95.67% of messages are body-searchable. The rest are
encrypted, not downloaded, or genuinely empty, and each is reported as its own
category rather than rounded away.

## Standing on `apple-mail-mcp`

This project exists because [Dinu Catalin-Mihai](https://github.com/imdinu) built
[`apple-mail-mcp`](https://github.com/imdinu/apple-mail-mcp) first and published
how it worked.

Two contributions in particular made this one possible. The first: establishing
that reading Apple's SQLite catalogue and `.emlx` files directly beats driving
Mail.app through AppleScript, which is the architectural decision everything
here is downstream of. The second: publishing real benchmarks with a stated
methodology, including a detailed report more conservative than the headline
figures. Comparable numbers are rare in this corner of the ecosystem, and
they're the only reason the table above is a comparison rather than an
assertion.

Sift is an independent implementation in TypeScript. No code was read or copied:
`apple-mail-mcp` is GPL-3.0 and this is MIT, so the boundary is deliberate and
kept. What was used is public: the feature list, the published timings, and the
benchmark methodology.

If you want a mature Python implementation, go and use theirs.

## Reading further

| | |
|---|---|
| **[Performance](docs/PERFORMANCE.md)** | Every number, how it was measured, what it doesn't prove |
| **[CLAUDE.md](CLAUDE.md)** | Architecture, conventions, divergences from the house stack |
| **[Daemon operations](docs/daemon-operations.md)** | Running it warm, and the permission model |
| **[Specs](docs/specs/)** | One per feature, with the review findings that changed it |
| **[Research](docs/deep-research/)** | The two research inputs the architecture came out of |

## What it doesn't do

It doesn't send, reply, delete or change anything. Sift only reads.

It doesn't work on private, unlisted or undownloaded mail, because that mail
isn't on your disk to read.

It doesn't do semantic or vector search. Dense retrieval was measured and the
numbers are in [Performance](docs/PERFORMANCE.md); it isn't shipped, and the
measurements are recorded so nobody has to redo them.

It has only ever run on one machine, macOS 26.6 with a `V10` mail store. Apple
owns that format and changes it between releases. Sift checks the schema at
startup and refuses with a legible error rather than returning partial results
that look complete.

## For developers

TypeScript, ESM, Node 22.23+. MCP over stdio, `better-sqlite3`, Zod at every
boundary, `exactOptionalPropertyTypes` on, no `any`.

```bash
npm install
npm run gate      # typecheck, lint, control-char scan, 411 tests, build
npm run bench     # benchmarks, needs no Apple Mail data
```

The gate never touches your real mailbox. Every test is hermetic, and a guard
fails the suite if one tries.

## Licence

MIT. Use it, fork it, ship it.

TDQS

A4.3/5.0

Scored across 10 tools

Disambiguation4/5

The tools are mostly distinct: list accounts/mailboxes, get emails vs single email vs body vs attachments vs links, search, thread, and index status. However, get_emails and search both return lists of messages with mailbox/date filters, which could cause an agent to choose the wrong one. The descriptions help but the boundary is not razor-sharp.

Naming Consistency4/5

Most tools follow a clear verb_noun pattern (get_email, list_mailboxes, get_thread), but two deviate: 'index_status' is a noun phrase and 'search' is a bare verb. This is a minor inconsistency in an otherwise readable set.

Tool Count5/5

10 tools is well-scoped for an Apple Mail inspection server. Each tool covers a distinct aspect (accounts, mailboxes, message listing, metadata, body, attachments, links, thread, search, index state) without redundancy or unnecessary sprawl.

Completeness5/5

The server covers the full read/search lifecycle for email: listing structure, retrieving messages, bodies, attachments, links, threads, and searching with index awareness. Since the server appears read-oriented ('sift'), no update/delete/send operations are expected, and index_status addresses the main dependency gap.

Maintenance

ActivitySlowing
ResponsivenessNo issues