mcp-mail-macos
Allows an agent to read, search, send, move, and delete email from Gmail accounts via macOS Mail.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-mail-macossearch my mailbox for invoices from last month"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-mail-macos
An MCP server that drives macOS Mail: read, search, send, organise. It runs over stdio, launched on demand by the MCP client — there is no long-lived process.
Two mechanisms live side by side, deliberately. Actions go through AppleScript, the only interface that can make Mail do anything. Search goes through a local SQLite index built from Mail's own storage, because AppleScript needs seconds per message and cannot search an archive of tens of thousands of mails in any usable time.
Tested on macOS 27, Python 3.14, Mail 16, against Gmail, IMAP and Exchange accounts, on a mailbox of roughly 50,000 messages spanning several years.
Read this before installing. This server grants an agent the right to read, send, move and delete mail on every account Mail is configured with, and the search index needs Full Disk Access, which macOS cannot scope to a single folder. See Before you trust it with your mail.
Contents
Related MCP server: email-mcp
Before you trust it with your mail
This is a local tool for one person on their own machine. It is not a service, and it was not designed to be exposed to several users. What it asks for is broad, and worth weighing before installing.
Automation lets an agent act on every account. One grant, once, and the server can read, send, reply, move and delete across all of them. There is no per-account allowlist — adding one means filtering in two places (see Scope).
Full Disk Access is all or nothing. The index reads ~/Library/Mail, which
macOS protects; there is no setting scoped to that folder. Granting it also
grants Messages, browser history and other applications' data to whatever
application you granted it to, and for every future session until you revoke it.
Message content reaches the agent unfiltered. Anyone can send you mail, and that mail lands in an agent's context as text. That is the classic prompt injection setup, and no permission dialog stands between the two.
Reasonable precautions, in rough order of value:
Try it on a secondary account first, before pointing it at anything that matters.
Keep the confirmation guard. Every send requires
confirm=trueand returns a preview otherwise. It makes each send deliberate and shows exactly what would leave.Only grant Full Disk Access if you need indexed search, and revoke it afterwards — everything already indexed stays searchable. Set
index_max_age_minuteshigh inconfig.jsonso the server stops trying to refresh.Decide whether an agent should send at all. Preparing drafts as
.emlfiles and sending them yourself is a perfectly good mode;write_drafttouches nothing but a folder.Run the checks before real use:
python3 -m unittest discover -s tests -t .for the logic, thentest_manual.py readagainst your own Mail.
Requirements
OS | macOS, with Mail configured and its accounts loaded. AppleScript and Mail's storage layout are the whole foundation, so there is no path to another platform. |
Python | 3.11 or later — the code uses |
Version 16 (macOS 13+). The AppleScript dictionary has been stable across these releases; the internal index schema has not (see below). |
Dependencies
One, declared in requirements.txt:
mcp>=1.2.0That is the official Model Context Protocol SDK. Both generations work and the import picks whichever is installed:
SDK | Class | Import |
1.x |
|
|
2.x |
|
|
The decorator API is identical between the two, so nothing else changes.
One more, optional, in requirements-semantic.txt:
sqlite-vec>=0.1.6It computes cosine similarity inside SQLite for search by meaning
(below), and is imported only when that feature runs.
Without it the server, the index and keyword search are untouched; semantic
search then compares vectors in plain Python, which is fine for a filtered
search and refused (with a hint) for a whole-mailbox one. Search by meaning also
needs Ollama running locally with the bge-m3 model
(brew install ollama, brew services start ollama, ollama pull bge-m3,
about 1.2 GB); it is reached over HTTP with urllib, nothing to install in
Python.
Everything else is standard library — sqlite3 for the index and its FTS5
tables, email for parsing .emlx containers and writing .eml drafts,
subprocess for osascript, unicodedata, urllib.parse, json, tempfile.
No compiled extension, no build step.
Running the unit tests needs nothing at all beyond the standard library: they never import the SDK.
Install
git clone https://github.com/beeraw/mcp-mail-macos.git
cd mcp-mail-macos
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/pip install -r requirements-semantic.txt # optional: search by meaningmacOS permissions
Two separate grants, for two different needs. The first is required. The second only concerns indexed search.
1. Automation — driving Mail
On the first call, macOS asks for permission to control Mail. The dialog appears once, and it is attributed to the application launching the server, not to Python.
If it was denied it will not ask again: restore it in System Settings →
Privacy & Security → Automation, unfold the application concerned and tick
Mail. Until then every tool returns permission_denied with that reminder.
To trigger the prompt at a quiet moment, before wiring anything up:
.venv/bin/python test_manual.py read2. Full Disk Access — building the index
The indexer reads ~/Library/Mail, which macOS protects through TCC. There is
no setting scoped to that folder: the only lever is Full Disk Access, all or
nothing, in System Settings → Privacy & Security.
The grant goes to the process's responsible application. For Claude Code
that is /Applications/Claude.app — not the nested claude-code binary, and
not Python. It is only read at launch, so the application has to be quit and
restarted.
Granted to | Consequence |
| The index refreshes itself from the MCP server. In exchange the grant covers every protected location, not just Mail, and applies to future sessions. |
Terminal only | The server can search the index but not refresh it. |
Revoking it later breaks nothing: everything already indexed stays searchable, only updates stop.
3. An account password — writing drafts and sending
Reading mail goes through Mail. Writing one does not: a draft is filed on the server over IMAP, and a send is submitted over SMTP. Mail holds the host, the port and the user name for both, so the only thing missing is a password it will not hand out.
One is stored per account in the login keychain, under the service name in
keychain_service, keyed by the account's own user name:
security add-generic-password -U -s mcp-mail-macos -a you@example.com -w 'the password'On an account with two step verification — every Google account, for one — that must be an app password, not the account password. Google offers them at myaccount.google.com/apppasswords. If the page answers that the setting is unavailable, two step verification is off on that account, or a Workspace administrator has disabled app passwords.
The lookup runs through /usr/bin/security, the same tool that stored the
password, so the keychain grants it without a prompt. A password added by hand
through Keychain Access is refused until it is allowed once.
An account with no IMAP server — an Exchange or Outlook account — cannot be written to this way, and says so rather than failing obscurely.
Add to Claude Code
Absolute paths, since the server can be launched from anywhere:
claude mcp add mail-macos -s user -- /path/to/mcp-mail-macos/.venv/bin/python /path/to/mcp-mail-macos/server.py-s user makes it available in every project; without the flag it stays scoped
to the current one. Check with claude mcp list. The tools appear once Claude
Code restarts.
Configuration
Nothing has to be configured: every setting falls back to something that works out of the box, and drafts and the index stay inside the repository directory.
To change any of it, copy the example and edit what you need:
cp config.example.json config.jsonconfig.json is gitignored, so local paths never end up in a commit. Any key
may be omitted. An environment variable of the form MAIL_MCP_<KEY> overrides
both the file and the default — convenient when the MCP client passes its own
configuration:
claude mcp add mail-macos -s user -e MAIL_MCP_DRAFTS_FOLDER="$HOME/Documents/Outgoing mail" -- /path/to/.venv/bin/python /path/to/server.pyKey | Default | What it does |
|
| Where |
|
| How long an unsent draft may sit on disk |
|
| How long a sent draft stays archived |
|
| The line above the quoted original in a reply, in the language of the person answered |
|
| How |
|
| Keychain service name the account passwords are stored under |
|
| The search index |
|
| Mail's storage, where the index is built from |
|
| Past this age, |
|
| Ceiling for a read call, in seconds |
|
| Ceiling for a send or move |
|
| Characters of body kept per message when indexing |
| beside | The attachment text index, |
|
| OCR of loose images (at least 50 KB, not named like a logo); scanned PDFs are always read |
|
| Largest attachment read: PDFs up to this, other types up to half |
|
| Characters kept per attachment |
|
| After a message sync, start the attachment sync in the background (once that index exists) |
| beside | The named searches of |
| beside | The embeddings for search by meaning, |
|
| Where Ollama listens |
|
| The model asked to embed text; changing it needs |
|
| Seconds allowed to embed a query before |
|
|
|
|
| After a message sync, start |
The launchd agent is the one place a path cannot come from configuration:
launchd needs absolute paths in the plist itself. Replace /ABSOLUTE/PATH/TO
in launchd/com.mcp-mail-macos.sync.plist before installing it.
The 29 tools
Search across everything
Tool | Purpose |
| Search every account, through the local index; |
| Count matching messages per sender, domain, month, year, account, mailbox or recipient ("who writes to me most", volumes per month) |
| Save a frequent search under a name ( |
| The whole conversation a message belongs to |
| The messages closest in meaning to a given one ("more like this"), from the stored vectors |
| What the index holds, how old it is, how many messages have a searchable body (per account and overall), and the state of the attachment and vector indexes |
| Bring the index up to date |
search_all covers the whole archive in milliseconds. Subject, sender,
To and Cc recipients, body and attachment names are all indexed. FTS5 syntax
works — subject: invoice, sender: jane (recipients: searches
To and Cc together), AND / OR / NOT, "exact phrase", NEAR(one two, 5).
A query that is not valid FTS5 (invoice 12/2025) is reinterpreted word by
word, which the answer reports in interpreted_as.
saved_search keeps a frequent search under a name. save stores the
search_all parameters you pass (the query, with any Gmail operators, is
required) plus an optional description; run executes it and returns exactly
what search_all returns, and any parameter passed to run overrides the
stored one for that run (saved_search("run", "unread from Jane", limit=5)).
Saving under an existing name (any case) replaces that search. Store relative
operators such as newer_than:7d rather than fixed dates: the query is kept as
text and evaluated again at each run. One tool with an action, rather than one
tool per action, keeps the tool list short. The file is written atomically and
holds queries only; aggregate does not take a saved search.
aggregate takes the same query, operators and filters but returns counts
instead of messages: aggregate("sender", "to:me -is:bulk", since="2026-01-01")
lists who writes most, aggregate("month", "from:@example.com") gives the
volume per month. Nothing is excluded by default; -is:bulk leaves out
newsletters. A message in several mailboxes counts once (except
group_by="mailbox"); text inside attachments is not searched.
Operators
Gmail-style operators are read first and become SQL filters (on messages,
locations and recipients); only the remaining text goes to FTS5. Write them
with no space after the colon, quote values that hold spaces (from:"jane doe"),
and prefix - to negate (-is:bulk). Keys are case-insensitive, operators
combine with AND only (also two of the same kind): OR next to an operator, or an operator sharing parentheses with OR or free text, is refused with a hint (use -key: to exclude, run two searches for alternatives). NOT from:x is -from:x, and parentheses wrapping operators only, (from:a has:attachment), are ignored. Operators also combine with the
account, mailbox, unread_only, flagged_only, since and until
parameters. A query made only of operators works and returns the newest first.
An unknown word: stays free text; a bad value is refused with a hint. The
answer's filters shows how the query was understood (original_query holds
what was typed).
Operator | Meaning |
| sender: full address, |
| same, on the recipients table; |
| a real attachment (same meaning as |
| attachment name, as typed or stemmed ( |
| size, |
| relative to now: |
|
|
| read state, flag ( |
| mailbox, case-insensitive: the whole name or its last path segment as Mail stores it, so a localised display name such as "Boîte de réception" is not mapped to INBOX ( |
A message filed in several mailboxes (Gmail's All Mail and Important, typically)
matches is:unread, is:read, is:flagged and in: when any of its copies
does; a negation is the exact opposite (-is:unread = unread in none).
to: and cc: are also names of FTS columns: the operator wins. To search
those columns as words, write a space (to: jane) or braces ({to}: jane,
{to cc}: jane). me is read from Mail's account list on first use (one
AppleScript call, cached for the life of the server); if Mail cannot answer,
only that query fails, with a hint to use the address itself.
Results are ranked by relevance by default (sort="relevance"): weighted bm25
with a subject match counting most, then sender, attachment names, and
To/Cc/body, times a moderate recency bonus (up to +30 % for a mail received
today, +15 % at one year old; a strong old match still beats a weak recent one).
Each result carries has_attachment (a real attachment: inline logos and
png/gif/bmp/svg names are ignored) and is_bulk (a List-Id or List-Unsubscribe
header: mailing lists and newsletters), and a snippet (snippets=false to skip): about 200
characters of the body around the first matched word, matched like the index
does (case and accents ignored, term* prefixes and quoted phrases handled), or
the start of the body when the word is only in the subject, sender or an
attachment name. The index cannot store text, so it is read from the message's
.emlx file; the file is located from the id alone (Mail shards its folders by
the id's digits), with no extra index column. It is null when the file is
missing or not downloaded. sort="date" returns the newest matches first instead. The weights and the
bonus are constants at the top of mail_search.py. The command-line
python3 mail_index.py --search QUERY ranks the same way, --sort date to
order by date.
When the attachment index exists, the text of PDFs, Word, Excel, PowerPoint and
scans is searched too (see Attachment text). A message found
only through an attachment carries attachment_match (filename, and a
snippet of the attachment when snippets are on), and its snippet starts with
[attachment: name]. Without attachments.sqlite nothing changes.
get_thread uses the conversation grouping Mail computes itself, carried in the
index. The whole exchange comes back, including replies filed in another mailbox
or sent from another account.
Read directly
Tool | Purpose |
| Accounts and mailboxes, with unread counts |
| Messages of one mailbox, newest first |
| Full message: body, headers, attachments |
| Unread counts, per mailbox or across accounts |
These ask Mail directly, so they see the real state including what has just arrived.
There is deliberately no tool that searches through Mail. Mail serves Apple
events on the thread that draws its interface, so any search wide enough to be
useful freezes the app for minutes — and the timeout does not rescue it: killing
osascript leaves Mail chewing on the event it already accepted, so the freeze
outlives the call. Search goes through the index, which reads the same store
from disk and needs the same Full Disk Access. If the index is missing,
search_all says so and sync_index builds it; there is no faster path worth
having.
Prepare a message for review
Tool | Purpose |
| Write a draft as an |
| Drafts waiting to be sent |
| Full content of one draft |
| Change the draft in place — see Edit a draft |
| Send the draft, then file it away |
| Delete a draft that will not be sent |
| Sweep forgotten drafts |
See Drafts are files for why they live outside Mail.
Send
Tool | Purpose |
| Compose and send |
| Save a draft in Mail |
| Change a draft Mail already holds — see Edit a draft |
| Send a draft Mail already holds |
| Reply, staying in the thread, with attachments if any; |
Confirmation is mandatory. Every tool that actually sends — send_email,
send_draft_file, send_draft and reply_to_message(send=True) — does nothing
unless confirm=true. Called without it they return confirmation_required
along with a preview block describing precisely what would go out: sender,
recipients, subject, body, and the attachments actually carried. It doubles as a
dry run.
The guard covers every send path rather than one of them: protecting only the
draft path would push a caller to recompose with send_email, which is the
behaviour worth avoiding in the first place.
to, cc and bcc accept one address, a comma-separated string, or a list.
attachments takes absolute paths to existing files, checked before anything is
built. Without sender, the account Mail lists first is used — worth being
explicit when several accounts coexist, since the sender decides which signature
is appended.
Every message goes out as HTML, laid out the way Mail lays out its own: the
message, then the account's signature, then a blank line, then the attachments.
A plain text body is converted — blank lines become paragraphs, single newlines
become breaks, and anything resembling markup is escaped — so a caller never has
to write HTML to get a well-formed message. Pass signature=false to leave the
signature off.
The signature is not configured here. It is read from Mail's own settings for the sending account: which signature is selected, its HTML, and the images it carries. Editing it in Mail is enough; there is no second copy to keep in step.
Edit a draft
edit_draft (a draft in Mail) and edit_draft_file (an .eml draft) take the
same changes, any of them in one call:
Argument | Change |
| New subject line |
| The whole new list for that field; an empty string clears Cc or Bcc |
| Addresses added to that field |
| Addresses taken off, whichever field holds them |
| New text; the signature and a quoted original below it stay |
| Targeted edits, |
| Absolute paths of files to attach |
| File names of attachments to take off |
The draft is edited where it stands, never composed again: only the headers asked about and the text are rewritten, and the formatting, the signature and its logo, the quoted original, the attachments and the thread headers stay byte for byte. An address placed in a field leaves the other two, so adding to To someone who was in Cc moves them. Every change is checked before anything is written; a refused one leaves the draft untouched.
A server cannot change a message in place, so edit_draft files the edited
version first and only then expunges the previous one: a failure in between
leaves two drafts, never none. The draft therefore gets a new message_id,
which list_messages(mailbox="drafts") gives once Mail has checked the account.
Close the draft if it is open in a Mail window, or Mail may save its own copy
over the edit.
Organise
Tool | Purpose |
| Create a mailbox, optionally nested |
| Move to another mailbox |
| Move to trash |
| Read status |
| red, orange, yellow, green, blue, purple, gray, or |
Drafts are files, not Mail drafts
write_draft writes a self-contained .eml file — attachments embedded — into
mails/. macOS renders an .eml in Mail on double-click, so it reads like a
real message. send_draft_file submits that file as it stands, so what
leaves is byte for byte what was reviewed, then moves it to mails/sent/.
This is not a stylistic choice. Mail cannot send a draft it holds. Its
send command only understands an outgoing message, not a message sitting in a
mailbox; opening a draft turns it into one, but only after an unpredictable
delay that exceeded a minute in testing; moving it to the Outbox does nothing at
all. Anything drafted inside Mail therefore has to be re-posted and the original
deleted — and on a Gmail account that delete is undone by the server unless it
is issued once the send has settled.
Keeping drafts out of Mail removes the problem rather than working around it. Sending becomes a single instant operation with nothing to clean up afterwards.
send_draft remains for drafts that already sit in Mail, whether written by
hand or filed there by create_draft. It fetches the message from the server,
submits it unchanged and expunges the draft — so the formatting, the signature
and every attachment survive, and the deleted draft does not come back.
Retention. An unsent draft is removed after 7 days, an archived one after
30. The sweep runs on every write, every listing, and from sync_index, so a
forgotten draft does not sit on disk indefinitely. Files live in
mcp-mail-macos/mails/ unless folder says otherwise, and are gitignored:
they hold real message content.
The search index
Why
Mail answers message by message. On a mailbox of around 20,000 messages, reading
metadata costs about 0.65 s per message and reading a body about 1.7 s. Mail's own search
(whose subject contains …) takes about 21 s over 2,500 messages, and searching
bodies exceeds 120 s — to the point of leaving Mail unresponsive to every
subsequent call for minutes.
Searching an archive of tens of thousands of messages that way would take hours. The index sidesteps it by reading Mail's storage directly.
What feeds it
Two sources, neither sufficient alone:
MailData/Envelope Index, Mail's internal SQLite database, for metadata, mailbox membership, and read and flag status. It is copied — together with its write-ahead log — then opened read-only.The
.emlxfiles, for body text and the RFCMessage-IDheader.
Mail's index holds no full text; the files do not say which mailboxes a message belongs to.
Indexed, not stored
Bodies go into an FTS5 table declared content='': searchable, never kept. The
database stores only what is needed to display a result and act on it — subject,
sender, date, Message-ID, locations. Reading a message goes back through
get_message. For roughly 50,000 messages the index weighs about 80 MB.
messages (id, account, rfc_id, subject, sender, date_received, size, conversation_id,
body_indexed, has_attachment, list_id, is_bulk)
locations (message, account, mailbox, read, flagged)
recipients (message, kind, address, domain, name) -- kind is 'to' or 'cc', address lower-cased
messages_fts(subject, sender, "to", cc, attachments, body) -- FTS5, content=''Splitting message from locations absorbs Gmail's duplication: a message exists once on disk, in All Mail, and labels are only views. A mailbox of some 50,000 distinct messages yields around 135,000 locations — which is exactly the figure AppleScript reports when its mailboxes are summed.
The durable key is the RFC Message-ID, not Mail's internal id, which changes
whenever a message moves.
Build and update
python3 mail_index.py --check # verify assumptions, build nothing
python3 mail_index.py --build # full backfill
python3 mail_index.py --sync # incremental
python3 mail_index.py --search "invoice acme"--check validates seven points, including that message ids map to files and —
most importantly — that membership rebuilt from both sources matches Mail's own
per-mailbox counts. --build refuses to start if any of them fails:
Envelope Index is undocumented and changes between macOS releases, so a clean
refusal beats a silently wrong index.
--sync diffs the set of messages Mail lists against the set the index holds.
Deletion is not a special case, and a move reads as a change of location at
constant Message-ID. A pass with nothing to do costs about two seconds.
Body coverage
Some messages end up with no searchable body: the file is missing, only a
partial download exists, or the message is empty. messages.body_indexed is
1 when a non-trivial body (at least a few word characters) was extracted and
0 otherwise, and index_status reports with_body and without_body for
each account and overall. Those messages still match on subject, sender,
recipients and attachment names.
For multipart/alternative messages the text/plain part is used, unless it
is empty, near-empty or a "this message contains HTML" placeholder: then the
HTML part is stripped of markup and indexed instead.
Quoted history is cut from the indexed body, so a reply no longer competes with
the messages it quotes. Only safe patterns count: lines starting with > (with
the On ... wrote: / Le ... a écrit : line that introduces them, and the
non-quoted lines after the block are kept, for bottom-posted and interleaved
replies); Outlook From:/De : header blocks (at least three header lines) and
-----Original Message----- markers, which cut what follows; the last --
signature line when the block after it is 15 lines or fewer; in HTML,
blockquote type=cite, Gmail quote containers and Outlook's reply header. If
fewer than 20 word characters would remain, the full text is kept, and a
forward keeps its content when the text above it is tiny. Bodies indexed before
this need a rebuild (python3 mail_index.py --build).
French stemming
Search understands French plurals, feminines and common verb forms: facture
finds factures, relancé finds relance and relancer, travail finds
travaux. Accents and case are ignored as before. Under the hood every text
column (subject, attachment names, body) is indexed twice, as written and as a
light stem (mail_stem.py, pure Python, no dependency), and a query word is
searched in both. A message holding the exact word ranks above one holding only
another form. Things to know:
Quotes mean exact.
"facture"and"les factures"match the words as written, in that order, with no stemming.fact*matches word starts in both forms. Sender, To and Cc are never stemmed, so names and addresses stay exact.Numbers, codes and words under four letters are left alone. The stemmer is deliberately light: it will not join a noun and its verb (
paiement,payer).The stems change the index content: an index from before needs
python3 mail_index.py --build(schema version 5, about 25 % larger).
Schema version
The index carries a schema_version in its meta table (an index that
predates versioning counts as version 1). When the code expects a newer one,
search_all, sync_index and mail_index.py --sync refuse with an
index_outdated error and the hint to run python3 mail_index.py --build;
they never mix formats. The version also moves when the indexed content changes,
not only the tables: version 5 (stemmed columns next to the raw ones) needs a
rebuild from version 4; version 4 (quoted history cut from bodies) needed one
from version 3; version 3 (To and Cc kept apart, has_attachment, list_id,
is_bulk) needed one from version 2. --build always starts from scratch, writing to
<index>.building and swapping the finished file in atomically, so the live
index keeps answering until the new one is ready and a crashed build loses
nothing. --sync, --build and the automatic sync share one lock file
(<index>.sync.lock, holding the owner's PID and refreshed while it works), so
they never run at the same time.
Freshness
search_all checks the index's age and runs the sync itself past
max_age_minutes (10 by default). The answer carries index_age_minutes and
synced, so the caller knows what was searched. If disk access was revoked in
the meantime the search still succeeds against the existing index and says so in
sync_note rather than failing — a stale result beats an error. A lock prevents
two concurrent syncs.
Background sync
launchd/com.mcp-mail-macos.sync.plist runs --sync every ten minutes,
independently of any client:
cp launchd/com.mcp-mail-macos.sync.plist ~/Library/LaunchAgents/
launchctl load ~/Library/LaunchAgents/com.mcp-mail-macos.sync.plistA warning before installing it: the agent reads ~/Library/Mail, so Full Disk
Access has to be granted to the program it runs, /usr/bin/python3. That hands
the grant to every Python script on the machine — wider than an app-scoped
one. A venv interpreter is no better: its path carries a version number and the
grant breaks on the first upgrade.
The agent is only worth it if the index must stay current with no client
running. Otherwise search_all's own freshness check is enough, and the grant
stays scoped to a single application.
Attachment text
Attachment names are always indexed. Their text lives in a second database,
attachments.sqlite, built by mail_attachments.py. It is separate on purpose:
it is large (text of tens of thousands of files), it is rebuilt on its own
schedule, and search works exactly as before when it is absent or broken.
python3 mail_attachments.py --sync # first run reads everything; later ones only what is new
python3 mail_attachments.py --build # start from scratch
python3 mail_attachments.py --status # rows by status, size, last run
python3 mail_attachments.py --measure # time the extractors on a random sampleWhere the files are: Mail keeps the attachments of a message it has not fully
downloaded (.partial.emlx) as plain files in
<mailbox>.mbox/<store>/Data/<shard>/Attachments/<id>/<part>/<file>, and those of
a full .emlx inside its MIME. Both are read; a MIME part is written to a private
temporary directory (mkdtemp, mode 0700, outside the repository) for the length
of one batch and removed after it, on exit and on SIGTERM. Nothing is written
under ~/Library/Mail, and Full Disk Access is needed as for the index.
What is read, and how:
Kind | Method | Rule |
PDFKit text layer ( | up to 50 pages, | |
Scanned PDF (under 20 characters of text) | Vision OCR, fr-FR + en-US, accurate ( | first 3 pages |
docx, xlsx, xlsm, pptx | zip + XML (numbers are left out of sheets) | half of |
doc |
| same |
png, jpg, jpeg, heic, tiff | Vision OCR | at least 50 KB, name not a logo, banner, signature, social icon or |
zip, rar, dwg, audio, video, gif, xls, anything else | skipped | recorded with the reason |
The Swift tools are compiled on first use into tools/build/ (gitignored) with
/usr/bin/swiftc, and again whenever their source is newer; a missing compiler
is reported with xcode-select --install. Each takes many files per call, so the
process start is paid once. A tool that stays silent for 40 s is killed; the file
it choked on is marked error and the rest of the batch goes again.
Resumable: discovery first records every attachment as a row (pending, or
skipped / too_big with a reason), then the pending rows are read in batches of
40 with a commit each. Kill the run at any point and run --sync again: it picks up
the remaining rows. A row is read again when its size or modification time
changes (--retry-errors retries failures); rows of messages that left Mail, and
of files that disappeared, are deleted. Changing a setting (OCR on or off)
reclassifies the skipped rows on the next run. A run takes attachments.sqlite.sync.lock,
independent of the index lock: the two syncs can run at once, since the
attachment one reads Mail's files, not the index.
Column | Meaning |
|
|
| why skipped or failed: |
|
|
| the extracted text, one line, capped at |
attachments_fts(filename, text, filename_stem, text_stem) is contentless like the
message table, with the same French stemming; the text itself stays in
attachments.text for snippets.
How search uses it: search_all runs the free text against the message index as
before, then against attachments_fts (bm25 with the file name counting three
times the text), and merges by message. A message found by both gets its own
recency-adjusted score plus ATTACHMENT_WEIGHT (0.5) times the attachment's; a
message found only through an attachment is ranked by that half score. Attachment
hits go through the same filters as any result (operators, dates, account,
mailbox, filename:), so an attachment never bypasses them. A query restricted
to a message field (subject:, sender:...) leaves attachments out;
sort="date" orders the merged set by date. has:attachment is unchanged and
filename: still matches the names in the message index.
Keeping it current: a search never waits for attachments (reading them takes
minutes). Instead, once attachments.sqlite exists, every successful
sync_index (and so every stale-index refresh from search_all) starts
mail_attachments.py --sync detached in the background; if one is already
running the new one exits at once. Set attachments_auto_sync to false to
turn that off and run it yourself, or from launchd next to the message sync
(copy launchd/com.mcp-mail-macos.sync.plist, replace its --sync program by
mail_attachments.py --sync, and the same Full Disk Access caveat applies).
index_status reports attachments: rows by status, size and last run.
Measured on a mailbox of about 54,000 messages and 34,000 attachment files on disk (plus about 13,500 named parts inside full messages), on Apple silicon: PDF text about 30 ms a file, scanned PDF OCR about 350 ms, Word and Excel a few ms, images about 60 ms. A full first run takes on the order of half an hour.
Search by meaning
Keyword search finds the words you typed. search_all(..., mode="semantic")
finds the messages that are about what you typed ("unpaid bills" finds a
"payment reminder"), and mode="hybrid" fuses both rankings. It is optional:
without the vectors file, Ollama or sqlite-vec, everything above behaves
exactly as before.
| What runs |
| The search described above (bm25, attachments, stemming). |
| The query is embedded and compared with the message chunks; no keyword involved. Fails with a hint when it cannot run. |
| Both, then Reciprocal Rank Fusion ( |
omitted | The |
Each result of semantic and hybrid carries match (keyword, semantic or
both) and, when meaning found it, similarity (cosine of its best chunk). A
result found only by meaning has no matched word to anchor a snippet on, so its
snippet is the passage that matched. Everything that narrows a keyword search
narrows this one too: operators, dates, account, mailbox. The filters are
applied before the nearest-neighbour cut, so a narrow filter never loses its hits
to the global top. sort="date" reorders the fused best matches newest first.
With free text absent (operators alone) there is nothing to embed and the search
is a keyword one.
Building the vectors, once, then keeping them current:
brew install ollama && brew services start ollama && ollama pull bge-m3
.venv/bin/pip install -r requirements-semantic.txt
python3 mail_vectors.py --build # from scratch; --sync resumes and only does what is new
python3 mail_vectors.py --statusWhat is embedded: the subject and the message's own text (quoted history already
cut, links and tokens over 40 characters dropped), in chunks of about 1,000
characters with 150 overlapping, at most 8 per message, so a long thread is
represented by its first 8,000 characters or so. The subject is added to the first
chunk. A message with no readable body is embedded by its subject alone. The file
holds the vector, the chunk number and a hash of the text, never the text: an
excerpt is cut again from the message when needed. --sync skips a message
already embedded without reading it (--verify reads them again and re-embeds
those whose text changed), deletes the vectors of messages the index dropped,
commits every batch of 32 chunks (an interrupted run, SIGTERM included, resumes
where it stopped) and takes vectors.sqlite.sync.lock, independent of the
index's. --sample PERCENT --seed N embeds a reproducible random subset, for
measurements. Once the file exists, every successful sync_index starts
--sync in the background (vectors_auto_sync); a search never waits for it.
Measured on a mailbox of about 50,700 messages (Apple silicon, bge-m3 in Ollama):
Chunks | 70,946: 1.4 per message; 91 % of messages take one or two, 1 % reach the cap of 8, 141 empty ones none |
Full build | 36 minutes: about 33 chunks a second whatever the batch size from 4 up (16 to 64 measured), 32 per request; the .emlx reading (about 2 minutes for the whole mailbox) is negligible next to it |
File size | 96 MB (int8, 1 KB a chunk); float32 would be about 4 times larger |
Quantisation | int8, one scale per vector. Against float32, 99.3 % of the top 10 neighbours are kept; a binary quantisation (128 bytes a chunk) keeps 66 % and was rejected |
Query latency | embedding a query 11 ms once the model is loaded (0.8 s when Ollama has to load it again, after 30 minutes idle); |
The vectors are plain int8 blobs in an ordinary table compared with sqlite-vec's
scalar functions, not a vec0 virtual table: vec0 does the same exact scan (95 ms
against 89 ms for 125,000 chunks), but its contents cannot be read without the
extension and it cannot be restricted by message id, both of which the Python
fallback and the filters need. The embedding call has a 5 second timeout
(ollama_timeout); after a failure Ollama is not asked again for a minute, so a
search never pays the timeout twice, and hybrid answers with keywords and says
why in semantic_note.
How well it works, on the pairs of Evaluating search quality
(550 pairs, mail_eval.py --run --mode ...; MRR / recall@10). These pairs are two or
three words of a message, so they measure exact words, which keyword search is built
for; semantic search is not expected to win there, and attachment text is not embedded.
Mode | Pairs as drawn | Query words inflected ( |
| 0.467 / 0.747 | 0.330 / 0.547 |
| 0.102 / 0.204 | 0.092 / 0.173 |
| 0.440 / 0.744 | 0.358 / 0.593 |
Hybrid uses the semantic ranking at a quarter of the weight of the keyword one
(RRF_SEMANTIC_WEIGHT). With equal weights, the loose neighbours of a short query
pushed exact matches down (MRR 0.410 on the pairs as drawn). Weights from 1 to 0.15,
a minimum similarity and a cap on the semantic list were tried on the same pairs;
0.25 without threshold kept recall@10 level with keyword and gained on inflected
queries (subject pairs: recall@10 0.465 to 0.580, 40 fewer misses out of 200). The
price is a lower first place on exact words (MRR minus 0.026), which is why
keyword stays the default: it is right for the most frequent kind of query, needs
no Ollama and answers in 24 ms. Search by meaning is where words differ, for
instance a query in English finds French mail about "unpaid bills" (a keyword search
finds nothing); use mode="hybrid" or "semantic" for those, or set
search_mode to auto if you prefer it everywhere.
Attachments are not embedded in this version. Their text is the largest volume
of the mailbox (a PDF alone can hold a hundred thousand characters), it is often
noise for a language model (letterheads, tables, OCR of stamps), and keyword
search already reaches it: in hybrid the keyword side still merges attachment
hits, so a message found through an attachment word stays in the fused list.
Similar messages
find_similar(message_id) returns the messages whose meaning is closest to a given
one. It needs the vectors file but neither Ollama nor the network: the query is the
message's own stored chunk vectors, averaged (each brought to unit length first). The
message itself is never returned, and with exclude_thread=True (the default) neither
is the rest of its conversation, so the answer is other mail about the subject; pass
false to keep the replies. message_id is a message_id reference or the numeric
mail_id of search_all. query narrows the candidates like search_all's: Gmail
operators, dates, account and mailbox filter them before the nearest-neighbour cut,
and plain words must appear in the message. Results have the shape of search_all's,
with score (cosine of the best chunk) and the matching passage as snippet. A
message without vectors (newer than the last mail_vectors.py --sync, or no text)
is an error saying so. On the 71,000 chunks above a search takes about 80 ms (100 ms with snippets),
almost all of it the exact comparison; a filter that leaves fewer chunks is faster.
Averaging the chunks was chosen against keeping the best score over each chunk as a query, on 12 random multi-chunk messages (top 10 each, thread excluded): same sender in 61 and 69 of 120, a subject word in common in 88 and 83 of 120, against 0 and 12 for random messages, at 68 ms against 168 ms. The two are as on-topic; the average costs one scan instead of up to eight.
Evaluating search quality
mail_eval.py measures where search puts the message you were looking for, so a
change to ranking or matching can be shown to help rather than just to differ.
It opens the index read-only and never syncs during a run.
python3 mail_eval.py --generate 200 --seed 1 # pairs from random messages
python3 mail_eval.py --attach 100 --seed 3 # add pairs from attachment text (keeps the other pairs)
python3 mail_eval.py --run # rank of each expected message
python3 mail_eval.py --run --save-baseline before # keep the numbers
python3 mail_eval.py --run --compare before # deltas, and pairs that got worse
python3 mail_eval.py --run --mode hybrid # keyword, semantic or hybrid (needs the vectors)A pair is a query and the message it should find. --generate derives the query
from two to four distinctive words of a random message's subject; the body is
not stored in the index, so it is not used. Pairs written by hand
("source": "manual") are kept when the automatic ones are regenerated;
eval.example.json shows the format with made-up data. The report gives MRR,
recall@1, recall@10 and the number of pairs not found in the top 50. Add --json
for machine output.
--attach N adds pairs of a third kind, attachment: two or three words that are
rare in one attachment's text (and not in its message's own fields), expected
result the message. Alone it redraws only the attachment pairs and keeps every
other one, so the subject and body numbers stay comparable across runs; the
report gives one line per kind.
Pairs and results are written under eval/, which is gitignored: they contain
real subjects.
Message identifiers
Every message carries an opaque message_id encoding the account, the mailbox
path and Mail's internal id. It is stable between calls and survives a Mail
restart — it is not a position in a list.
Two caveats. The internal id is only unique within a mailbox, hence the account
and path travelling with it. And a move creates a new one: move_message
returns the new message_id when it can find the moved copy through its
Message-ID header, and says so when it cannot.
A reference from search_all or get_thread has the same shape and works
directly in get_message, reply_to_message or move_message. It points at
the smallest mailbox holding the message, because Mail resolves an id by
walking the mailbox it is given: aiming at a folder of a few thousand messages
rather than one holding tens of thousands changes the response time by an order
of magnitude.
Response format
Every function returns a dictionary. On failure:
{
"ok": false,
"error_code": "mailbox_not_found",
"error": "mailbox not found: Drafts (account Work)",
"hint": "Call list_mailboxes to see the exact mailbox paths."
}Errors are data, not protocol exceptions, so a caller can correct itself from the code and the hint.
Code | Cause |
| macOS refuses control of Mail, or access to its storage |
| Mail did not answer in time |
| Mail is closed and could not be started |
| Target not found |
| Malformed identifier |
| Attachment missing from disk |
| A send was requested without |
|
|
| Attachments could not be recovered; nothing was sent |
|
|
| An edit's |
| An edit removes someone absent, would leave no recipient, or names something that is not an address |
| An edit removes a file the draft does not carry, or changes nothing |
| Index absent, incomplete, or not refreshable |
Known limitations
All of these come from Mail, not from this server. The figures were measured on an M4 Pro MacBook Pro against real accounts.
Speed
Operation | ~2,500-message mailbox | ~20,000-message mailbox |
Metadata for 20 messages | ~1 s | ~13 s |
Metadata for 200 messages | ~13 s | — |
One message body | ~1.7 s | ~1.7 s |
A mailbox's | instant | instant |
This dictates the defaults: include_preview and include_totals are off, a
single list_messages call reads at most twenty previews however many messages
it returns and says so in the answer, and search goes through the index rather
than through Mail.
What Mail cannot do at all
Send a draft it holds.
sendonly understands an outgoing message (-1708). Opening the draft produces one only after an unpredictable delay, sometimes over a minute. Moving it to the Outbox does nothing. The only faithful alternative reported by the community is GUI scripting (Cmd+Shift+Dthrough System Events), which needs Accessibility permission and breaks with any interface change — deliberately not taken here. So a send does not go through Mail at all: the message is submitted over SMTP, through the account's own outgoing server.Compose a message without rewriting it. Setting
html contenton an outgoing message makes Mail wrap the body in its share wrapper — a stray<br>above the first line, inside a<blockquote type="cite">— and an attachment made throughmake new attachmentlands inside that body, ahead of the signature. Neither is reachable from AppleScript, and both survive every ordering of the calls, a full HTML document, and setting the property aftersave. So messages are built as MIME here and handed to the server.Delete a mailbox.
delete mailboxfails with -10000 whatever the syntax. A mailbox created bycreate_mailboxhas to be removed by hand.Export an attachment. Mail refuses to write the file anywhere (-10004). This no longer matters for sending:
send_draftfetches the whole message from the server and submits it unchanged, so the attachments never have to be read back out of Mail's storage.Set headers on an outgoing message.
In-Reply-ToandReferencescannot be set on a message Mail composes, which used to forcereply_to_messagethrough Mail's ownreplycommand — opening a compose window and rewriting the body. Building the message here sets them directly instead.Create a mailbox with an
accountproperty. It has to happen inside atellblock targeting the account, or -10000.Delete a draft for good. A Gmail account pushes back a draft deleted through Mail a few seconds after the delete reports success. Expunged on the server, it is gone — which is how
send_draftremoves it.
Behaviours worth knowing
Mail never composes here any more, so the autosaves it used to leave behind after a send, and the outgoing messages that accumulated invisibly in its internal list, no longer happen at all. The sweep that hunted them down is gone with them.
Mail counts a signature image among the attachments, so a preview built from Mail's own list announces an image the recipient never receives as a file — and lists nothing at all before Mail has downloaded the parts. What a preview describes is therefore the message on the server, with each part settled by its disposition: an attachment is offered, an inline part belongs to the body. For a message written elsewhere that says neither, a part the HTML shows with
<img src="cid:…">is taken as part of the body. What was left behind is reported underkept_inline.Gmail labels are mailboxes, and one message appears in several.
INBOXcan resolve to All Mail: a message'smailboxfield reports where Mail sees it, which is not always what was queried.every mailbox of accountreturns leaf names, but lookup by slash-separated path works. The server rebuilds full paths by walking thecontainerproperty.An attachment's name is sometimes inconsistent between calls; its size is reliable.
A disabled account disappears from Mail's list without an error.
AppleScript calls are wrapped in an explicit
with timeout; without it any call over 60 s fails, which a large mailbox reaches easily.Numbers and dates coerced to text follow the machine's locale — a date becomes
1,785863539E+9. The server assembles ISO 8601 dates digit by digit to avoid it.
Message content reaches the client unfiltered
Everything these tools return — bodies, subjects, sender names, attachment names — is whatever arrived in the mailbox, passed through untouched. A message can therefore contain text that reads like an instruction, and an agent consuming this server will see it alongside its own. Treat mail content as data, never as direction, and be wary of a tool call whose arguments were lifted verbatim from a message. This is not specific to this server, but it is worth stating: reading mail on an agent's behalf is exactly the situation prompt injection targets.
read_draft_file takes a path and parses whatever is there as an email, so any
readable file on the machine can be turned into a body and handed back. That is
deliberate — folder would be pointless otherwise, and attachments already
require arbitrary paths — but it means the server is as trusted as the client
driving it. It is meant to run locally, for one user.
Scope
The server exposes every account Mail knows about, for reading and writing
alike. There is no account allowlist. Adding one means filtering in two places —
MessageReference.decode and resolveMailbox on the AppleScript side, plus a
WHERE account IN (...) on the index — because the AppleScript tools reach Mail
directly and would otherwise still see everything.
The index reflects Mail's local store. What an account has not synced does not exist for Mail, and therefore not for search either.
Testing
Unit tests cover everything that does not need Mail: identifier encoding,
address parsing, error classification, AppleScript assembly, .eml round-trips,
retention, message layout and reply threading. They run anywhere, in under a
second:
python3 -m unittest discover -s tests -t .Manual checks exercise the live path against a real Mail install:
.venv/bin/python test_manual.py read # read-only
.venv/bin/python test_manual.py read --account Work --mailbox INBOX
.venv/bin/python test_manual.py write --to you@example.com # draft + mailbox
.venv/bin/python test_manual.py write --to you@example.com --send # really sendsread changes nothing: eight checks, two of which verify that errors surface
cleanly. write creates a draft and a test mailbox in Mail; the mailbox has to
be deleted by hand, since Mail cannot do it through AppleScript.
Project layout
mcp-mail-macos/
├── server.py # MCP entry point, the 29 tool definitions
├── mail_tools.py # driving Mail through AppleScript
├── mail_message.py # building the message: body, signature, attachments
├── mail_signature.py # the signature Mail would have used, from its settings
├── mail_draft.py # drafting and sending, on top of the two above
├── mail_imap.py # the account's own server, for filing and sending
├── mail_files.py # .eml drafts, retention, leftover sweep
├── mail_edit.py # editing an existing draft in place, in Mail or as .eml
├── mail_search.py # querying the index
├── mail_saved.py # named searches (saved_search tool)
├── mail_index.py # building and updating the index
├── mail_stem.py # French light stemmer and query rewrite
├── mail_attachments.py # attachment text: extractors, attachments.sqlite, sync
├── mail_vectors.py # embeddings for search by meaning: vectors.sqlite, Ollama client, sync
├── mail_eval.py # search relevance evaluation (pairs, MRR, recall)
├── test_manual.py # manual checks against a real Mail install
├── tools/ # Swift sources (PDFKit text, Vision OCR), compiled into tools/build/
├── tests/ # unit tests, no Mail required
├── applescript/ # one script per operation, plus shared handlers
│ ├── _common.applescript
│ └── …
├── launchd/ # optional periodic sync agent
├── requirements.txt
├── requirements-semantic.txt # optional: sqlite-vec
└── README.mdAppleScript files are assembled at run time: _common.applescript is prepended
to each script, and a with timeout wrapper is added around the run handler.
Parameters travel through argv rather than string interpolation, which rules
out injection, and -- protects values starting with a dash. Results are
serialised with ASCII separators 31 and 30, which never appear in real mail and
are stripped from values before joining — hence no escaping when parsing.
License
MIT. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Your IMAP mailbox as an MCP server: read, search and (if you allow it) organize mail. Open source.
Your mailbox for MCP clients: search, read, draft, send, rules and notes. Sending is off by default.
Email infrastructure for AI agents — send, receive, search, and reply to email over MCP.
MCP server for Nylas — read email, calendars, events and contacts, and send email or create events.
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server that gives Claude and other MCP hosts full access to Mail.app on macOS — search, read, send, reply, flag, move, and more across all accounts configured in Mail.app.2453 npm1MIT
- AlicenseAqualityDmaintenanceLocal MCP server for multi-account IMAP/SMTP email (iCloud + Gmail via app-specific passwords). Never marks mail read. Cross-folder search, idempotent sends, TLS verified.8MIT
- AlicenseNot gradedqualityDmaintenanceLocal MCP server for macOS Mail reads plus visible unsent compose, reply, and forward drafts, and constrained single-message moves.MIT
- AlicenseBqualityCmaintenanceLocal MCP server for macOS native apps: Mail, Calendar, Reminders, Notes, Messages, and Contacts. Enables reading and organizing your Mac life through a single stdio process using AppleScript/JXA.408 npmMIT