Skip to main content
Glama
beeraw

mcp-mail-macos

by beeraw

mcp-mail-macos

An MCP server that drives macOS Mail: read, search, send, organise. It runs over stdio, launched on demand by the MCP client — there is no long-lived process.

Two mechanisms live side by side, deliberately. Actions go through AppleScript, the only interface that can make Mail do anything. Search goes through a local SQLite index built from Mail's own storage, because AppleScript needs seconds per message and cannot search an archive of tens of thousands of mails in any usable time.

Tested on macOS 27, Python 3.14, Mail 16, against Gmail, IMAP and Exchange accounts, on a mailbox of roughly 50,000 messages spanning several years.

Read this before installing. This server grants an agent the right to read, send, move and delete mail on every account Mail is configured with, and the search index needs Full Disk Access, which macOS cannot scope to a single folder. See Before you trust it with your mail.


Contents


Related MCP server: email-mcp

Before you trust it with your mail

This is a local tool for one person on their own machine. It is not a service, and it was not designed to be exposed to several users. What it asks for is broad, and worth weighing before installing.

Automation lets an agent act on every account. One grant, once, and the server can read, send, reply, move and delete across all of them. There is no per-account allowlist — adding one means filtering in two places (see Scope).

Full Disk Access is all or nothing. The index reads ~/Library/Mail, which macOS protects; there is no setting scoped to that folder. Granting it also grants Messages, browser history and other applications' data to whatever application you granted it to, and for every future session until you revoke it.

Message content reaches the agent unfiltered. Anyone can send you mail, and that mail lands in an agent's context as text. That is the classic prompt injection setup, and no permission dialog stands between the two.

Reasonable precautions, in rough order of value:

  1. Try it on a secondary account first, before pointing it at anything that matters.

  2. Keep the confirmation guard. Every send requires confirm=true and returns a preview otherwise. It makes each send deliberate and shows exactly what would leave.

  3. Only grant Full Disk Access if you need indexed search, and revoke it afterwards — everything already indexed stays searchable. Set index_max_age_minutes high in config.json so the server stops trying to refresh.

  4. Decide whether an agent should send at all. Preparing drafts as .eml files and sending them yourself is a perfectly good mode; write_draft touches nothing but a folder.

  5. Run the checks before real use: python3 -m unittest discover -s tests -t . for the logic, then test_manual.py read against your own Mail.


Requirements

OS

macOS, with Mail configured and its accounts loaded. AppleScript and Mail's storage layout are the whole foundation, so there is no path to another platform.

Python

3.11 or later — the code uses X | None annotations and tomllib-era stdlib behaviour. Developed on 3.14.

Mail

Version 16 (macOS 13+). The AppleScript dictionary has been stable across these releases; the internal index schema has not (see below).

Dependencies

One, declared in requirements.txt:

mcp>=1.2.0

That is the official Model Context Protocol SDK. Both generations work and the import picks whichever is installed:

SDK

Class

Import

1.x

FastMCP

mcp.server.fastmcp

2.x

MCPServer

mcp.server.mcpserver

The decorator API is identical between the two, so nothing else changes.

One more, optional, in requirements-semantic.txt:

sqlite-vec>=0.1.6

It computes cosine similarity inside SQLite for search by meaning (below), and is imported only when that feature runs. Without it the server, the index and keyword search are untouched; semantic search then compares vectors in plain Python, which is fine for a filtered search and refused (with a hint) for a whole-mailbox one. Search by meaning also needs Ollama running locally with the bge-m3 model (brew install ollama, brew services start ollama, ollama pull bge-m3, about 1.2 GB); it is reached over HTTP with urllib, nothing to install in Python.

Everything else is standard library — sqlite3 for the index and its FTS5 tables, email for parsing .emlx containers and writing .eml drafts, subprocess for osascript, unicodedata, urllib.parse, json, tempfile. No compiled extension, no build step.

Running the unit tests needs nothing at all beyond the standard library: they never import the SDK.


Install

git clone https://github.com/beeraw/mcp-mail-macos.git
cd mcp-mail-macos
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/pip install -r requirements-semantic.txt   # optional: search by meaning

macOS permissions

Two separate grants, for two different needs. The first is required. The second only concerns indexed search.

1. Automation — driving Mail

On the first call, macOS asks for permission to control Mail. The dialog appears once, and it is attributed to the application launching the server, not to Python.

If it was denied it will not ask again: restore it in System Settings → Privacy & Security → Automation, unfold the application concerned and tick Mail. Until then every tool returns permission_denied with that reminder.

To trigger the prompt at a quiet moment, before wiring anything up:

.venv/bin/python test_manual.py read

2. Full Disk Access — building the index

The indexer reads ~/Library/Mail, which macOS protects through TCC. There is no setting scoped to that folder: the only lever is Full Disk Access, all or nothing, in System Settings → Privacy & Security.

The grant goes to the process's responsible application. For Claude Code that is /Applications/Claude.app — not the nested claude-code binary, and not Python. It is only read at launch, so the application has to be quit and restarted.

Granted to

Consequence

/Applications/Claude.app

The index refreshes itself from the MCP server. In exchange the grant covers every protected location, not just Mail, and applies to future sessions.

Terminal only

The server can search the index but not refresh it. sync_index returns permission_denied with instructions; updates happen by hand or through a launchd agent.

Revoking it later breaks nothing: everything already indexed stays searchable, only updates stop.

3. An account password — writing drafts and sending

Reading mail goes through Mail. Writing one does not: a draft is filed on the server over IMAP, and a send is submitted over SMTP. Mail holds the host, the port and the user name for both, so the only thing missing is a password it will not hand out.

One is stored per account in the login keychain, under the service name in keychain_service, keyed by the account's own user name:

security add-generic-password -U -s mcp-mail-macos -a you@example.com -w 'the password'

On an account with two step verification — every Google account, for one — that must be an app password, not the account password. Google offers them at myaccount.google.com/apppasswords. If the page answers that the setting is unavailable, two step verification is off on that account, or a Workspace administrator has disabled app passwords.

The lookup runs through /usr/bin/security, the same tool that stored the password, so the keychain grants it without a prompt. A password added by hand through Keychain Access is refused until it is allowed once.

An account with no IMAP server — an Exchange or Outlook account — cannot be written to this way, and says so rather than failing obscurely.


Add to Claude Code

Absolute paths, since the server can be launched from anywhere:

claude mcp add mail-macos -s user -- /path/to/mcp-mail-macos/.venv/bin/python /path/to/mcp-mail-macos/server.py

-s user makes it available in every project; without the flag it stays scoped to the current one. Check with claude mcp list. The tools appear once Claude Code restarts.


Configuration

Nothing has to be configured: every setting falls back to something that works out of the box, and drafts and the index stay inside the repository directory.

To change any of it, copy the example and edit what you need:

cp config.example.json config.json

config.json is gitignored, so local paths never end up in a commit. Any key may be omitted. An environment variable of the form MAIL_MCP_<KEY> overrides both the file and the default — convenient when the MCP client passes its own configuration:

claude mcp add mail-macos -s user -e MAIL_MCP_DRAFTS_FOLDER="$HOME/Documents/Outgoing mail" -- /path/to/.venv/bin/python /path/to/server.py

Key

Default

What it does

drafts_folder

mails/ in the repo

Where .eml drafts are written, and where sent ones are filed

pending_retention_days

7

How long an unsent draft may sit on disk

archive_retention_days

30

How long a sent draft stays archived

reply_attribution

On {date}, {sender} wrote:

The line above the quoted original in a reply, in the language of the person answered

reply_date_format

%Y-%m-%d %H:%M

How {date} above is written out, in strftime terms

keychain_service

mcp-mail-macos

Keychain service name the account passwords are stored under

index_path

index.sqlite in the repo

The search index

mail_root

~/Library/Mail

Mail's storage, where the index is built from

index_max_age_minutes

10

Past this age, search_all refreshes the index before answering

applescript_timeout

120

Ceiling for a read call, in seconds

applescript_write_timeout

180

Ceiling for a send or move

body_limit

200000

Characters of body kept per message when indexing

attachments_path

beside index_path

The attachment text index, attachments.sqlite

attachments_ocr_images

true

OCR of loose images (at least 50 KB, not named like a logo); scanned PDFs are always read

attachments_max_mb

20

Largest attachment read: PDFs up to this, other types up to half

attachments_char_limit

100000

Characters kept per attachment

attachments_auto_sync

true

After a message sync, start the attachment sync in the background (once that index exists)

saved_searches_path

beside index_path

The named searches of saved_search, saved_searches.json (gitignored)

vectors_path

beside index_path

The embeddings for search by meaning, vectors.sqlite (gitignored)

ollama_url

http://localhost:11434

Where Ollama listens

embedding_model

bge-m3

The model asked to embed text; changing it needs mail_vectors.py --build

ollama_timeout

5

Seconds allowed to embed a query before search_all answers with keywords alone

search_mode

keyword

search_all's mode when the caller gives none: keyword, or auto (hybrid once the vectors exist and cover the index)

vectors_auto_sync

true

After a message sync, start mail_vectors.py --sync in the background (once the vectors exist)

The launchd agent is the one place a path cannot come from configuration: launchd needs absolute paths in the plist itself. Replace /ABSOLUTE/PATH/TO in launchd/com.mcp-mail-macos.sync.plist before installing it.


The 29 tools

Search across everything

Tool

Purpose

search_all(query, account, mailbox, unread_only, flagged_only, since, until, limit, sort, snippets, mode)

Search every account, through the local index; mode adds search by meaning

aggregate(group_by, query, account, mailbox, unread_only, flagged_only, since, until, limit, order)

Count matching messages per sender, domain, month, year, account, mailbox or recipient ("who writes to me most", volumes per month)

saved_search(action, name, description, query, account, mailbox, unread_only, flagged_only, since, until, limit, sort, snippets)

Save a frequent search under a name (save), re-run it (run, with optional overrides), or list / show / delete saved ones

get_thread(message_id, limit)

The whole conversation a message belongs to

find_similar(message_id, query, account, mailbox, unread_only, flagged_only, since, until, limit, exclude_thread, snippets)

The messages closest in meaning to a given one ("more like this"), from the stored vectors

index_status()

What the index holds, how old it is, how many messages have a searchable body (per account and overall), and the state of the attachment and vector indexes

sync_index()

Bring the index up to date

search_all covers the whole archive in milliseconds. Subject, sender, To and Cc recipients, body and attachment names are all indexed. FTS5 syntax works — subject: invoice, sender: jane (recipients: searches To and Cc together), AND / OR / NOT, "exact phrase", NEAR(one two, 5). A query that is not valid FTS5 (invoice 12/2025) is reinterpreted word by word, which the answer reports in interpreted_as.

saved_search keeps a frequent search under a name. save stores the search_all parameters you pass (the query, with any Gmail operators, is required) plus an optional description; run executes it and returns exactly what search_all returns, and any parameter passed to run overrides the stored one for that run (saved_search("run", "unread from Jane", limit=5)). Saving under an existing name (any case) replaces that search. Store relative operators such as newer_than:7d rather than fixed dates: the query is kept as text and evaluated again at each run. One tool with an action, rather than one tool per action, keeps the tool list short. The file is written atomically and holds queries only; aggregate does not take a saved search.

aggregate takes the same query, operators and filters but returns counts instead of messages: aggregate("sender", "to:me -is:bulk", since="2026-01-01") lists who writes most, aggregate("month", "from:@example.com") gives the volume per month. Nothing is excluded by default; -is:bulk leaves out newsletters. A message in several mailboxes counts once (except group_by="mailbox"); text inside attachments is not searched.

Operators

Gmail-style operators are read first and become SQL filters (on messages, locations and recipients); only the remaining text goes to FTS5. Write them with no space after the colon, quote values that hold spaces (from:"jane doe"), and prefix - to negate (-is:bulk). Keys are case-insensitive, operators combine with AND only (also two of the same kind): OR next to an operator, or an operator sharing parentheses with OR or free text, is refused with a hint (use -key: to exclude, run two searches for alternatives). NOT from:x is -from:x, and parentheses wrapping operators only, (from:a has:attachment), are ignored. Operators also combine with the account, mailbox, unread_only, flagged_only, since and until parameters. A query made only of operators works and returns the newest first. An unknown word: stays free text; a bad value is refused with a hint. The answer's filters shows how the query was understood (original_query holds what was typed).

Operator

Meaning

from:

sender: full address, @example.com (domain and subdomains) or part of the name or address (from:example.com, from:jane); from:me = your accounts

to: cc:

same, on the recipients table; to:me / cc:me = your accounts

has:attachment

a real attachment (same meaning as has_attachment)

filename:

attachment name, as typed or stemmed (filename:plan.pdf, filename:budget)

larger: smaller:

size, 500K, 2M, 1G or bytes (K = 1024)

older_than: newer_than:

relative to now: 30d, 2w, 6m (30 days), 1y (365 days)

after: before:

YYYY-MM-DD or YYYY/MM/DD, local midnight; after inclusive, before exclusive

is:unread is:read is:flagged is:starred is:bulk

read state, flag (starred = flagged), newsletter / mailing list

in:

mailbox, case-insensitive: the whole name or its last path segment as Mail stores it, so a localised display name such as "Boîte de réception" is not mapped to INBOX (in:inbox, in:archive for [Work]/Archive)

A message filed in several mailboxes (Gmail's All Mail and Important, typically) matches is:unread, is:read, is:flagged and in: when any of its copies does; a negation is the exact opposite (-is:unread = unread in none). to: and cc: are also names of FTS columns: the operator wins. To search those columns as words, write a space (to: jane) or braces ({to}: jane, {to cc}: jane). me is read from Mail's account list on first use (one AppleScript call, cached for the life of the server); if Mail cannot answer, only that query fails, with a hint to use the address itself.

Results are ranked by relevance by default (sort="relevance"): weighted bm25 with a subject match counting most, then sender, attachment names, and To/Cc/body, times a moderate recency bonus (up to +30 % for a mail received today, +15 % at one year old; a strong old match still beats a weak recent one). Each result carries has_attachment (a real attachment: inline logos and png/gif/bmp/svg names are ignored) and is_bulk (a List-Id or List-Unsubscribe header: mailing lists and newsletters), and a snippet (snippets=false to skip): about 200 characters of the body around the first matched word, matched like the index does (case and accents ignored, term* prefixes and quoted phrases handled), or the start of the body when the word is only in the subject, sender or an attachment name. The index cannot store text, so it is read from the message's .emlx file; the file is located from the id alone (Mail shards its folders by the id's digits), with no extra index column. It is null when the file is missing or not downloaded. sort="date" returns the newest matches first instead. The weights and the bonus are constants at the top of mail_search.py. The command-line python3 mail_index.py --search QUERY ranks the same way, --sort date to order by date.

When the attachment index exists, the text of PDFs, Word, Excel, PowerPoint and scans is searched too (see Attachment text). A message found only through an attachment carries attachment_match (filename, and a snippet of the attachment when snippets are on), and its snippet starts with [attachment: name]. Without attachments.sqlite nothing changes.

get_thread uses the conversation grouping Mail computes itself, carried in the index. The whole exchange comes back, including replies filed in another mailbox or sent from another account.

Read directly

Tool

Purpose

list_mailboxes(include_totals)

Accounts and mailboxes, with unread counts

list_messages(mailbox, account, limit, unread_only, include_preview, scan_limit)

Messages of one mailbox, newest first

get_message(message_id, max_body_chars)

Full message: body, headers, attachments

count_unread(mailbox, account)

Unread counts, per mailbox or across accounts

These ask Mail directly, so they see the real state including what has just arrived.

There is deliberately no tool that searches through Mail. Mail serves Apple events on the thread that draws its interface, so any search wide enough to be useful freezes the app for minutes — and the timeout does not rescue it: killing osascript leaves Mail chewing on the event it already accepted, so the freeze outlives the call. Search goes through the index, which reads the same store from disk and needs the same Full Disk Access. If the index is missing, search_all says so and sync_index builds it; there is no faster path worth having.

Prepare a message for review

Tool

Purpose

write_draft(to, subject, body, cc, bcc, attachments, sender, folder)

Write a draft as an .eml file, outside Mail

list_drafts(folder)

Drafts waiting to be sent

read_draft_file(path)

Full content of one draft

edit_draft_file(path, …)

Change the draft in place — see Edit a draft

send_draft_file(path, confirm, keep_file)

Send the draft, then file it away

discard_draft_file(path)

Delete a draft that will not be sent

purge_drafts(folder)

Sweep forgotten drafts

See Drafts are files for why they live outside Mail.

Send

Tool

Purpose

send_email(to, subject, body, cc, bcc, attachments, sender, confirm)

Compose and send

create_draft(to, subject, body, cc, bcc, attachments, sender, signature)

Save a draft in Mail

edit_draft(message_id, …)

Change a draft Mail already holds — see Edit a draft

send_draft(message_id, confirm)

Send a draft Mail already holds

reply_to_message(message_id, body, reply_all, attachments, send, confirm, add_to, add_cc, bcc)

Reply, staying in the thread, with attachments if any; add_to, add_cc and bcc add recipients to the computed ones without duplicates

Confirmation is mandatory. Every tool that actually sends — send_email, send_draft_file, send_draft and reply_to_message(send=True) — does nothing unless confirm=true. Called without it they return confirmation_required along with a preview block describing precisely what would go out: sender, recipients, subject, body, and the attachments actually carried. It doubles as a dry run.

The guard covers every send path rather than one of them: protecting only the draft path would push a caller to recompose with send_email, which is the behaviour worth avoiding in the first place.

to, cc and bcc accept one address, a comma-separated string, or a list. attachments takes absolute paths to existing files, checked before anything is built. Without sender, the account Mail lists first is used — worth being explicit when several accounts coexist, since the sender decides which signature is appended.

Every message goes out as HTML, laid out the way Mail lays out its own: the message, then the account's signature, then a blank line, then the attachments. A plain text body is converted — blank lines become paragraphs, single newlines become breaks, and anything resembling markup is escaped — so a caller never has to write HTML to get a well-formed message. Pass signature=false to leave the signature off.

The signature is not configured here. It is read from Mail's own settings for the sending account: which signature is selected, its HTML, and the images it carries. Editing it in Mail is enough; there is no second copy to keep in step.

Edit a draft

edit_draft (a draft in Mail) and edit_draft_file (an .eml draft) take the same changes, any of them in one call:

Argument

Change

subject

New subject line

to, cc, bcc

The whole new list for that field; an empty string clears Cc or Bcc

add_to, add_cc, add_bcc

Addresses added to that field

remove

Addresses taken off, whichever field holds them

body

New text; the signature and a quoted original below it stay

replacements

Targeted edits, [{"old": "...", "new": "..."}], each found exactly once

add_attachments

Absolute paths of files to attach

remove_attachments

File names of attachments to take off

The draft is edited where it stands, never composed again: only the headers asked about and the text are rewritten, and the formatting, the signature and its logo, the quoted original, the attachments and the thread headers stay byte for byte. An address placed in a field leaves the other two, so adding to To someone who was in Cc moves them. Every change is checked before anything is written; a refused one leaves the draft untouched.

A server cannot change a message in place, so edit_draft files the edited version first and only then expunges the previous one: a failure in between leaves two drafts, never none. The draft therefore gets a new message_id, which list_messages(mailbox="drafts") gives once Mail has checked the account. Close the draft if it is open in a Mail window, or Mail may save its own copy over the edit.

Organise

Tool

Purpose

create_mailbox(name, parent, account)

Create a mailbox, optionally nested

move_message(message_id, target_mailbox, target_account)

Move to another mailbox

delete_message(message_id)

Move to trash

mark_as_read(message_id) / mark_as_unread(message_id)

Read status

flag_message(message_id, flag_color)

red, orange, yellow, green, blue, purple, gray, or none


Drafts are files, not Mail drafts

write_draft writes a self-contained .eml file — attachments embedded — into mails/. macOS renders an .eml in Mail on double-click, so it reads like a real message. send_draft_file submits that file as it stands, so what leaves is byte for byte what was reviewed, then moves it to mails/sent/.

This is not a stylistic choice. Mail cannot send a draft it holds. Its send command only understands an outgoing message, not a message sitting in a mailbox; opening a draft turns it into one, but only after an unpredictable delay that exceeded a minute in testing; moving it to the Outbox does nothing at all. Anything drafted inside Mail therefore has to be re-posted and the original deleted — and on a Gmail account that delete is undone by the server unless it is issued once the send has settled.

Keeping drafts out of Mail removes the problem rather than working around it. Sending becomes a single instant operation with nothing to clean up afterwards.

send_draft remains for drafts that already sit in Mail, whether written by hand or filed there by create_draft. It fetches the message from the server, submits it unchanged and expunges the draft — so the formatting, the signature and every attachment survive, and the deleted draft does not come back.

Retention. An unsent draft is removed after 7 days, an archived one after 30. The sweep runs on every write, every listing, and from sync_index, so a forgotten draft does not sit on disk indefinitely. Files live in mcp-mail-macos/mails/ unless folder says otherwise, and are gitignored: they hold real message content.


The search index

Why

Mail answers message by message. On a mailbox of around 20,000 messages, reading metadata costs about 0.65 s per message and reading a body about 1.7 s. Mail's own search (whose subject contains …) takes about 21 s over 2,500 messages, and searching bodies exceeds 120 s — to the point of leaving Mail unresponsive to every subsequent call for minutes.

Searching an archive of tens of thousands of messages that way would take hours. The index sidesteps it by reading Mail's storage directly.

What feeds it

Two sources, neither sufficient alone:

  • MailData/Envelope Index, Mail's internal SQLite database, for metadata, mailbox membership, and read and flag status. It is copied — together with its write-ahead log — then opened read-only.

  • The .emlx files, for body text and the RFC Message-ID header.

Mail's index holds no full text; the files do not say which mailboxes a message belongs to.

Indexed, not stored

Bodies go into an FTS5 table declared content='': searchable, never kept. The database stores only what is needed to display a result and act on it — subject, sender, date, Message-ID, locations. Reading a message goes back through get_message. For roughly 50,000 messages the index weighs about 80 MB.

messages    (id, account, rfc_id, subject, sender, date_received, size, conversation_id,
             body_indexed, has_attachment, list_id, is_bulk)
locations   (message, account, mailbox, read, flagged)
recipients  (message, kind, address, domain, name)   -- kind is 'to' or 'cc', address lower-cased
messages_fts(subject, sender, "to", cc, attachments, body)   -- FTS5, content=''

Splitting message from locations absorbs Gmail's duplication: a message exists once on disk, in All Mail, and labels are only views. A mailbox of some 50,000 distinct messages yields around 135,000 locations — which is exactly the figure AppleScript reports when its mailboxes are summed.

The durable key is the RFC Message-ID, not Mail's internal id, which changes whenever a message moves.

Build and update

python3 mail_index.py --check    # verify assumptions, build nothing
python3 mail_index.py --build    # full backfill
python3 mail_index.py --sync     # incremental
python3 mail_index.py --search "invoice acme"

--check validates seven points, including that message ids map to files and — most importantly — that membership rebuilt from both sources matches Mail's own per-mailbox counts. --build refuses to start if any of them fails: Envelope Index is undocumented and changes between macOS releases, so a clean refusal beats a silently wrong index.

--sync diffs the set of messages Mail lists against the set the index holds. Deletion is not a special case, and a move reads as a change of location at constant Message-ID. A pass with nothing to do costs about two seconds.

Body coverage

Some messages end up with no searchable body: the file is missing, only a partial download exists, or the message is empty. messages.body_indexed is 1 when a non-trivial body (at least a few word characters) was extracted and 0 otherwise, and index_status reports with_body and without_body for each account and overall. Those messages still match on subject, sender, recipients and attachment names.

For multipart/alternative messages the text/plain part is used, unless it is empty, near-empty or a "this message contains HTML" placeholder: then the HTML part is stripped of markup and indexed instead.

Quoted history is cut from the indexed body, so a reply no longer competes with the messages it quotes. Only safe patterns count: lines starting with > (with the On ... wrote: / Le ... a écrit : line that introduces them, and the non-quoted lines after the block are kept, for bottom-posted and interleaved replies); Outlook From:/De : header blocks (at least three header lines) and -----Original Message----- markers, which cut what follows; the last -- signature line when the block after it is 15 lines or fewer; in HTML, blockquote type=cite, Gmail quote containers and Outlook's reply header. If fewer than 20 word characters would remain, the full text is kept, and a forward keeps its content when the text above it is tiny. Bodies indexed before this need a rebuild (python3 mail_index.py --build).

French stemming

Search understands French plurals, feminines and common verb forms: facture finds factures, relancé finds relance and relancer, travail finds travaux. Accents and case are ignored as before. Under the hood every text column (subject, attachment names, body) is indexed twice, as written and as a light stem (mail_stem.py, pure Python, no dependency), and a query word is searched in both. A message holding the exact word ranks above one holding only another form. Things to know:

  • Quotes mean exact. "facture" and "les factures" match the words as written, in that order, with no stemming.

  • fact* matches word starts in both forms. Sender, To and Cc are never stemmed, so names and addresses stay exact.

  • Numbers, codes and words under four letters are left alone. The stemmer is deliberately light: it will not join a noun and its verb (paiement, payer).

  • The stems change the index content: an index from before needs python3 mail_index.py --build (schema version 5, about 25 % larger).

Schema version

The index carries a schema_version in its meta table (an index that predates versioning counts as version 1). When the code expects a newer one, search_all, sync_index and mail_index.py --sync refuse with an index_outdated error and the hint to run python3 mail_index.py --build; they never mix formats. The version also moves when the indexed content changes, not only the tables: version 5 (stemmed columns next to the raw ones) needs a rebuild from version 4; version 4 (quoted history cut from bodies) needed one from version 3; version 3 (To and Cc kept apart, has_attachment, list_id, is_bulk) needed one from version 2. --build always starts from scratch, writing to <index>.building and swapping the finished file in atomically, so the live index keeps answering until the new one is ready and a crashed build loses nothing. --sync, --build and the automatic sync share one lock file (<index>.sync.lock, holding the owner's PID and refreshed while it works), so they never run at the same time.

Freshness

search_all checks the index's age and runs the sync itself past max_age_minutes (10 by default). The answer carries index_age_minutes and synced, so the caller knows what was searched. If disk access was revoked in the meantime the search still succeeds against the existing index and says so in sync_note rather than failing — a stale result beats an error. A lock prevents two concurrent syncs.

Background sync

launchd/com.mcp-mail-macos.sync.plist runs --sync every ten minutes, independently of any client:

cp launchd/com.mcp-mail-macos.sync.plist ~/Library/LaunchAgents/
launchctl load ~/Library/LaunchAgents/com.mcp-mail-macos.sync.plist

A warning before installing it: the agent reads ~/Library/Mail, so Full Disk Access has to be granted to the program it runs, /usr/bin/python3. That hands the grant to every Python script on the machine — wider than an app-scoped one. A venv interpreter is no better: its path carries a version number and the grant breaks on the first upgrade.

The agent is only worth it if the index must stay current with no client running. Otherwise search_all's own freshness check is enough, and the grant stays scoped to a single application.

Attachment text

Attachment names are always indexed. Their text lives in a second database, attachments.sqlite, built by mail_attachments.py. It is separate on purpose: it is large (text of tens of thousands of files), it is rebuilt on its own schedule, and search works exactly as before when it is absent or broken.

python3 mail_attachments.py --sync      # first run reads everything; later ones only what is new
python3 mail_attachments.py --build     # start from scratch
python3 mail_attachments.py --status    # rows by status, size, last run
python3 mail_attachments.py --measure   # time the extractors on a random sample

Where the files are: Mail keeps the attachments of a message it has not fully downloaded (.partial.emlx) as plain files in <mailbox>.mbox/<store>/Data/<shard>/Attachments/<id>/<part>/<file>, and those of a full .emlx inside its MIME. Both are read; a MIME part is written to a private temporary directory (mkdtemp, mode 0700, outside the repository) for the length of one batch and removed after it, on exit and on SIGTERM. Nothing is written under ~/Library/Mail, and Full Disk Access is needed as for the index.

What is read, and how:

Kind

Method

Rule

PDF

PDFKit text layer (tools/pdftext.swift)

up to 50 pages, attachments_max_mb

Scanned PDF (under 20 characters of text)

Vision OCR, fr-FR + en-US, accurate (tools/ocr.swift)

first 3 pages

docx, xlsx, xlsm, pptx

zip + XML (numbers are left out of sheets)

half of attachments_max_mb

doc

textutil

same

png, jpg, jpeg, heic, tiff

Vision OCR

at least 50 KB, name not a logo, banner, signature, social icon or imageNNN; off with attachments_ocr_images

zip, rar, dwg, audio, video, gif, xls, anything else

skipped

recorded with the reason

The Swift tools are compiled on first use into tools/build/ (gitignored) with /usr/bin/swiftc, and again whenever their source is newer; a missing compiler is reported with xcode-select --install. Each takes many files per call, so the process start is paid once. A tool that stays silent for 40 s is killed; the file it choked on is marked error and the rest of the batch goes again.

Resumable: discovery first records every attachment as a row (pending, or skipped / too_big with a reason), then the pending rows are read in batches of 40 with a commit each. Kill the run at any point and run --sync again: it picks up the remaining rows. A row is read again when its size or modification time changes (--retry-errors retries failures); rows of messages that left Mail, and of files that disappeared, are deleted. Changing a setting (OCR on or off) reclassifies the skipped rows on the next run. A run takes attachments.sqlite.sync.lock, independent of the index lock: the two syncs can run at once, since the attachment one reads Mail's files, not the index.

Column

Meaning

status

pending, ok, empty (no text found), skipped, too_big, error

reason

why skipped or failed: type, small_image, decoration, ocr_off, encrypted, tool_failed...

method

text, ocr, office, textutil

text

the extracted text, one line, capped at attachments_char_limit

attachments_fts(filename, text, filename_stem, text_stem) is contentless like the message table, with the same French stemming; the text itself stays in attachments.text for snippets.

How search uses it: search_all runs the free text against the message index as before, then against attachments_fts (bm25 with the file name counting three times the text), and merges by message. A message found by both gets its own recency-adjusted score plus ATTACHMENT_WEIGHT (0.5) times the attachment's; a message found only through an attachment is ranked by that half score. Attachment hits go through the same filters as any result (operators, dates, account, mailbox, filename:), so an attachment never bypasses them. A query restricted to a message field (subject:, sender:...) leaves attachments out; sort="date" orders the merged set by date. has:attachment is unchanged and filename: still matches the names in the message index.

Keeping it current: a search never waits for attachments (reading them takes minutes). Instead, once attachments.sqlite exists, every successful sync_index (and so every stale-index refresh from search_all) starts mail_attachments.py --sync detached in the background; if one is already running the new one exits at once. Set attachments_auto_sync to false to turn that off and run it yourself, or from launchd next to the message sync (copy launchd/com.mcp-mail-macos.sync.plist, replace its --sync program by mail_attachments.py --sync, and the same Full Disk Access caveat applies). index_status reports attachments: rows by status, size and last run.

Measured on a mailbox of about 54,000 messages and 34,000 attachment files on disk (plus about 13,500 named parts inside full messages), on Apple silicon: PDF text about 30 ms a file, scanned PDF OCR about 350 ms, Word and Excel a few ms, images about 60 ms. A full first run takes on the order of half an hour.

Search by meaning

Keyword search finds the words you typed. search_all(..., mode="semantic") finds the messages that are about what you typed ("unpaid bills" finds a "payment reminder"), and mode="hybrid" fuses both rankings. It is optional: without the vectors file, Ollama or sqlite-vec, everything above behaves exactly as before.

mode

What runs

keyword

The search described above (bm25, attachments, stemming).

semantic

The query is embedded and compared with the message chunks; no keyword involved. Fails with a hint when it cannot run.

hybrid

Both, then Reciprocal Rank Fusion (k = 60): a message scores the sum of 1 / (60 + rank) over the rankings that hold it. Falls back to keyword, with a semantic_note in the answer, when it cannot run.

omitted

The search_mode setting: keyword (the default), or auto = hybrid when the vectors are fresh (built with this recipe, covering at least 95 % of the index, sqlite-vec loadable), keyword otherwise.

Each result of semantic and hybrid carries match (keyword, semantic or both) and, when meaning found it, similarity (cosine of its best chunk). A result found only by meaning has no matched word to anchor a snippet on, so its snippet is the passage that matched. Everything that narrows a keyword search narrows this one too: operators, dates, account, mailbox. The filters are applied before the nearest-neighbour cut, so a narrow filter never loses its hits to the global top. sort="date" reorders the fused best matches newest first. With free text absent (operators alone) there is nothing to embed and the search is a keyword one.

Building the vectors, once, then keeping them current:

brew install ollama && brew services start ollama && ollama pull bge-m3
.venv/bin/pip install -r requirements-semantic.txt
python3 mail_vectors.py --build    # from scratch; --sync resumes and only does what is new
python3 mail_vectors.py --status

What is embedded: the subject and the message's own text (quoted history already cut, links and tokens over 40 characters dropped), in chunks of about 1,000 characters with 150 overlapping, at most 8 per message, so a long thread is represented by its first 8,000 characters or so. The subject is added to the first chunk. A message with no readable body is embedded by its subject alone. The file holds the vector, the chunk number and a hash of the text, never the text: an excerpt is cut again from the message when needed. --sync skips a message already embedded without reading it (--verify reads them again and re-embeds those whose text changed), deletes the vectors of messages the index dropped, commits every batch of 32 chunks (an interrupted run, SIGTERM included, resumes where it stopped) and takes vectors.sqlite.sync.lock, independent of the index's. --sample PERCENT --seed N embeds a reproducible random subset, for measurements. Once the file exists, every successful sync_index starts --sync in the background (vectors_auto_sync); a search never waits for it.

Measured on a mailbox of about 50,700 messages (Apple silicon, bge-m3 in Ollama):

Chunks

70,946: 1.4 per message; 91 % of messages take one or two, 1 % reach the cap of 8, 141 empty ones none

Full build

36 minutes: about 33 chunks a second whatever the batch size from 4 up (16 to 64 measured), 32 per request; the .emlx reading (about 2 minutes for the whole mailbox) is negligible next to it

File size

96 MB (int8, 1 KB a chunk); float32 would be about 4 times larger

Quantisation

int8, one scale per vector. Against float32, 99.3 % of the top 10 neighbours are kept; a binary quantisation (128 bytes a chunk) keeps 66 % and was rejected

Query latency

embedding a query 11 ms once the model is loaded (0.8 s when Ollama has to load it again, after 30 minutes idle); keyword 24 ms, semantic 110 ms, hybrid 124 ms, of which about 75 ms is the comparison with the 71,000 chunks (sqlite-vec, exact, no approximate index)

The vectors are plain int8 blobs in an ordinary table compared with sqlite-vec's scalar functions, not a vec0 virtual table: vec0 does the same exact scan (95 ms against 89 ms for 125,000 chunks), but its contents cannot be read without the extension and it cannot be restricted by message id, both of which the Python fallback and the filters need. The embedding call has a 5 second timeout (ollama_timeout); after a failure Ollama is not asked again for a minute, so a search never pays the timeout twice, and hybrid answers with keywords and says why in semantic_note.

How well it works, on the pairs of Evaluating search quality (550 pairs, mail_eval.py --run --mode ...; MRR / recall@10). These pairs are two or three words of a message, so they measure exact words, which keyword search is built for; semantic search is not expected to win there, and attachment text is not embedded.

Mode

Pairs as drawn

Query words inflected (--inflect)

keyword

0.467 / 0.747

0.330 / 0.547

semantic

0.102 / 0.204

0.092 / 0.173

hybrid

0.440 / 0.744

0.358 / 0.593

Hybrid uses the semantic ranking at a quarter of the weight of the keyword one (RRF_SEMANTIC_WEIGHT). With equal weights, the loose neighbours of a short query pushed exact matches down (MRR 0.410 on the pairs as drawn). Weights from 1 to 0.15, a minimum similarity and a cap on the semantic list were tried on the same pairs; 0.25 without threshold kept recall@10 level with keyword and gained on inflected queries (subject pairs: recall@10 0.465 to 0.580, 40 fewer misses out of 200). The price is a lower first place on exact words (MRR minus 0.026), which is why keyword stays the default: it is right for the most frequent kind of query, needs no Ollama and answers in 24 ms. Search by meaning is where words differ, for instance a query in English finds French mail about "unpaid bills" (a keyword search finds nothing); use mode="hybrid" or "semantic" for those, or set search_mode to auto if you prefer it everywhere.

Attachments are not embedded in this version. Their text is the largest volume of the mailbox (a PDF alone can hold a hundred thousand characters), it is often noise for a language model (letterheads, tables, OCR of stamps), and keyword search already reaches it: in hybrid the keyword side still merges attachment hits, so a message found through an attachment word stays in the fused list.

Similar messages

find_similar(message_id) returns the messages whose meaning is closest to a given one. It needs the vectors file but neither Ollama nor the network: the query is the message's own stored chunk vectors, averaged (each brought to unit length first). The message itself is never returned, and with exclude_thread=True (the default) neither is the rest of its conversation, so the answer is other mail about the subject; pass false to keep the replies. message_id is a message_id reference or the numeric mail_id of search_all. query narrows the candidates like search_all's: Gmail operators, dates, account and mailbox filter them before the nearest-neighbour cut, and plain words must appear in the message. Results have the shape of search_all's, with score (cosine of the best chunk) and the matching passage as snippet. A message without vectors (newer than the last mail_vectors.py --sync, or no text) is an error saying so. On the 71,000 chunks above a search takes about 80 ms (100 ms with snippets), almost all of it the exact comparison; a filter that leaves fewer chunks is faster.

Averaging the chunks was chosen against keeping the best score over each chunk as a query, on 12 random multi-chunk messages (top 10 each, thread excluded): same sender in 61 and 69 of 120, a subject word in common in 88 and 83 of 120, against 0 and 12 for random messages, at 68 ms against 168 ms. The two are as on-topic; the average costs one scan instead of up to eight.

Evaluating search quality

mail_eval.py measures where search puts the message you were looking for, so a change to ranking or matching can be shown to help rather than just to differ. It opens the index read-only and never syncs during a run.

python3 mail_eval.py --generate 200 --seed 1        # pairs from random messages
python3 mail_eval.py --attach 100 --seed 3          # add pairs from attachment text (keeps the other pairs)
python3 mail_eval.py --run                          # rank of each expected message
python3 mail_eval.py --run --save-baseline before   # keep the numbers
python3 mail_eval.py --run --compare before         # deltas, and pairs that got worse
python3 mail_eval.py --run --mode hybrid            # keyword, semantic or hybrid (needs the vectors)

A pair is a query and the message it should find. --generate derives the query from two to four distinctive words of a random message's subject; the body is not stored in the index, so it is not used. Pairs written by hand ("source": "manual") are kept when the automatic ones are regenerated; eval.example.json shows the format with made-up data. The report gives MRR, recall@1, recall@10 and the number of pairs not found in the top 50. Add --json for machine output.

--attach N adds pairs of a third kind, attachment: two or three words that are rare in one attachment's text (and not in its message's own fields), expected result the message. Alone it redraws only the attachment pairs and keeps every other one, so the subject and body numbers stay comparable across runs; the report gives one line per kind.

Pairs and results are written under eval/, which is gitignored: they contain real subjects.


Message identifiers

Every message carries an opaque message_id encoding the account, the mailbox path and Mail's internal id. It is stable between calls and survives a Mail restart — it is not a position in a list.

Two caveats. The internal id is only unique within a mailbox, hence the account and path travelling with it. And a move creates a new one: move_message returns the new message_id when it can find the moved copy through its Message-ID header, and says so when it cannot.

A reference from search_all or get_thread has the same shape and works directly in get_message, reply_to_message or move_message. It points at the smallest mailbox holding the message, because Mail resolves an id by walking the mailbox it is given: aiming at a folder of a few thousand messages rather than one holding tens of thousands changes the response time by an order of magnitude.


Response format

Every function returns a dictionary. On failure:

{
  "ok": false,
  "error_code": "mailbox_not_found",
  "error": "mailbox not found: Drafts (account Work)",
  "hint": "Call list_mailboxes to see the exact mailbox paths."
}

Errors are data, not protocol exceptions, so a caller can correct itself from the code and the hint.

Code

Cause

permission_denied

macOS refuses control of Mail, or access to its storage

apple_event_timeout, timeout

Mail did not answer in time

mail_not_running

Mail is closed and could not be started

account_not_found, mailbox_not_found, message_not_found

Target not found

invalid_message_id

Malformed identifier

attachment_not_found

Attachment missing from disk

confirmation_required

A send was requested without confirm=true; nothing left, preview describes it

not_a_draft

send_draft aimed at a message that is not in a Drafts mailbox

attachments_unreachable, attachments_incomplete

Attachments could not be recovered; nothing was sent

draft_file_not_found, draft_file_unreadable, folder_not_found

.eml draft or its folder missing

text_not_found, ambiguous_replacement

An edit's old text is absent, or appears more than once

not_a_recipient, no_recipient, invalid_address

An edit removes someone absent, would leave no recipient, or names something that is not an address

attachment_not_in_draft, nothing_to_change

An edit removes a file the draft does not carry, or changes nothing

index_missing, not_indexed, sync_failed, sync_timeout

Index absent, incomplete, or not refreshable


Known limitations

All of these come from Mail, not from this server. The figures were measured on an M4 Pro MacBook Pro against real accounts.

Speed

Operation

~2,500-message mailbox

~20,000-message mailbox

Metadata for 20 messages

~1 s

~13 s

Metadata for 200 messages

~13 s

—

One message body

~1.7 s

~1.7 s

A mailbox's unread count

instant

instant

This dictates the defaults: include_preview and include_totals are off, a single list_messages call reads at most twenty previews however many messages it returns and says so in the answer, and search goes through the index rather than through Mail.

What Mail cannot do at all

  • Send a draft it holds. send only understands an outgoing message (-1708). Opening the draft produces one only after an unpredictable delay, sometimes over a minute. Moving it to the Outbox does nothing. The only faithful alternative reported by the community is GUI scripting (Cmd+Shift+D through System Events), which needs Accessibility permission and breaks with any interface change — deliberately not taken here. So a send does not go through Mail at all: the message is submitted over SMTP, through the account's own outgoing server.

  • Compose a message without rewriting it. Setting html content on an outgoing message makes Mail wrap the body in its share wrapper — a stray <br> above the first line, inside a <blockquote type="cite"> — and an attachment made through make new attachment lands inside that body, ahead of the signature. Neither is reachable from AppleScript, and both survive every ordering of the calls, a full HTML document, and setting the property after save. So messages are built as MIME here and handed to the server.

  • Delete a mailbox. delete mailbox fails with -10000 whatever the syntax. A mailbox created by create_mailbox has to be removed by hand.

  • Export an attachment. Mail refuses to write the file anywhere (-10004). This no longer matters for sending: send_draft fetches the whole message from the server and submits it unchanged, so the attachments never have to be read back out of Mail's storage.

  • Set headers on an outgoing message. In-Reply-To and References cannot be set on a message Mail composes, which used to force reply_to_message through Mail's own reply command — opening a compose window and rewriting the body. Building the message here sets them directly instead.

  • Create a mailbox with an account property. It has to happen inside a tell block targeting the account, or -10000.

  • Delete a draft for good. A Gmail account pushes back a draft deleted through Mail a few seconds after the delete reports success. Expunged on the server, it is gone — which is how send_draft removes it.

Behaviours worth knowing

  • Mail never composes here any more, so the autosaves it used to leave behind after a send, and the outgoing messages that accumulated invisibly in its internal list, no longer happen at all. The sweep that hunted them down is gone with them.

  • Mail counts a signature image among the attachments, so a preview built from Mail's own list announces an image the recipient never receives as a file — and lists nothing at all before Mail has downloaded the parts. What a preview describes is therefore the message on the server, with each part settled by its disposition: an attachment is offered, an inline part belongs to the body. For a message written elsewhere that says neither, a part the HTML shows with <img src="cid:…"> is taken as part of the body. What was left behind is reported under kept_inline.

  • Gmail labels are mailboxes, and one message appears in several. INBOX can resolve to All Mail: a message's mailbox field reports where Mail sees it, which is not always what was queried.

  • every mailbox of account returns leaf names, but lookup by slash-separated path works. The server rebuilds full paths by walking the container property.

  • An attachment's name is sometimes inconsistent between calls; its size is reliable.

  • A disabled account disappears from Mail's list without an error.

  • AppleScript calls are wrapped in an explicit with timeout; without it any call over 60 s fails, which a large mailbox reaches easily.

  • Numbers and dates coerced to text follow the machine's locale — a date becomes 1,785863539E+9. The server assembles ISO 8601 dates digit by digit to avoid it.

Message content reaches the client unfiltered

Everything these tools return — bodies, subjects, sender names, attachment names — is whatever arrived in the mailbox, passed through untouched. A message can therefore contain text that reads like an instruction, and an agent consuming this server will see it alongside its own. Treat mail content as data, never as direction, and be wary of a tool call whose arguments were lifted verbatim from a message. This is not specific to this server, but it is worth stating: reading mail on an agent's behalf is exactly the situation prompt injection targets.

read_draft_file takes a path and parses whatever is there as an email, so any readable file on the machine can be turned into a body and handed back. That is deliberate — folder would be pointless otherwise, and attachments already require arbitrary paths — but it means the server is as trusted as the client driving it. It is meant to run locally, for one user.

Scope

The server exposes every account Mail knows about, for reading and writing alike. There is no account allowlist. Adding one means filtering in two places — MessageReference.decode and resolveMailbox on the AppleScript side, plus a WHERE account IN (...) on the index — because the AppleScript tools reach Mail directly and would otherwise still see everything.

The index reflects Mail's local store. What an account has not synced does not exist for Mail, and therefore not for search either.


Testing

Unit tests cover everything that does not need Mail: identifier encoding, address parsing, error classification, AppleScript assembly, .eml round-trips, retention, message layout and reply threading. They run anywhere, in under a second:

python3 -m unittest discover -s tests -t .

Manual checks exercise the live path against a real Mail install:

.venv/bin/python test_manual.py read                              # read-only
.venv/bin/python test_manual.py read --account Work --mailbox INBOX
.venv/bin/python test_manual.py write --to you@example.com        # draft + mailbox
.venv/bin/python test_manual.py write --to you@example.com --send # really sends

read changes nothing: eight checks, two of which verify that errors surface cleanly. write creates a draft and a test mailbox in Mail; the mailbox has to be deleted by hand, since Mail cannot do it through AppleScript.


Project layout

mcp-mail-macos/
├── server.py           # MCP entry point, the 29 tool definitions
├── mail_tools.py       # driving Mail through AppleScript
├── mail_message.py     # building the message: body, signature, attachments
├── mail_signature.py   # the signature Mail would have used, from its settings
├── mail_draft.py       # drafting and sending, on top of the two above
├── mail_imap.py        # the account's own server, for filing and sending
├── mail_files.py       # .eml drafts, retention, leftover sweep
├── mail_edit.py        # editing an existing draft in place, in Mail or as .eml
├── mail_search.py      # querying the index
├── mail_saved.py       # named searches (saved_search tool)
├── mail_index.py       # building and updating the index
├── mail_stem.py        # French light stemmer and query rewrite
├── mail_attachments.py # attachment text: extractors, attachments.sqlite, sync
├── mail_vectors.py     # embeddings for search by meaning: vectors.sqlite, Ollama client, sync
├── mail_eval.py        # search relevance evaluation (pairs, MRR, recall)
├── test_manual.py      # manual checks against a real Mail install
├── tools/              # Swift sources (PDFKit text, Vision OCR), compiled into tools/build/
├── tests/              # unit tests, no Mail required
├── applescript/        # one script per operation, plus shared handlers
│   ├── _common.applescript
│   └── …
├── launchd/            # optional periodic sync agent
├── requirements.txt
├── requirements-semantic.txt   # optional: sqlite-vec
└── README.md

AppleScript files are assembled at run time: _common.applescript is prepended to each script, and a with timeout wrapper is added around the run handler. Parameters travel through argv rather than string interpolation, which rules out injection, and -- protects values starting with a dash. Results are serialised with ASCII separators 31 and 30, which never appear in real mail and are stripped from values before joining — hence no escaping when parsing.


License

MIT. See LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    MCP server that gives Claude and other MCP hosts full access to Mail.app on macOS — search, read, send, reply, flag, move, and more across all accounts configured in Mail.app.
    24
    53 npm
    1
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Local MCP server for multi-account IMAP/SMTP email (iCloud + Gmail via app-specific passwords). Never marks mail read. Cross-folder search, idempotent sends, TLS verified.
    8
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Local MCP server for macOS Mail reads plus visible unsent compose, reply, and forward drafts, and constrained single-message moves.
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    Local MCP server for macOS native apps: Mail, Calendar, Reminders, Notes, Messages, and Contacts. Enables reading and organizing your Mac life through a single stdio process using AppleScript/JXA.
    40
    8 npm
    MIT