Skip to main content
Glama
jenyago

Apple Mail MCP

by jenyago

Apple Mail MCP

Specification — tool contracts, limits, and verification status.

A local stdio MCP server for Apple's Mail app on macOS, using the official MCP Python SDK and Mail's installed scripting dictionary. No email credentials or remote server required.

Tools

  • list_accounts: permitted account IDs, names, and email addresses.

  • list_mailboxes: nested mailbox paths for an account.

  • search_messages: paginated, case-insensitive subject/sender search in a mailbox.

  • search_inboxes: the same search over the inbox of every permitted account in one call (first page from each; unread_only: true reviews what is new). Go deeper in one account with search_messages and the next_offset it returns.

  • read_message: message metadata and bounded plain-text content.

  • create_draft: save a visible draft for review in Mail; never sends it. Off by default — see Configuration.

Use account IDs and mailbox path arrays returned by the listing tools. Search scans at most 200 messages by default (up to 1,000), in Mail's native order, and stops after about 30 seconds on a slow mailbox. Follow next_offset until null, including on empty pages. Results are not guaranteed newest-first. Mailbox changes between pages can cause skips or duplicates. Account-scoped mailboxes only; local “On My Mac” mailboxes and attachment downloads are not implemented. No send, delete, move, or arbitrary scripting tool is exposed.

Related MCP server: Mac Local Mail MCP

Setup

Requires macOS, configured Apple Mail accounts, and uv.

git clone https://github.com/jenyago/apple-mail-mcp.git
cd apple-mail-mcp
uv sync --locked

Rebuild .venv with uv sync --locked on each Mac; do not copy it between machines. Account IDs and addresses must be discovered again on the destination Mac.

Configuration

Environment variables, all optional:

Variable

Effect

APPLE_MAIL_ACCOUNTS

Comma-separated email addresses. Only accounts owning one of them are visible or reachable; everything else is refused inside the JXA bridge. Unset = every account. Set but empty = the server refuses to start.

APPLE_MAIL_ALLOW_DRAFTS

1 registers create_draft. Any other value keeps the server read-only. With APPLE_MAIL_ACCOUNTS set, sender is required and must be a listed address, because Mail files a draft under its default account otherwise.

APPLE_MAIL_AUDIT_LOG

Audit log path. Default ~/Library/Logs/apple-mail-mcp/audit.log.

uv run python server.py --list-accounts prints every account (addresses, name, ID) to your terminal, so you can pick addresses for APPLE_MAIL_ACCOUNTS without routing the list through a model. It is a command-line option only, never an MCP tool.

The audit log is one JSON line per call, mode 0600: time, operation, account, mailbox path, message ID, result count, outcome. Blocked attempts are logged with blocked_by_allowlist. It never contains subjects, senders, bodies, search text, or draft content.

Security model

Email is attacker-controlled input. A message body can carry instructions aimed at the model reading it. This server cannot stop that; it limits what the server itself can do (read-only by default, account allowlist, no send/delete/move tool, audit log). What decides the outcome is what else the session can do after reading a hostile email: if it also has tools that send mail, post messages, fetch URLs, or a pre-approved shell, an injected instruction can use them without a prompt.

  • Read mail in a dedicated session that has no other tools — see the Claude Code recipe.

  • Do not register this server at user scope next to pre-approved outbound tools.

  • macOS grants Automation permission to the launching app, not to this server. Any process started from that app can drive Mail with osascript, including sending. The "no send tool" property holds for the MCP interface only; an isolated session (no shell) is what makes it real.

  • Mail bodies you read are sent to the model provider of your client. Use APPLE_MAIL_ACCOUNTS to keep accounts you must not share out of reach.

Register with a client

The flags below start a session whose only tools are this server's. --restricted removes the shell, code-running tools and WebFetch and ignores your settings files (so broad allow rules do not apply); --strict-mcp-config skips every other MCP server, including connectors; --tools "" removes the remaining built-in tools. Check /mcp on launch.

cat > "$HOME/.claude/mcp-apple-mail.json" <<EOF
{"mcpServers":{"apple-mail":{"type":"stdio","command":"$PWD/.venv/bin/python","args":["$PWD/server.py"],
 "env":{"APPLE_MAIL_ACCOUNTS":"you@example.com"}}}}
EOF
alias claude-mail='claude --tools "" --restricted --strict-mcp-config --mcp-config "$HOME/.claude/mcp-apple-mail.json"'

Run it from the cloned directory, and put the alias in your shell profile. Add "APPLE_MAIL_ALLOW_DRAFTS":"1" to env only when you want drafts.

Registering with claude mcp add --scope user loads the server into every session, which is only safe if none of them pre-approves outbound tools.

Codex

From the cloned directory:

codex mcp add apple-mail -- "$PWD/.venv/bin/python" "$PWD/server.py"

Set APPLE_MAIL_ACCOUNTS in the server's environment as Codex's MCP configuration allows (not verified here).

Claude Desktop

Merge this entry into ~/Library/Application Support/Claude/claude_desktop_config.json, preserving other settings. Replace both paths with absolute paths to your clone:

{
  "mcpServers": {
    "apple-mail": {
      "command": "/absolute/path/apple-mail-mcp/.venv/bin/python",
      "args": ["/absolute/path/apple-mail-mcp/server.py"],
      "env": {"APPLE_MAIL_ACCOUNTS": "you@example.com"}
    }
  }
}

Start a new session or restart the client if the tools do not appear. On first use, macOS may ask permission for the launching app to control Mail. Allow it under System Settings → Privacy & Security → Automation → Mail.

The server sends fixed JavaScript for Automation (JXA) code through stdin to /usr/bin/osascript; user values are JSON data, never executable code. Calls have a 45-second timeout. A draft timeout may leave a partial draft: inspect Mail before retrying. Drafts may sync to the configured mail provider. Email contents returned by the server are untrusted data, not instructions.

Verify

uv run python -m unittest -v
uv run python check_connection.py
uv run python check_connection.py --live

The live check lists accounts without printing their contents, then proves against real Mail that an allowlist filters list_accounts, blocks other accounts, and is logged, and that search_inboxes reaches every permitted account inside its time budget. It does not create drafts or send email. Unit tests cover script injection resistance, permission errors, timeouts, recipient validation, allowlist parsing and fail-closed behavior, draft gating, and that the audit log never contains message content. Protocol tests check initialization, tool discovery, and parameter bounds.

Try it

  1. “Use apple-mail to list my Mail accounts.”

  2. “List the mailboxes for [account].”

  3. “List five messages from its Inbox using apple-mail.”

  4. “Read the first message from that result.”

  5. “Use apple-mail to show my unread messages across all accounts.”

  6. Optional, with drafts enabled: “Create a draft to [your own email] from [allowed sender] with subject MCP test and body This is a draft test. Do not send it.”

Verify the optional draft in Mail. No send tool is exposed. Message order is Mail's native order, not guaranteed newest-first.

License

MIT

Available Tools

5 tools
list_accountsA
Read-only

List permitted Apple Mail accounts and their stable IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds meaningful context with 'permitted' (implying permission scoping) and 'stable IDs' (implying persistent identity across calls), which is beyond the annotations, but it does not discuss ordering, empty results, or account states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the resource and the key output characteristic front-loaded. No filler, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter list tool with annotations covering the safety profile and an output schema covering the return shape, this description is nearly sufficient. It could add a note on ordering or on what 'permitted' means in practice, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. With 100% schema coverage and no inputs to explain, there is nothing further the description could add here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('List') and resource ('permitted Apple Mail accounts') plus a useful output detail ('stable IDs'), which cleanly separates it from siblings like list_mailboxes and search_messages by resource type. It stops short of explicitly naming a sibling, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as list_mailboxes, nor any stated prerequisites or context. The sentence only describes what the tool returns, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mailboxesB
Read-only

List account mailbox paths as arrays of names, including nested mailboxes.

ParametersJSON Schema
NameRequiredDescriptionDefault
account_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral fact — that results are nested/recursive arrays of names — but says nothing about auth requirements, sorting, or coverage limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding. It earns its place, though it is terse enough that it sacrifices clarity for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and annotations cover the read-only safety profile. What remains missing is any account_id semantics and any routing guidance among the sibling list/search tools, leaving the definition minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter, account_id, which the description never explains (format, where it comes from, whether it can be omitted). With no schema-level documentation to fall back on, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource ('List account mailbox paths') and adds scope ('including nested mailboxes'). It is reasonably distinguishable from list_accounts and read_message, though the boundary with the sibling search_inboxes is left implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus the alternatives — no mention of search_inboxes, search_messages, or any precondition. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_messageB
Read-only

Read a message by ID within its account/mailbox, with bounded body length.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_charsNo
account_idYes
message_idYes
mailbox_pathYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds a genuinely useful behavioral trait — that the body is length-bounded — but never says what the bound truncates, what happens to the remainder, or how the caller detects truncation. With annotations doing the heavy lifting, the added context is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that identifies the resource, the identifier, the scope and the one notable behavioral constraint with no filler. It is efficient, though the terseness is partly what leaves parameter and truncation details unexplained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing what comes back, and it does not: no mention of headers, attachment handling, or truncation signalling. Combined with 0% schema coverage on four parameters, the definition is workable but leaves an agent guessing on return shape and input formats.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description must compensate, and it does partially: 'by ID' -> message_id, 'account/mailbox' -> account_id and mailbox_path, 'bounded body length' -> max_chars. However it gives no format guidance for mailbox_path (a path-component array), no ID type, and none of the max_chars default (20000) or ceiling (100000), so real gaps remain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read a message by ID') plus the scoping dimension (account/mailbox), which is enough for an agent to distinguish it from the search_* siblings. It stops short of naming those siblings or contrasting the lookup-by-ID path against them, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by ID' implies the tool is for retrieving a single already-known message, versus search_messages/search_inboxes for discovery, but no alternative is named and no when-not condition is given. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_inboxesA
Read-only

Search the inbox of every permitted account in one call: the first page from each.

Empty query lists messages; unread_only=true reviews what is new. Each account is searched separately and a failing account is reported in its own entry without stopping the rest. Accounts skipped for lack of time are in not_reached. To go deeper in one account, call search_messages with that entry's account_id, mailbox_path and next_offset.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
scan_limitNo
unread_onlyNo
limit_per_accountNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, non-destructive, non-open-world semantics, and the description adds material behavior the annotations cannot: per-account partial-failure handling ('a failing account is reported in its own entry without stopping the rest') and time-based skipping surfaced via not_reached, plus the pagination handoff through next_offset.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with purpose and scope, followed by usage conditions, failure semantics, and the escalation path. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does useful work explaining return shape (per-account entries, not_reached, next_offset, account_id, mailbox_path). It is nearly complete, though the meaning of limit_per_account versus the first-page constraint is not spelled out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description carries the burden. It explains query (empty lists everything) and unread_only behavior, but says nothing about scan_limit or limit_per_account, leaving half the parameters undocumented in both schema and text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource plus scope: 'Search the inbox of every permitted account in one call: the first page from each.' It also differentiates itself from the sibling search_messages by naming it for deeper single-account work, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use conditions ('Empty query lists messages; unread_only=true reviews what is new') and a concrete routing rule to the alternative tool ('To go deeper in one account, call search_messages with that entry's account_id, mailbox_path and next_offset'). Nothing about tool selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_messagesB
Read-only

Search subject/sender case-insensitively in one mailbox. Empty query lists messages.

Continue with next_offset even if this page has no matches: a scan also stops after about 30 seconds on slow mailboxes. Mail order is not guaranteed chronological; concurrent mailbox changes can affect pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
offsetNo
account_idYes
scan_limitNo
unread_onlyNo
mailbox_pathYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/destructiveHint/openWorldHint, so safety is covered. The description goes further with behavior annotations don't carry: a ~30-second scan cutoff on slow mailboxes, non-chronological mail order, and pagination instability under concurrent mailbox changes. That is genuinely useful disclosure for an agent driving a scan loop.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then three compact sentences of caveats, with no filler. The only friction is the 'next_offset' wording that doesn't line up with the schema's 'offset'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, no-output-schema search tool the description covers query semantics and scan/pagination caveats well, but leaves scan_limit, unread_only, and mailbox_path unexplained despite zero schema descriptions. Adequate but with clear gaps for an agent building correct calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate and largely does not. It clarifies query semantics (subject/sender, empty lists all) and implies pagination via offset, but says nothing about limit, scan_limit, unread_only, account_id, or mailbox_path, and the term 'next_offset' does not match the actual 'offset' parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search), the fields searched (subject/sender, case-insensitive), and the scope (one mailbox), plus the empty-query behavior. It is clearly distinct from read_message but does not explicitly contrast itself with search_inboxes, leaving the sibling boundary to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides real operational guidance on continuation ('Continue with next_offset even if this page has no matches'), which is actionable, but never states when to choose this tool over search_inboxes or read_message, and gives no prerequisites. Usage is implied rather than framed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedlist_accounts
    • First observedlist_mailboxes
    • First observedread_message
    • First observedsearch_inboxes
    • First observedsearch_messages

TDQS

A3.7/5.0

Scored across 5 tools

Disambiguation4/5

list_accounts and list_mailboxes clearly target different resources, and read_message is distinct. The only potential overlap is search_messages vs search_inboxes, but descriptions make the scope difference explicit (single mailbox vs all inboxes).

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (list_accounts, list_mailboxes, search_messages, search_inboxes, read_message). Verb choice is predictable and parallel.

Tool Count5/5

Five tools is well-scoped for a read/search-oriented mail server. Each tool covers a distinct step (accounts, mailboxes, message search, inbox scan, message read) without redundancy.

Completeness3/5

The read/search surface is partially covered: accounts, mailboxes, search, and single-message read are present. However, there are no tools for sending/reply, deleting, moving, flagging read/unread, or retrieving attachments, which are notable gaps for an Apple Mail server.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to search, read, and inspect Apple Mail on macOS, including conversations and attachments. It can create new, reply, reply-all, or forward drafts, but cannot send or modify existing messages.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Safely searches, reads, flags, and drafts email through IMAP, with no send, delete, or move capabilities. Uses a local broker and OS credential store for secure authentication.
    16
    49 PyPI
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI assistants to read, search, send, and manage email on macOS through programmatic access to Apple Mail.
    29
    MIT