Stalwart Mail MCP
Allows using local Ollama-hosted vision models as an OCR provider for scanned PDF pages, accepting page images and returning transcribed markdown without sending content to a cloud service.
Allows using OpenAI-compatible vision chat APIs as an OCR provider for scanned PDF pages, accepting page images and returning transcribed markdown.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Stalwart Mail MCPsearch my inbox for unread emails from this week"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Stalwart Mail MCP
An MCP server that lets Claude Desktop (or any MCP client that speaks stdio) work with a mailbox on your own Stalwart mail server: search and read mail including attachments and scans, reply in a thread, send, keep drafts, and look up or add contacts.
It talks JMAP with the mailbox's own credentials. Nothing is installed on the server.
Claude Desktop ──stdio──▶ dist/index.cjs (Node, this MCP server)
│ HTTPS · JMAP (RFC 8620 / 8621 / 9610)
▼
https://mail.example.com/jmap (reverse proxy → Stalwart)There is also a remote mode: the same server run next to Stalwart, added to Claude as a custom connector, so the mailbox works in claude.ai, the mobile apps and every desktop chat — people sign in through Stalwart's own OAuth and the server holds no credentials. See docs/remote.md.
How it fits into a small self-hosted setup — reverse proxy, what to expose, shared mailboxes, branded builds for a family or a team — is described in docs/small-infrastructure.md.
Tools
Names carry a prefix, mail_ by default (a branded build can change it).
Tool | What it does |
| accounts (own + shared), folders with counts, allowed senders, address books |
| full text / from / to / subject / folder / date / unread / has attachment, or a whole thread |
| a whole message by id (HTML → text) with a numbered list of attachments |
| an attachment's content: text, PDF page by page, OCR of scans and photographed documents, images; saves the file to disk |
| send a new mail or a reply ( |
| the same, but only saved to Drafts |
| send / delete a draft by id |
| address books of every account plus senders and recipients from the mail history |
| new contact (own or shared address book) |
Sending is immediate and cannot be undone, so the tool descriptions tell the model to send only on the user's explicit instruction and to create a draft otherwise.
Related MCP server: jmap-mcp
Install
Claude Desktop (extension)
Download stalwart-mail.mcpb from the
latest release and open it —
Claude Desktop offers to install it. Or build it yourself:
npm install
./pack.sh # → stalwart-mail.mcpb
open stalwart-mail.mcpb # Claude Desktop → InstallFill in the server address, the mailbox e-mail and password. Optional: the language and a Mistral API key for OCR. The password is kept in the operating system's keychain.
Any MCP client (stdio)
The server is on npm as stalwart-mail-mcp
and in the MCP Registry as
io.github.cybersmurf/stalwart-mail-mcp, so no checkout is needed:
{
"mcpServers": {
"stalwart-mail": {
"command": "npx",
"args": ["-y", "stalwart-mail-mcp"],
"env": {
"STALWART_URL": "https://mail.example.com",
"STALWART_USER": "jane@example.com",
"STALWART_PASSWORD": "…"
}
}
}
}Claude Code: claude mcp add stalwart-mail --env STALWART_URL=https://mail.example.com --env STALWART_USER=jane@example.com --env STALWART_PASSWORD=… -- npx -y stalwart-mail-mcp
Claude on the web, desktop chat and phone
Run the server next to Stalwart (stalwart-mail-mcp --http, or the Docker image) and add its
URL as a custom connector — step by step in docs/remote.md.
Configuration
Variable | Meaning |
| public address of the server, e.g. |
| the mailbox you sign in as (required) |
| mailbox or app password — sent as Basic auth |
| an OAuth access token instead of the password — sent as Bearer |
|
|
| prefix of the tool names, default |
| display name of the server, default |
| where attachments are saved, default |
| IANA zone for dates in the output, default the machine's zone |
| who reads scans — see OCR providers |
| shortcut: with only this set, scans go to Mistral OCR |
|
|
|
|
|
|
|
|
|
|
What it is allowed to do
Every capability is a switch in the extension settings (or an env variable above), all on by default. A capability that is off is not offered as a tool at all, so it holds regardless of what the client's approval prompts remember:
sending off → a mailbox Claude can read and draft in, but never send from;
sending, drafts and contact edits off → a read-only mailbox;
saving attachments off →
get_attachmentreads the file from a temporary copy and is annotated read-only; with saving on it writes to the download folder and is annotated as a writing tool, which clients may treat differently when asking for approval.
Approvals themselves ("allow once / always allow") belong to the client, not to this server. In Claude Desktop they are set per tool in the extension's settings; a client may ask again after an update that changes a tool's definition.
Attachments and OCR
mail_get_attachment downloads the file and returns what the model can read:
text files as text, HTML converted to text;
PDFs as text page by page (
page_from/page_tofor long ones);PDF pages without a text layer (scans) and photos go to the OCR provider you choose — the result is markdown including tables, and those pages are marked
(OCR). A mixed PDF sends only its scanned pages;images are returned as images; anything above ~600 kB or in HEIC is downscaled on macOS (
sips);other types (docx, xlsx, zip…) are only saved, and the path is returned.
Without a provider, or with ocr: false, a scan comes back as an image of page 1 (macOS) with
a note. OCR is the only thing in this server that sends content anywhere besides your mail
server — to the provider you picked, or nowhere at all with a local model.
OCR providers
| What it is | Needs | Reads PDFs |
| Mistral when | — | — |
| Mistral OCR ( | key | directly |
| Claude through the official SDK (default model | key | directly |
| hosted OpenAI-compatible vision chat APIs | key + model | page images |
| local models on | model | page images |
| any other OpenAI-compatible server |
| page images |
| no OCR | — | — |
MAIL_OCR_API_KEY, MAIL_OCR_MODEL and MAIL_OCR_BASE_URL complete the choice. Examples:
MAIL_OCR_PROVIDER=ollama MAIL_OCR_MODEL=llama3.2-vision # fully local
MAIL_OCR_PROVIDER=openrouter MAIL_OCR_API_KEY=… MAIL_OCR_MODEL=<a vision model>
MAIL_OCR_PROVIDER=anthropic MAIL_OCR_API_KEY=…
MAIL_OCR_PROVIDER=custom MAIL_OCR_BASE_URL=http://nas.lan:8000/v1 MAIL_OCR_MODEL=…Providers that take images only get each scanned page as a PNG taken out of the PDF (the scan itself, scaled to 2000 px). A page that is not one big picture cannot be handed to them and is reported as unread; Mistral and Anthropic read any PDF. Pages are sent three at a time.
Things to expect: a local model can need a minute or more per dense page, which may exceed
your client's tool timeout — read long scans in page ranges. General vision models transcribe
well but, like every OCR, can misplace cells in tables with graphics; preview: true adds the
page image so the model can check. With anthropic, a declined request is retried
server-side on a fallback model (fallbacks: "default") on the current Claude models.
node test/live-ocr.mjs runs the provider configured in the environment against a scanned
fixture (or your own file) and prints the result.
Languages
Tool titles, descriptions, output and error messages are localized. English is the source, Czech is written by hand, and German, Spanish, French, Italian, Dutch, Polish, Portuguese and Slovak are machine translations that no native speaker has reviewed yet — corrections are welcome.
To add or fix a language edit src/locales/<code>.ts (copy en.ts, keep the {placeholders}
and line breaks) and register it in src/locales/index.ts. A locale may be partial; missing
keys fall back to English. npm run test:offline checks every locale against the English keys.
Branded builds (presets)
For a family or a team you can ship an extension where the server address is pre-filled and
people only type their e-mail and password. A preset is a folder with its own manifest.json
(and optionally icon.png); see presets/example.
./pack.sh --preset /path/to/preset # → /path/to/preset/<name>.mcpbThe preset's manifest sets MAIL_TOOL_PREFIX, MAIL_BRAND, MAIL_LANG or
MAIL_DOWNLOAD_DIR through env, and gives server_url a default. Keep the preset's name
stable so Claude Desktop treats new builds as updates.
Tests
npm run test:offline # no mailbox needed: fake JMAP + fake OCR, the real server over stdio
MISTRAL_API_KEY=… node test/offline.mjs --live-ocr # also sends the fixtures to the real OCR
STALWART_URL=… STALWART_USER=… STALWART_PASSWORD=… node test/smoke.mjs # real mailbox
STALWART_URL=… STALWART_USER=… STALWART_PASSWORD=… node test/smoke.mjs --send # also sends a mail to yourselfThe offline test covers attachments, OCR, the tool prefix, language selection and locale
consistency. The smoke test creates a draft and deletes it; with --send it leaves one test
message in the mailbox.
Good to know
A wrong password gets the IP banned by Stalwart after a few attempts. The server signs in on the first tool call, not at start, so restarting the client does not cause bans; calling tools repeatedly with a wrong password does.
Built for and used with Stalwart 0.16. The mail part is plain RFC 8620/8621 and may work with other JMAP servers, but that is untested; contacts need JMAP for Contacts (RFC 9610).
Email/querywith"inMailbox": nullis rejected by Stalwart — the filter must be absent or carry an id.The extension bundle is one CommonJS file (~3.8 MB, most of it pdf.js from
unpdf); the.cjsextension matters becausepackage.jsonsays"type": "module".A tool result in Claude Desktop is capped at about 1 MB, hence the 600 kB limit for inline images.
License
MIT
Available Tools
10 toolsmail_add_contactAdd a contactA
Creates a contact in an address book. Default is the signed-in mailbox's own address book; pass the account of a shared mailbox to store it in the shared one (everyone with access sees it, on phones too through CardDAV). Check mail_search_contacts first so you do not create a duplicate.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Note. | |
| given | No | First name. | |
| title | No | Job title. | |
| emails | No | E-mail addresses. | |
| phones | No | Phone numbers (international format, e.g. +1 202 555 0123). | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. | |
| surname | No | Surname. | |
| nickname | No | Nickname. | |
| full_name | No | Full name, when splitting it makes no sense (e.g. a company as a contact). | |
| address_book | No | Address book name (default = the account's default address book). | |
| organization | No | Company / organization. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a write (readOnlyHint=false), not idempotent, non-destructive, and open-world. The description adds real behavioral context beyond that: the default target is the signed-in mailbox's own address book, and choosing a shared mailbox makes the contact visible to everyone with access, including via CardDAV. It stops short of covering permission requirements or what happens if a duplicate is still created.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded with the action and the constraint; no filler, no repetition of the title. The scoping detail about shared-mailbox visibility is placed exactly where the account decision is made.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, zero-required mutation tool with no output schema, the description covers the destination and the dedupe prerequisite well. Minor gaps remain: it does not state the minimum viable contact (all fields optional — can an empty contact be created?), nor what the tool returns or whether duplicates are rejected server-side.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the account default/shared semantics that the schema's 'account' description already covers in more detail, and adds no new field-level information (requiredness, email format, or address_book behavior) beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a contact in an address book') and immediately scopes the default destination, which distinguishes it from the read-only sibling mail_search_contacts. An agent can tell what this tool does and what it does not do without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Check mail_search_contacts first so you do not create a duplicate') and gives the selection rule for the account parameter (empty for the signed-in mailbox, shared mailbox name otherwise). Both the duplicate-avoidance workflow and the destination choice are spelled out rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_create_draftCreate a draftA
Saves a message as a draft in the Drafts folder (does not send). The user can then review it in webmail or on the phone and send it themselves, or mail_send_draft sends it by id. Same parameters as mail_send_email (including replying via in_reply_to_id and attachments).
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | Cc — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| to | No | Recipients — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| bcc | No | Bcc — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| from | No | Sender address; default is the address of the signed-in mailbox. Must be one of the account's allowed senders. | |
| html | No | Body as HTML (optional; when text is missing it is derived from the HTML). | |
| text | No | Message body as plain text (preferred). Write the signature yourself, the server adds none. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. | |
| subject | No | Subject. A reply gets “Re: …” automatically when it is missing. | |
| reply_to | No | Reply-To (where replies should go) — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| reply_all | No | For a reply: also copy the other original recipients. | |
| attachments | No | Full paths of files on this computer to attach (max. 25 MB each). | |
| in_reply_to_id | No | id of the message you are replying to — sets the thread (In-Reply-To/References), fills in the recipient and subject, marks the original as answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true), so the bar is lower. The description adds genuine context the annotations don't: the write is non-communicating (draft is not sent) and the message remains user-reviewable. It stops short of noting that repeated calls create duplicate drafts (idempotentHint=false) or where the resulting draft id comes from for mail_send_draft.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero waste, and the most important fact (does not send) is front-loaded right after the name. The alternative-tool pointer is placed directly after the workflow explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param, zero-required mutation tool with no output schema, the description covers the essential semantics: it is a non-sending write, parameters mirror mail_send_email, and there are two follow-up paths. The one real gap is that it never says what comes back (a draft id) which mail_send_draft would need, but overall it is sufficient to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 12 parameters in detail. The description only adds a cross-reference ('same parameters as mail_send_email') and highlights in_reply_to_id and attachments, which is helpful orientation but not new parameter-level meaning. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ('Saves a message as a draft in the Drafts folder') and immediately scopes it against the send path ('does not send'). An agent can distinguish this from mail_send_email and mail_send_draft without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the workflow (user reviews in webmail/phone and sends it themselves) and the alternative tool by name ('mail_send_draft sends it by id'), giving the condition that selects each path. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_delete_draftDelete a draftADestructiveIdempotent
Permanently deletes a draft by id. Works only on messages flagged as drafts — it never deletes sent or received mail.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Draft id. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true. The description adds useful behavioral context beyond annotations by stating the deletion is permanent and limited to drafts, preventing accidental attempts on non-draft messages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and scope. Every sentence earns its place with no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple destructive delete-by-id tool, the description is complete: it states the action, permanence, and draft-only restriction. Annotations cover safety traits, the schema covers parameters, and no output schema is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both the 'id' and 'account' parameters are fully documented in the schema. The description only says 'by id' and adds no meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Permanently deletes a draft by id.' It also scopes the operation by noting it works only on messages flagged as drafts, which distinguishes it from operations on sent or received mail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear context and a when-not condition: 'Works only on messages flagged as drafts — it never deletes sent or received mail.' However, it does not explicitly name alternative tools or spell out when to delete versus send or edit a draft.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_get_attachmentOpen an attachmentAIdempotent
Downloads an attachment of a received message, saves it to disk and returns its content: text for text files, the text of every page for PDFs, the picture for images. Scans (PDF pages without a text layer) and photos of documents are read through OCR, tables included; pages that came from OCR are marked “(OCR)” — treat numbers from them with some caution and ask for a preview when a figure matters. Other types (docx, xlsx, zip…) only return the path of the saved file. Pick the attachment by name or by its number from the list in mail_get_email; a message with a single attachment needs neither. The message itself is not changed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Message id (from mail_search_emails / mail_get_email). | |
| ocr | No | false = do not send scans and photos to the OCR provider, return the image only. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. | |
| page_to | No | PDF: last page to return (default: the last one). | |
| preview | No | PDF: also attach an image of page 1 (chart, stamp, signature, checking the OCR). | |
| save_dir | No | Folder to save the file to; default /root/Downloads/Mail-Attachments. | |
| page_from | No | PDF: first page to return (default 1). | |
| attachment | No | Attachment name (a part is enough, e.g. “invoice”) or its number starting at 1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give a coarse profile (readOnlyHint=false because it writes to disk, destructiveHint=false, idempotentHint=true). The description adds substantive behavior: OCR is applied to scans/photos, OCR pages are marked '(OCR)' with a caution about their numbers, other types return only a saved path, and the source message is not modified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the return format, then the OCR caveat, then selection guidance — a logical order. Sentences are dense with parentheticals but each carries actionable content; slightly longer than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so thoroughly (text/PDF/image/OCR/other types). Combined with 100% schema coverage and annotations covering the safety profile, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the 'attachment' parameter can be a partial name or a 1-based number, and OCR marking/preview behavior is tied to how results should be interpreted. It does not restate the page/save_dir/account parameters, which the schema already documents well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with clear scope: downloads an attachment of a received message, saves it to disk, and returns its content. It also specifies the per-type return behavior (text/PDF/image), which lets an agent distinguish it from mail_get_email at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tells the agent how to select the attachment (by name or number from mail_get_email, unnecessary for single-attachment messages) and routes to mail_get_email as the source of the list. It also advises asking for a preview when a figure matters. It stops short of stating explicit when-not-to-use conditions (e.g. attachments that cannot be retrieved).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_get_emailRead a messageARead-onlyIdempotent
Returns a whole message by id (from mail_search_emails): headers, body text (HTML is converted to text) and the list of attachments. mail_get_attachment reads an attachment's content (PDF, image, text). To reply use mail_send_email with in_reply_to_id. Changes nothing (does not mark the message as read).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Message id. | |
| html | No | true = return the original HTML instead of converted text. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, so the bar is lower, yet the description still adds genuinely useful behavior: HTML is silently converted to text unless html=true, and the call 'changes nothing (does not mark the message as read)'—a non-obvious behavior in mail clients. It omits any mention of size limits or permission requirements for shared mailboxes, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and return payload, then alternatives, then the side-effect constraint. No filler and nothing repeated from structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description properly compensates by naming the returned fields (headers, text body, attachment list) and the HTML conversion behavior. Coverage is strong, though an agent still lacks any note on size limits or how attachments must be fetched separately in bulk.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (id, html, account) are already documented in the schema; the baseline is 3. The description only adds provenance for the required id (comes from mail_search_emails) and implies the html behavior, without new meaning for the account parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns a whole message by id') and enumerates exactly what comes back: headers, converted body text, and the attachment list. It also distinguishes itself from mail_get_attachment and mail_send_email, so an agent can place it among its siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says where the id comes from ('from mail_search_emails'), routes attachment reading to mail_get_attachment, and directs replies to mail_send_email with in_reply_to_id. Both the when-to-use and the alternative-routing conditions are stated outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_list_mailboxesFolders, accounts and sendersARead-onlyIdempotent
Returns the available accounts (your mailbox plus shared ones), folders with message counts (unread) and the addresses you can send from. Call it first when you do not know which folders or accounts exist. Changes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. 'Changes nothing' largely restates readOnlyHint, and the description adds little beyond annotations apart from noting unread counts are included in the response. With annotations carrying the main burden, 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler; the discovery guidance ('call it first') is front-loaded and the read-only reassurance closes it efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description enumerates what is returned (accounts, folders with unread counts, sendable addresses) and the zero-required-parameter discovery role. Nothing essential for calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional 'account' parameter is fully documented in the schema, including a pointer back to this tool. The description adds no syntax or format detail beyond that, so the baseline 3 for high schema coverage applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Returns) and three concrete resources (accounts, folders with unread counts, send-from addresses). This is clearly distinguishable from siblings like mail_search_emails or mail_get_email, which operate on messages rather than mailbox structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to call it: 'Call it first when you do not know which folders or accounts exist.' This is a clear discovery-use condition. It doesn't name alternative siblings or state when NOT to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_search_contactsFind a contact / addressARead-onlyIdempotent
Finds an e-mail address or phone number: searches the address books (your own and shared ones) and also the mail history (who wrote to you and whom you wrote to, matching the query). Use it whenever you need a person's or company's address before sending anything. Without a query it lists the whole address book.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max. contacts from the address book per account. | |
| query | No | Name, part of an address, company… | |
| account | No | Limit to the address books of one account; default = all. | |
| include_history | No | Also search senders/recipients of mail (only with a query). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive/openWorld, so safety is covered. The description adds genuinely useful behavioral context: results merge address books with mail history, history is query-dependent, and an omitted query dumps the entire address book rather than failing. It stops short of describing result ordering or result caps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the two sources, then the usage trigger and default behavior. Three sentences with no filler, though the parenthetical about shared address books and the closing default-behavior sentence could be compressed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup with no output schema and fully documented parameters, the description covers what the tool searches, when to reach for it, and the no-query fallback. Missing only result shape/ordering details, which are minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented, including that include_history requires a query. The description reinforces the query-optional semantics but adds no syntax or value-format detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Finds an e-mail address or phone number') and names the two data sources it spans: address books (own and shared) plus mail history. This clearly separates it from siblings like mail_search_emails (message content) and mail_add_contact (mutation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use trigger ('whenever you need a person's or company's address before sending anything') and explains the no-query behavior. It does not explicitly contrast with mail_search_emails for the ambiguous case of looking someone up by message content, but the context is otherwise unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_search_emailsSearch messagesARead-onlyIdempotent
Searches mail (full text, sender, recipient, subject, folder, date, unread, attachments) or lists a whole thread. Returns message ids (for mail_get_email), thread, sender, date and a preview. Newest first, paging through limit/offset. Without a filter it returns the latest messages from all folders (trash included) — for “what is new” use mailbox="inbox", unread=true.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Recipient (part of the name or address). | |
| from | No | Sender (part of the name or address). | |
| text | No | Full text across subject, body and headers. | |
| after | No | Only messages received from this date on (YYYY-MM-DD or ISO 8601). | |
| limit | No | How many messages to return (1–100). | |
| before | No | Only messages received up to this date (YYYY-MM-DD or ISO 8601). | |
| offset | No | How many messages to skip (paging). | |
| unread | No | true = unread only. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. | |
| mailbox | No | Folder: inbox, sent, drafts, trash, junk, archive, or a folder name. | |
| subject | No | Part of the subject. | |
| thread_id | No | List every message of this thread (threadId from an earlier result); other filters are then ignored. | |
| has_attachment | No | true = only messages with an attachment. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive/openWorld, so the description's job is to add beyond that — and it does: default scope includes trash, results are newest-first, paging is via limit/offset, and the returned fields (ids, thread, sender, date, preview) are named. It doesn't discuss rate limits or result caps beyond the schema's limit=100.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with no filler: capability, return shape plus ordering/paging, then the default-scope caveat and its remedy. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value information, and it does (ids for mail_get_email, thread, sender, date, preview, newest-first ordering). For a 13-parameter, all-optional search tool, an agent has everything needed to form a correct first call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema carries the per-parameter meaning. The description adds cross-parameter guidance not present in the schema: the no-filter default, the inbox+unread combination for new mail, and thread_id suppressing other filters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (searches mail) and enumerates the searchable facets plus the thread-listing mode, which is a distinct capability. It references sibling mail_get_email as the natural follow-up, so an agent can place it in the workflow without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete conditional: with no filter it returns the latest messages from all folders including trash, and it directs 'what is new' queries to mailbox="inbox", unread=true. It also notes thread_id overrides other filters. It stops short of naming a competing tool to prefer for other cases, but the context provided is operationally clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_send_draftSend a draftA
Sends an existing draft by id (from mail_create_draft or from a search in the drafts folder). Sends immediately — only on the user's explicit instruction.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Draft id. | |
| from | No | Sending address, when it differs from the draft's From header. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds that the send is immediate, which is useful behavioral context beyond the annotations, though it does not discuss irreversibility, auth requirements, or after-send effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The core action and the explicit-instruction constraint are front-loaded and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 100% schema coverage, rich annotations, and no output schema, the description supplies the essential context: what it sends, by id, with source and an immediate-send warning. Minor omissions such as explicit differentiation from mail_send_email and any after-send behavior keep it from being fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both optional parameters (from and account) and the required id are fully documented in the schema. The description adds only that the id comes from mail_create_draft or a drafts-folder search, which is helpful but does not meaningfully extend the other parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Sends) and resource (an existing draft) scoped by id. It distinguishes drafting from the sibling mail_send_email by operating on an existing draft, and it names mail_create_draft as a source of the id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition for use: only on the user's explicit instruction. It also specifies that the draft id comes from mail_create_draft or a drafts-folder search. It does not explicitly compare this tool against mail_send_email or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_send_emailSend an e-mailA
Sends an e-mail from the mailbox — a new one, or a reply (in_reply_to_id). A copy is stored in Sent. It sends IMMEDIATELY and cannot be undone. Send only when the user explicitly said the message should go out; otherwise create a draft (mail_create_draft). Verify the recipient before sending — never guess an unknown address, look it up with mail_search_contacts.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | Cc — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| to | No | Recipients — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| bcc | No | Bcc — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| from | No | Sender address; default is the address of the signed-in mailbox. Must be one of the account's allowed senders. | |
| html | No | Body as HTML (optional; when text is missing it is derived from the HTML). | |
| text | No | Message body as plain text (preferred). Write the signature yourself, the server adds none. | |
| account | No | Which account: leave empty for the signed-in mailbox, or name a shared mailbox / group you have access to (the part before @ is enough; a full address works too). mail_list_mailboxes shows what exists. | |
| subject | No | Subject. A reply gets “Re: …” automatically when it is missing. | |
| reply_to | No | Reply-To (where replies should go) — a list of addresses as “name@domain” or “First Last <name@domain>”. | |
| reply_all | No | For a reply: also copy the other original recipients. | |
| attachments | No | Full paths of files on this computer to attach (max. 25 MB each). | |
| in_reply_to_id | No | id of the message you are replying to — sets the thread (In-Reply-To/References), fills in the recipient and subject, marks the original as answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, openWorld=true, idempotent=false, so safety is partly covered; the description still adds high-value behavior: it sends IMMEDIATELY, cannot be undone, and leaves a copy in Sent. It does not cover failure behavior or how the send is confirmed, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and the irreversibility warning before the routing rules. Every clause carries actionable information; nothing is redundant with the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, zero-required, no-output-schema tool, the description covers the essentials an agent needs: immediacy, irreversibility, Sent archival, and which sibling to use instead. It leaves only minor gaps such as post-send confirmation or partial-failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents recipient formats, account selection, attachment paths, and the in_reply_to_id thread behavior. The description only nods at in_reply_to_id in passing, adding essentially no semantics beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a precise verb+resource ('Sends an e-mail from the mailbox') and immediately disambiguates the reply case via 'in_reply_to_id'. Combined with the named siblings (mail_create_draft, mail_send_draft), an agent can tell this apart from the drafting tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use ('only when the user explicitly said the message should go out'), an explicit when-not with the alternative ('otherwise create a draft (mail_create_draft)'), and a prerequisite ('never guess an unknown address, look it up with mail_search_contacts'). This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v2.3.0- First observed
mail_add_contact - First observed
mail_create_draft - First observed
mail_delete_draft - First observed
mail_get_attachment - First observed
mail_get_email - First observed
mail_list_mailboxes - First observed
mail_search_contacts - First observed
mail_search_emails - First observed
mail_send_draft - First observed
mail_send_email
TDQS
Scored across 10 tools
Each tool maps to a distinct resource+action: listing mailboxes, searching, fetching a message, fetching an attachment, sending, draft lifecycle (create/send/delete), and contact search/add. Boundaries are clear, with draft-sending vs email-sending distinguished by tool name and description.
Every tool follows a strict mail_verb_noun pattern (mail_list_mailboxes, mail_search_emails, mail_get_email, mail_create_draft, etc.). No mixing of conventions or verb styles.
10 tools is well-scoped for a mail server spanning folders, messages, attachments, drafts, and contacts. Each tool clearly earns its place with no redundancy.
Core lifecycle is well covered: read/search/attachments, send, full draft lifecycle (create/send/delete), and contact search/add. Gaps remain around managing received mail (move, archive, mark read/unread) and contact update/delete, but agents can work around most of these.
Maintenance
Related MCP Connectors
- Lettio MCPOAutheu.lettio
Private, EU-hosted email for AI agents over JMAP: read, search, reply, organize, send.
Email infrastructure for AI agents — send, receive, search, and reply to email over MCP.
Your mailbox for MCP clients: search, read, draft, send, rules and notes. Sending is off by default.
Your IMAP mailbox as an MCP server: read, search and (if you allow it) organize mail. Open source.
Related MCP Servers
- AlicenseBqualityBmaintenanceEnables users to manage email accounts via IMAP/SMTP, including reading, searching, sending emails with attachments and calendar invites, all through natural language interactions with MCP-compatible clients.14MIT
- FlicenseNot gradedqualityDmaintenanceMCP server exposing JMAP email and Sieve script operations as tools, enabling mailbox management, email creation, search, flagging, and Sieve script management.2-
- AlicenseNot gradedqualityCmaintenanceConnects AI agents to self-hosted Stalwart mail servers via a Cloudflare Worker and JMAP, enabling mailbox search, reading, listing, and two-step draft-and-send email operations through MCP.MIT
- AlicenseNot gradedqualityBmaintenanceEnables searching, reading full conversations, and sending email across multiple IMAP/SMTP mailboxes from any MCP client, with multi-user support and per-user API tokens.MIT