icloud-mcp
Provides tools for managing iCloud calendars (CalDAV), contacts (CardDAV), and email (IMAP/SMTP).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@icloud-mcplist my upcoming events for today"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
iCloud MCP Server
Model Context Protocol server for iCloud. Gives Claude (Desktop, Code, claude.ai connectors) and any other MCP client full access to an iCloud account:
Calendar (CalDAV): list/search/create/update/delete events, recurring events (RRULE), alerts (VALARM), IANA timezones, all-day events, iTIP invitations and cancellations by email.
Contacts (CardDAV): list/search/get/create/update/delete, including notes; labels (home/work/custom) survive updates.
Mail (IMAP/SMTP): folders, listing, server-side search, full messages with attachments, attachment download, sending (HTML, attachments, threaded replies), drafts, move/delete/read flags.
Runs in two modes with the same code:
Mode | Transport | Typical use | Credentials |
Local ("offline") | stdio | Claude Desktop / Claude Code launch it as a subprocess on your machine |
|
Server | Streamable HTTP, stateless | Docker / Cloud Run / any host, many users | Per-request headers or HTTP Basic auth, env fallback |
Tools
Tool | Description |
| Calendars with IDs; Reminders lists are flagged |
| Events in a date range, recurring series expanded per occurrence, sorted by start |
| Text search over summary/description/location |
| Create an event: timezone, all-day, |
| Partial update; change timezone/recurrence/alerts; re-sends invitations when attendees change |
| Delete and email cancellations to attendees |
| Read contacts (name, phones, emails, addresses, organization, title, notes) |
| Write contacts |
| Folders with IMAP flags |
| Newest messages of a folder with readable |
| Server-side IMAP search: |
| Full message(s): headers, |
| Download an attachment: save to a local directory (stdio) or return it inline |
| Send mail: multiple recipients, CC/BCC, HTML with auto plain-text alternative, local file attachments, threaded reply ( |
| Same inputs as |
| Move between folders; delete to Trash ( |
| Toggle |
Every tool carries MCP annotations (readOnlyHint, destructiveHint) so clients can ask for confirmation before writes. Errors are returned as MCP tool errors with actionable messages (e.g. wrong password vs. missing credentials).
Behaviour worth knowing
IDs: calendar/contact/event IDs are full URLs, pass them back verbatim. Email IDs are IMAP UIDs and are only valid inside the folder they came from.
Times: pass
YYYY-MM-DDTHH:MM:SSplus an IANAtimezone(default:DEFAULT_TIMEZONE, which falls back to the machine's zone, then UTC) orYYYY-MM-DDfor all-day events. Returned times are ISO 8601 with offset.Recurring events:
calendar_list_eventsreturns one entry per occurrence; update/delete apply to the whole series.Email bodies are converted from HTML to readable text and truncated to
EMAIL_BODY_MAX_CHARS(20 000) so newsletters do not flood the model context.full_html=truereturns the raw HTML too.Attachments on disk (
save_dir,attachment_paths) are enabled by default in stdio mode and disabled in HTTP mode. Override withICLOUD_MCP_LOCAL_FILESand restrict to a folder withICLOUD_MCP_LOCAL_FILES_ROOT.Rich text that users paste into event notes/locations or contact notes comes back from iCloud as HTML; it is converted to Markdown so links and lists survive (
ICLOUD_HTML_MODE=markdown, default), rendered to plain text (text) or passed through (raw).Tool groups can be switched off per instance with
ICLOUD_ENABLED_CATEGORIES=calendar,contacts,email(e.g. a read-only mail agent getsemailonly).Only iCloud hosts are ever contacted with the account's credentials: calendar, event and contact URLs on any other host are rejected, and mail headers are validated against injection.
Outbound recipients can be restricted with
EMAIL_SEND_ALLOWLIST=@mycompany.com,partner@example.com; it applies toemail_send,email_save_draftand calendar invitations, so a prompt-injected agent cannot mail data to arbitrary addresses.
Related MCP server: orchard-mcp
Requirements
Python 3.11+ (tested on 3.12 and 3.14) or Docker
An iCloud account and an app-specific password: https://account.apple.com/account/manage → Sign-In and Security → App-Specific Passwords. iCloud Mail additionally needs an active
@icloud.com/@me.com/@mac.comaddress.
Local mode (Claude Desktop / Claude Code)
1. Install
git clone https://github.com/mike-tih/icloud-mcp.git
cd icloud-mcp
uv venv && uv pip install -e . # or: python3 -m venv .venv && .venv/bin/pip install -e .
cp .env.example .env # add ICLOUD_EMAIL / ICLOUD_APP_SPECIFIC_PASSWORDTry it:
.venv/bin/icloud-mcp --help2. Claude Code
claude mcp add icloud -- /absolute/path/to/icloud-mcp/.venv/bin/icloud-mcpor, without a checkout at all:
claude mcp add icloud -e ICLOUD_EMAIL=you@icloud.com -e ICLOUD_APP_SPECIFIC_PASSWORD=xxxx-xxxx-xxxx-xxxx \
-- uvx --from git+https://github.com/mike-tih/icloud-mcp icloud-mcp3. Claude Desktop
Config file: macOS ~/Library/Application Support/Claude/claude_desktop_config.json, Windows %APPDATA%\Claude\claude_desktop_config.json.
{
"mcpServers": {
"icloud": {
"command": "/absolute/path/to/icloud-mcp/.venv/bin/icloud-mcp",
"env": {
"ICLOUD_EMAIL": "you@icloud.com",
"ICLOUD_APP_SPECIFIC_PASSWORD": "xxxx-xxxx-xxxx-xxxx",
"DEFAULT_TIMEZONE": "Europe/Berlin"
}
}
}
}python /absolute/path/to/icloud-mcp/run.py works as the command too. Restart Claude Desktop completely after editing the file; the server shows up under the tools icon.
Server mode (Streamable HTTP)
icloud-mcp --http # 0.0.0.0:8000/mcp, stateless
icloud-mcp --http --port 9000 --path /icloud --statefulOr with Docker:
docker compose up -d # reads .env for optional fallback credentials
curl http://localhost:8000/healthThe image runs icloud-mcp --http with PORT from the environment, so it works unchanged on Cloud Run, Fly.io, Railway etc. The server is stateless: every request carries its own credentials, no sessions are kept, and instances can be scaled horizontally.
Authentication (per request)
Checked in order:
Headers
X-Apple-EmailandX-Apple-App-Specific-PasswordAuthorization: Basic base64(email:app-specific-password)Environment
ICLOUD_EMAIL/ICLOUD_APP_SPECIFIC_PASSWORD(single-user deployments)
Protecting the endpoint
Set MCP_AUTH_TOKEN to a long random secret. Every MCP request must then carry it as Authorization: Bearer <token> or X-MCP-Token: <token> (/health stays open). Without a token the server logs a warning and ignores the environment credentials over HTTP, so an unprotected port can never hand out the operator's account; per-request credentials still work. Set ICLOUD_MCP_ALLOW_ENV_CREDENTIALS=true only if you really want an open endpoint bound to one account (e.g. on a private network). docker-compose.yml publishes the port on 127.0.0.1 only.
MCP_AUTH_TOKEN=$(openssl rand -hex 32) icloud-mcp --http
claude mcp add --transport http icloud https://mcp.example.com/mcp \
-H "Authorization: Bearer <token>" \
-H "X-Apple-Email: you@icloud.com" -H "X-Apple-App-Specific-Password: xxxx-xxxx-xxxx-xxxx"Example with Claude Code against a remote server:
claude mcp add --transport http icloud https://mcp.example.com/mcp \
-H "X-Apple-Email: you@icloud.com" -H "X-Apple-App-Specific-Password: xxxx-xxxx-xxxx-xxxx"Always put the server behind HTTPS: app-specific passwords travel in headers.
Configuration
All settings are environment variables (a .env file next to the checkout is loaded). See .env.example for the full list. The important ones:
Variable | Default | Purpose |
| – | Fallback credentials |
| machine zone or | Timezone for naive event times; set explicitly on servers |
|
| Body truncation |
|
| Outgoing attachment budget |
| – | Allowed outbound addresses/domains (send, drafts, invitations) |
| stdio: on, HTTP: off | Allow reading/writing attachments on the server's disk |
| – | Confine those files to a directory |
| – | Shared secret required on HTTP requests |
| true with token, else false | Serve the env account over HTTP |
|
| Tool groups to expose |
|
| Rich text in event/contact fields: |
| stdio, | HTTP transport |
|
| Logging (always to stderr, stdout is reserved for stdio) |
Troubleshooting
"Authentication required": no credentials reached the server. Check the
envblock / headers."iCloud rejected the credentials" / HTTP 401: wrong app-specific password, or the Apple ID password was used.
"IMAP login failed" but calendar works: the Apple ID has no iCloud Mail address (Apple IDs created with a third-party email cannot use iCloud Mail).
Events land at the wrong time: set
DEFAULT_TIMEZONEor passtimezoneexplicitly.HTTP 401
unauthorizedfrom the server itself:MCP_AUTH_TOKENis set and the request did not carry it."Refusing to use ... URL on untrusted host": an ID was not one returned by the list tools; pass the URL verbatim.
Server logs: everything goes to stderr; in Claude Desktop see Help → Show Logs.
Development
uv pip install -e ".[dev]"
pytest # unit tests + end-to-end through the MCP protocol with mocked IMAP
ruff check src testsProject layout:
src/icloud_mcp/
├── server.py # FastMCP app, tool definitions, CLI entrypoint
├── auth.py # per-request credential resolution
├── config.py # environment configuration and logging
├── calendar.py # CalDAV operations, RRULE/timezone handling, iTIP mail
├── contacts.py # CardDAV operations
├── mail.py # IMAP/SMTP operations
├── mail_utils.py # MIME parsing, HTML→text, attachments, special folders, header validation
├── html_render.py # HTML→Markdown for rich text in event/contact fields
└── urls.py # iCloud host allow-list for calendar/contact URLsLicense
MIT, see LICENSE.
Available Tools
24 toolscalendar_create_eventCalendar Create EventA
Create a calendar event (optionally recurring, with alerts) and email iTIP invitations to attendees.
Returns the created event including its id/url. If some invitations could not be sent,
invitations_failed lists those addresses.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | End, same format as start. For all-day events the end date is inclusive. | |
| rrule | No | RFC 5545 recurrence rule for a repeating event, e.g. 'FREQ=WEEKLY;BYDAY=TU', 'FREQ=WEEKLY;INTERVAL=2;BYDAY=MO', 'FREQ=MONTHLY;BYMONTHDAY=15;COUNT=6'. Do not include DTSTART (the event start is the series start). | |
| start | Yes | Start: 'YYYY-MM-DDTHH:MM:SS' (interpreted in `timezone`) or 'YYYY-MM-DD' for an all-day event. | |
| summary | Yes | Event title. | |
| location | No | Event location. | |
| timezone | No | IANA timezone for naive start/end values, e.g. 'Europe/Berlin' or 'UTC'. Default: Etc/UTC. | |
| attendees | No | Attendee email addresses. Each receives an iTIP invitation by email. | |
| reminders | No | Alerts: minutes before the event start (e.g. [60, 10]; 0 = at start, negative = after start) or a time of day on the event's date as 'HH:MM' (e.g. ['09:00'] for an all-day event). | |
| calendar_id | No | Calendar URL/ID from calendar_list_calendars. Omit to use all event calendars (or the first one when creating). | |
| description | No | Event notes. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate a write operation (readOnlyHint=false) and openWorldHint=true, implying external side effects. The description adds valuable context by explicitly mentioning emailing iTIP invitations and the 'invitations_failed' field, which clarifies the external behavior and error handling beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary purpose and then the return behavior. Every word is informative, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core action, return value, and failure handling. It does not explain usage selection, but that is covered under usage_guidelines. Given the complexity (10 parameters, output schema exists), the description is adequate and does not omit critical behavioral information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-specific details beyond what the schema already provides; it mentions the return field but that is not parameter semantics. Thus it meets the baseline without extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Create' with a clear resource 'calendar event' and mentions optional features (recurring, alerts) and side effects (iTIP invitations). It also states the return value and failure handling, distinguishing it from siblings like calendar_update_event and calendar_delete_event.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states what the tool does but does not provide explicit guidance on when to use it versus alternatives like calendar_update_event or calendar_list_events. There are no 'when to use' or 'when not to use' directives, leaving selection to the agent's inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_delete_eventCalendar Delete EventADestructive
Delete an event (the whole series for recurring events) and email cancellations to its attendees.
| Name | Required | Description | Default |
|---|---|---|---|
| event_id | Yes | Event URL/ID as returned by calendar_list_events or calendar_search_events. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and readOnlyHint=false, which indicate mutation. The description adds the key behavioral detail that it deletes the entire series for recurring events and sends email cancellations to attendees, which goes beyond the annotations. This is valuable context for the agent to understand side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence, clear and impactful. It front-loads the core action and includes the critical side effects concisely. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is destructive and has an output schema, the description covers the essential side effects and parameter source. It might benefit from mentioning what the output contains or any error conditions, but the output schema can handle that. Overall, it's sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage, meaning the parameter event_id is fully described in the input schema. The description reinforces that the event_id is from calendar_list_events or calendar_search_events, which adds useful sourcing context. However, it doesn't add much else, so a 4 is fair given the schema already covers it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (delete an event) and the resource (event), and specifies that it deletes the whole series for recurring events. It distinguishes itself from calendar_update_event and other siblings by noting the email cancellation behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, but the action is self-evident. It implies usage for deleting events, and mentions it applies to recurring events, which helps. However, it lacks guidance on when not to use it (e.g., for non-recurring events or for deleting a single occurrence).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_calendarsCalendar List CalendarsARead-onlyIdempotent
List the account's calendars with their IDs (URLs).
Reminders lists are included with read_only=true; they cannot hold events.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, destructiveHint. Description adds relevant behavioral detail: reminders lists are returned with read_only=true and cannot hold events, which is beyond the generic readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded purpose, no redundancy. Every sentence adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only listing tool with an output schema, the description covers the only notable edge case (reminders lists) and provides enough context to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters defined, so baseline 4 applies. No parameter documentation needed since the tool takes zero arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List), resource (account's calendars), and output (IDs/URLs). The note about reminders lists distinguishes it from event-listing siblings like calendar_list_events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage: if you need calendar IDs, use this. But it doesn't explicitly state when to choose it over siblings like calendar_search_events or when not to use it. The context of listing all calendars is clear but no explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_eventsCalendar List EventsARead-onlyIdempotent
List events in a date range, sorted by start.
Each entry has id/url, summary, description, location, start, end (ISO 8601 with offset, or a date for all_day events), start_timezone, recurring, rrule, attendees and calendar. Recurring series are expanded into one entry per occurrence; all occurrences share the series id, so update/delete apply to the whole series.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | Range end, inclusive for dates (YYYY-MM-DD or ISO datetime). Default: 365 days ahead. | |
| start_date | No | Range start (YYYY-MM-DD or ISO datetime). Default: 90 days ago. | |
| calendar_id | No | Calendar URL/ID from calendar_list_calendars. Omit to use all event calendars (or the first one when creating). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint. The description adds critical context: recurring series are expanded into per-occurrence entries sharing a series id, with update/delete applying to the whole series. Also clarifies date formats for all-day events. This goes well beyond annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise paragraphs: the first front-loads the main action, the second details entry structure and recurring behavior. Every sentence earns its place, with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, parameters are fully documented, and the description covers edge cases (all-day event dates, recurring expansion, series ID semantics). An agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents each parameter with defaults and formats. The description mentions 'date range' but adds no additional parameter-specific meaning. Baseline of 3 is appropriate when schema carries the detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb and resource: 'List events in a date range, sorted by start.' It states exactly what the tool does and distinguishes it from search, creation, and other calendar operations. The scope (date range) is explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context that this lists events in a range, distinguishing it from creation/update/delete. However, it does not explicitly contrast with calendar_search_events or state when to prefer one over the other. No explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_search_eventsCalendar Search EventsARead-onlyIdempotent
Search events by text within an optional date range (same output as calendar_list_events).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Case-insensitive text matched against summary, description and location. | |
| end_date | No | Date YYYY-MM-DD or ISO datetime. | |
| start_date | No | Date YYYY-MM-DD or ISO datetime. | |
| calendar_id | No | Calendar URL/ID from calendar_list_calendars. Omit to use all event calendars (or the first one when creating). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, so the description correctly avoids repeating those. It adds useful behavioral context by stating that results are equivalent to calendar_list_events and that filtering can be bounded by an optional date range. No statement contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one focused sentence: action, target, optional scope, and an output-reference parenthetical. Every phrase earns its place and the main behavior is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The rich input schema, output schema, and read-only/idempotence annotations cover the details a caller needs, and the description ties them together with the search and date-range semantics. The only gap is the absence of an explicit usage rule versus calendar_list_events, so it falls just short of fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with explicit descriptions for query (case-insensitive match on summary, description, and location), start/end dates (YYYY-MM-DD or ISO datetime), and calendar_id. The description contributes no additional parameter semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the operation ('Search'), the resource ('events'), and the distinguishing scope ('by text' within an optional date range). Referencing the same output as calendar_list_events further disambiguates it from the plain list sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys that this tool is the text-search entry point for events, which is a clear use context. However, it never explicitly says when to choose it over calendar_list_events or when an unfiltered list is the better alternative; the 'same output' note only addresses return format, not selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_update_eventCalendar Update EventADestructive
Update fields of an existing event. Only provided fields change; recurring series are updated as a whole.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | New end, same format as start. | |
| rrule | No | New recurrence rule for the whole series (e.g. 'FREQ=WEEKLY;BYDAY=TU'). Pass '' to make the event non-recurring. | |
| start | No | New start ('YYYY-MM-DDTHH:MM:SS' or 'YYYY-MM-DD'). | |
| summary | No | New title. | |
| event_id | Yes | Event URL/ID as returned by calendar_list_events or calendar_search_events. | |
| location | No | New location (empty string clears). | |
| timezone | No | IANA timezone for the new start/end. Omit to keep the event's current timezone. | |
| attendees | No | New attendee list; replaces the existing one and sends updated invitations. | |
| reminders | No | New alerts as minutes before start or 'HH:MM' (replaces existing alerts; [] removes all). | |
| description | No | New notes (empty string clears). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, so the description doesn't need to restate that this is a mutating operation. The description adds valuable behavioral context: partial updates ('Only provided fields change') and the whole-series behavior for recurring events. It also implicitly warns that updates to recurring series affect the entire series, which is important for an agent to know. It doesn't mention side effects like invitation sending, but the schema covers that for attendees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The first sentence states the action and the partial-update behavior; the second sentence adds the critical recurring-series caveat. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 10 parameters, the description is concise but the schema carries the parameter details. The description covers the two most important behavioral nuances: partial updates and whole-series recurrence. It doesn't mention that attendees replacement sends invitations, but that is in the schema's parameter description. Given the output schema exists and annotations cover the destructive nature, this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters thoroughly. The description adds the key semantic that only provided fields change, which is critical for understanding the null-default pattern. However, it doesn't add much beyond that because the schema already explains each parameter's meaning and special values (e.g., empty string clears). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Update') and resource ('fields of an existing event'), and immediately distinguishes itself from create/delete siblings by clarifying that only provided fields change and that recurring series are updated as a whole. This is clear and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: to modify an existing event, as opposed to creating or deleting. It does not explicitly name alternatives or exclusions, but the context of sibling tools and the phrase 'existing event' provide clear usage context. A slightly stronger statement about when not to use it (e.g., for creating events) would earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_createContacts CreateA
Create a contact in the default address book and return it with its id/url.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Full name, e.g. 'Jane Doe'. | |
| notes | No | Free-text notes. | |
| title | No | Job title. | |
| emails | No | Email addresses. | |
| phones | No | Phone numbers. | |
| addresses | No | Postal addresses, one string each. | |
| organization | No | Company / organization. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey that this is a non-read-only, non-idempotent write operation. The description adds useful context about the default address book and the returned id/url, but does not disclose duplicate behavior, permission requirements, or other side effects. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler. The verb, scope, and return behavior are front-loaded, and every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a fully described schema, an output schema, and annotations covering write/idempotency behavior, the description is largely sufficient. It could briefly mention duplicate handling or prerequisites, but these gaps are minor for this simple create tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter already documented. The description adds no per-parameter detail beyond what the schema provides, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Create'), names the resource ('contact'), specifies the scope ('default address book'), and states the outcome ('return it with its id/url'). This clearly distinguishes the tool from siblings like contacts_update, contacts_get, or contacts_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call it when creating a new contact. However, it does not explicitly contrast it with contacts_update, mention prerequisites, or state when not to use it. The context is clear but relies on inference from the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_deleteContacts DeleteADestructive
Delete a contact permanently.
| Name | Required | Description | Default |
|---|---|---|---|
| contact_id | Yes | Contact URL/ID from contacts_list or contacts_search. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is established. The description adds the word 'permanently', which conveys irreversibility, but it does not disclose side effects, cascading deletions, or any other behavioral context beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundant information. It earns its place by stating the action and permanence without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool, the combination of description, schema, and annotations is nearly sufficient. The output schema exists, the parameter source is documented, and the destructive nature is both annotated and verbally reinforced; a small usage-guidance note would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single contact_id parameter is already described as coming from contacts_list or contacts_search. The tool description itself adds no parameter-level meaning, so the schema carries the burden and justifies the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Delete a contact permanently.' It clearly separates this from sibling tools like contacts_get, contacts_update, and contacts_search, and the word 'permanently' reinforces that this is the destructive delete operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's action obvious, but it does not explicitly state when to use it or when to avoid it. It does not mention alternatives, exclusions, or any prerequisite such as retrieving the contact_id first, though the schema partially covers that.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_getContacts GetARead-onlyIdempotent
Get one contact with all fields.
| Name | Required | Description | Default |
|---|---|---|---|
| contact_id | Yes | Contact URL/ID from contacts_list or contacts_search. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds that the result includes 'all fields,' but it does not disclose behavior like not-found errors or permission requirements; with annotations covering the main concerns, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short, front-loaded sentence with no filler. Every word adds meaning, and it is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter get operation with a full output schema and comprehensive annotations, the description plus schema covers everything needed. The ID source is documented in the schema, and the read-only/idempotent behavior is already annotated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents contact_id and its source. The description adds no extra meaning beyond the schema, matching the baseline for fully covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Get') and resource ('one contact'), and 'all fields' distinguishes it from list/search siblings. It is unambiguous about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The schema notes that contact_id comes from contacts_list or contacts_search, which implies the intended flow: find first, then fetch one. It does not explicitly exclude alternatives, but the singular scope and ID-source hint provide clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_listContacts ListARead-onlyIdempotent
List contacts from the default address book: id/url, name, phones, emails, addresses, organization, title, notes.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of contacts to return. Default: all. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering the safety profile. The description adds value by specifying the returned fields (id/url, name, phones, emails, etc.), which is not in the annotations or schema, and does not contradict any annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the verb and resource, then lists the fields returned. There is no redundancy or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter, comprehensive annotations, and an output schema, the description fully covers what the tool does and what it returns. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully documents the only parameter (limit) with a description and default value, so schema coverage is 100%. The tool description adds no additional meaning to the parameter, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'contacts from the default address book', and enumerates the fields returned. This distinguishes it from siblings like contacts_get (single contact) and contacts_search (filtered search) without needing to read the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing all contacts from the default address book, but it does not explicitly mention alternatives or when not to use it. There is no guidance on when to prefer contacts_search or contacts_get, so the usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_searchContacts SearchARead-onlyIdempotent
Find contacts by name, organization, email or phone number.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of matches. | |
| query | Yes | Text matched against name, organization, emails and phone digits. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, idempotent, non-destructive profile, so the description does not need to restate safety. The description adds context that matching spans name, organization, emails, and phone digits, but it does not disclose behavior such as partial matching, result limits, ordering, or empty-result handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-structured sentence with no filler. The verb, resource, and scope are front-loaded, making the tool's purpose immediately understandable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only search tool with a complete input schema and an output schema, the description provides the essential behavior. The only notable gap is explicit usage differentiation from sibling tools, which is a usage-guideline concern rather than a completeness issue for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both query and limit are already documented. The description's mention of searched fields mirrors the query parameter description and adds no substantive new meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Find') and resource ('contacts'), and lists searchable fields: name, organization, email, phone. It clearly communicates a match-based lookup rather than direct retrieval, though it does not explicitly name sibling tools like contacts_list or contacts_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: use this tool when you need to locate contacts by partial information such as name, organization, email, or phone. However, there is no explicit guidance about when to prefer it over contacts_list or contacts_get, and no exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_updateContacts UpdateADestructive
Update a contact. Only provided fields change; list fields replace the existing values (entries that keep the same number/address keep their labels such as home/work).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New full name. | |
| notes | No | New notes (empty string clears). | |
| title | No | New job title. | |
| emails | No | New email list (replaces existing). | |
| phones | No | New phone list (replaces existing). | |
| addresses | No | New address list (replaces existing). | |
| contact_id | Yes | Contact URL/ID from contacts_list or contacts_search. | |
| organization | No | New organization. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description reveals key behaviors: only provided fields are modified, list fields replace existing values, and matching entries preserve labels. These details go beyond the annotations (destructiveHint=true) and clarify the partial-update semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences immediately deliver the core action and the most important behavioral caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an update tool with a full schema, output schema, and annotations, the description covers the essential behavioral semantics. It doesn't need to enumerate parameters; the schema handles that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are documented. Description adds important semantics that only provided fields change and list replacement behavior, which supplements the schema's per-field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Update') and resource ('a contact'), and clarifies that only provided fields change. Does not explicitly name sibling tools for differentiation, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives such as contacts_create or contacts_delete. It is implied by the name and sibling set, but there is no condition or exclusion stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_deleteEmail DeleteADestructive
Move a message to Trash, or delete it permanently.
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| permanent | No | True: delete permanently. False (default): move to the Trash folder ('Deleted Messages' on iCloud). | |
| message_id | Yes | Message ID (IMAP UID) from email_list_messages/email_search, valid within `folder` only. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description elaborates on the annotations by explaining that the default behavior is a soft delete (move to Trash) while setting permanent to true causes permanent deletion. This adds meaningful nuance beyond the destructiveHint and readOnlyHint annotations, though it could further note irreversibility of permanent deletion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that directly states the core behavior and both operational modes. Every word contributes meaning, with no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema fully documents the parameters, and the description explains the primary behavioral choice. Given the output schema exists and annotations already mark the operation as destructive, the description is sufficient for an agent to invoke the tool correctly, though a brief note on permanent-deletion irreversibility would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are already documented in the input schema. The description itself adds no new parameter-level detail; it only reiterates the behavioral distinction between permanent and non-permanent deletion, which the schema already captures.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Move a message to Trash, or delete it permanently') and a clear resource ('a message'). It also distinguishes between soft and hard delete, which separates it from sibling tools like email_move, making the tool's purpose immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the two modes of operation: moving to Trash by default or permanent deletion when requested. It does not explicitly mention alternative tools or when not to use it, but the context is clear enough for an agent to choose this tool for deletion-related tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_get_attachmentEmail Get AttachmentARead-onlyIdempotent
Download an email attachment.
With save_dir the file is written to disk and its path returned; otherwise the file is returned inline (images as image content, other types as an embedded resource).
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| save_dir | No | Directory on the machine running this server to save the file into. Recommended when running locally (stdio). When omitted the file content is returned inline. | |
| attachment | Yes | Attachment index ('1') or (partial) file name as listed by email_get_message. | |
| message_id | Yes | Message ID (IMAP UID) from email_list_messages/email_search, valid within `folder` only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the mutation safety profile. The description adds concrete behavior: it writes to disk when save_dir is provided (returning a path) and otherwise returns inline content (with type-specific handling). This goes beyond the annotations by explaining the side effect of saving a file, which is a meaningful addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero filler. The primary purpose is front-loaded, followed by a concise explanation of the two operating modes. Every sentence contributes essential information, and the structure is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool without an output schema, the description adequately explains return values (path, inline content with type-specific handling). It covers the both modes of operation and implicitly states the core use case. No critical information required for calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already well-documented. The description adds value by explaining the semantic distinction of save_dir (disk vs inline) and the return format (path vs image vs embedded resource). This supplements the schema's basic field definitions, making the tool's behavior clearer for an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Download an email attachment.' This unambiguously identifies the action and distinguishes it from siblings like email_get_message (which retrieves the full message) and email_get_messages. The two modes (save to disk vs inline) are briefly noted, further clarifying scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides conditional guidance on when to use save_dir versus inline return, which directly informs selection of parameters. However, it does not explicitly contrast with alternative tools or state when not to use this tool. Since the purpose is self-evident and the mode decision is clear, this is helpful but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_get_messageEmail Get MessageARead-onlyIdempotent
Get one message: headers (from, to, cc, reply_to, date, Message-ID), flags, body_text and the attachments list.
Attachments are described by index, name, mime_type and size; download one with email_get_attachment.
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| full_html | No | Also return body_html (raw HTML, size-capped). | |
| message_id | Yes | Message ID (IMAP UID) from email_list_messages/email_search, valid within `folder` only. | |
| include_body | No | Include the body. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds useful behavioral context beyond annotations by specifying the exact return composition and the attachment metadata shape, helping the agent understand what the tool actually produces.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. The primary purpose and return fields are front-loaded, and the attachment note is placed second as a natural follow-up. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich input schema, output schema, and strong annotations, the description is largely complete for a read-only single-message fetch. It covers the main return contents and points to the correct sibling for attachment downloads. It could be slightly stronger by explicitly contrasting with email_get_messages, but nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already fully documents all four parameters and their defaults. The description adds no additional parameter-level semantics, but it does not need to; the baseline of 3 applies because the schema carries the burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Get one message" and enumerates the exact returned fields (headers, flags, body_text, attachments list). It clearly distinguishes itself from the sibling email_get_messages by emphasizing "one message," and from email_get_attachment by describing attachments as metadata only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies single-message retrieval and explicitly directs attachment downloads to email_get_attachment. However, it does not state when to prefer this tool over email_get_messages, email_list_messages, or email_search, nor does it provide exclusion criteria. Usage context is mostly implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_get_messagesEmail Get MessagesARead-onlyIdempotent
Fetch several messages in one round trip (same fields as email_get_message). Missing IDs yield an error entry.
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| full_html | No | Also return body_html. | |
| message_ids | Yes | Message IDs (IMAP UIDs) from the same folder. | |
| include_body | No | Include bodies. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds valuable behavioral context by noting that missing IDs yield an error entry, which is not captured by annotations. It also implies a performance benefit (one round trip) without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no fluff. The primary purpose is front-loaded, and the error behavior is stated in a second sentence. Every word earns its place, and the description is easily skimmable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the parameter schema covers all inputs, the description is adequate. It addresses the key differentiator (batch) and a potential failure mode (missing IDs). No critical information is missing for an agent to call this tool correctly, though it could mention that IDs must belong to the same folder, but that is already in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with each parameter described in detail, so the description adds little beyond referencing the return fields. It does not clarify parameter semantics beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches multiple messages in a single round trip, referencing the singular email_get_message to indicate identical return fields. It distinguishes itself from the sibling email_get_message by the explicit batch nature, leaving no ambiguity about what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'several messages in one round trip' implies the tool is for batch retrieval, and the reference to email_get_message suggests that tool is for single messages. However, it does not explicitly state 'use email_get_message for a single message' or provide exclusions, though the openWorldHint annotation adds context that any input is acceptable. The guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_list_foldersEmail List FoldersARead-onlyIdempotent
List mail folders with their IMAP flags (e.g. \Sent, \Trash).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds the return detail of IMAP flags, enriching behavioral context beyond what annotations provide. It does not contradict any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that leads with the action and resource, then provides a concrete example of the output. Every word earns its place, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless list tool with an output schema, the description fully conveys what will be returned (folders with IMAP flags). There are no operational details missing, and the annotations cover side effects and idempotency. The description is sufficient for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is trivially 100%. Per guidelines, a baseline of 4 applies since there is nothing to explain about parameters. The description does not need to compensate for any coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'mail folders', and adds specific detail about IMAP flags with examples (\Sent, \Trash). This makes it distinct from sibling tools like email_list_messages, which operate on messages rather than folders.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, nor does it mention exclusions. However, the purpose is unambiguous and the scope is clearly folders, so usage is implied. It could benefit from a note about when to prefer this over email_list_messages, but the domain makes it obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_list_messagesEmail List MessagesARead-onlyIdempotent
List the newest messages in a folder: id, subject, from, to, date, flags, unread, has_attachments and body_text.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of messages, newest first. | |
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| unread_only | No | Only unread messages. | |
| include_body | No | Include body_text (readable text, truncated). Set false for a fast header-only listing. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is read-only, non-destructive, idempotent, and open-world. The description adds 'newest' ordering and lists returned message fields, but it does not mention body_text truncation or that body content is included by default, though the schema documents those details. Overall this is safe and non-contradictory, with only moderate added behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence: it states the action and scope first, then gives a compact enumeration of returned fields. There is no filler, repetition, or unnecessary context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich input schema, output schema, and read-only annotations, the description orients an agent well for a simple folder-listing call. It could be more complete by suggesting when to use email_search or email_get_message instead, but nothing essential for invoking this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully documents all four parameters with descriptions, defaults, and constraints (100% coverage), so the description does not need to compensate. The description's field list is about output, not parameters, and adds no meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List') on a concrete resource ('newest messages in a folder') and names the main returned fields. It is clear and unambiguous, but it does not explicitly differentiate from sibling tools like email_search or email_get_message(s), so it stops short of the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrasing implies this tool is for browsing recent folder contents, and the folder parameter references email_list_folders for valid folder names. However, there is no explicit guidance on when to choose this over email_search or email_get_message(s), nor any exclusions for search-style lookups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_mark_readEmail Mark ReadB
Mark a message as read (set \Seen).
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| message_id | Yes | Message ID (IMAP UID) from email_list_messages/email_search, valid within `folder` only. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnly=false and destructive=false, so the description need not repeat that. 'set \Seen' adds the useful detail that this is a flag mutation rather than a move or delete. It does not disclose whether marking an already-read message has any side effects or what a successful call returns, but the output schema presumably covers the latter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is six words and contains no filler; the action and mechanism are front-loaded. It earns a high score for efficiency, even though more usage detail could have been included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter mutation with full schema coverage and existing output schema, the description gives enough to invoke the tool correctly. The main gap is the lack of explicit routing to email_mark_unread or any caution about when not to mark a message read, so it is not fully contextualized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All parameters (folder and message_id) are fully documented in the JSON schema, including defaults and scope, so the description doesn't need to repeat them. It adds no extra parameter semantics, which is acceptable but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action and target: 'Mark a message as read' with the IMAP flag mechanism. This is specific enough to distinguish from email_mark_unread by the seen/unseen state, though it does not explicitly name that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance about when to choose this over email_mark_unread, no prerequisites, and no mention of folder/UID validity (left to schema). The only context is the action itself; decision-making is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_mark_unreadEmail Mark UnreadA
Mark a message as unread (clear \Seen).
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| message_id | Yes | Message ID (IMAP UID) from email_list_messages/email_search, valid within `folder` only. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, indicating a mutation. The description adds the technical detail 'clear \Seen', which explains the underlying mechanism. However, it does not disclose potential side effects, error conditions, or the requirement for the message to exist in the specified folder. It adds some value but not rich behavioral context beyond what annotations and schema already provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and result. No wasted words; it is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, the presence of an output schema (as indicated by context), and annotations covering mutation behavior, the description is sufficient for an agent to call it correctly. It could mention the need to specify a non-default folder, but the schema already covers that. Overall, minimal but complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both folder and message_id having descriptive text in the schema. The description does not add further meaning to the parameters; it only mentions 'message' generically. Since the schema already documents parameters well, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Mark a message as unread (clear \Seen)' states a specific verb (mark), a clear resource (a message), and the precise state change (unread). It clearly distinguishes from sibling email_mark_read, which does the opposite. No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives like email_mark_read or email_move. It implies the action but does not provide guidance on choosing it over the sibling. The context is clear enough for an agent to infer, but it lacks explicit routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_moveEmail MoveADestructive
Move a message to another folder. The message gets a new UID in the destination folder.
| Name | Required | Description | Default |
|---|---|---|---|
| to_folder | Yes | Destination folder (must exist; see email_list_folders). | |
| message_id | Yes | Message ID (IMAP UID) from email_list_messages/email_search, valid within `folder` only. | |
| from_folder | Yes | Current folder of the message. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag the operation as destructive and non-idempotent. The description adds a key behavioral consequence beyond that: the message gets a new UID in the destination folder, which is critical for subsequent references. It does not cover error cases, but the annotations lower the bar and this added context is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words: the first states the action, the second states the important consequence. The most critical information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter move operation with full schema coverage, annotations, and an output schema, the description is nearly complete. It could explicitly state that the original message_id becomes invalid after the move, but the new-UID statement already implies this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all three parameters. The description does not add additional parameter-level meaning beyond what the schema already provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Move') and resource ('a message') with a destination folder, making the operation unambiguous. The additional UID note adds precision and helps distinguish it from other email actions like delete or send.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: when a message needs to be relocated. It also provides a prerequisite ('destination folder must exist; see email_list_folders'), which is actionable guidance. It does not discuss alternatives, but no sibling tool offers a move/copy operation, so this is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_save_draftEmail Save DraftA
Save a message to the Drafts folder without sending it, so the user can review and send it from any mail client.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | CC address(es). | |
| to | No | Recipient address(es), comma-separated. May be omitted for a draft. | |
| bcc | No | BCC address(es). | |
| body | Yes | Message body (plain text, or HTML when html=true). | |
| html | No | Treat body as HTML. | |
| subject | Yes | Subject line. | |
| reply_to_folder | No | Folder of reply_to_message_id. | INBOX |
| attachment_paths | No | Local file paths to attach (see email_send). | |
| reply_to_message_id | No | UID of a message this draft replies to. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds the key behavioral fact that the message is not sent and is saved to Drafts, which is useful. However, it doesn't disclose details like whether an existing draft is overwritten, how the draft is identified in responses, or any side effects beyond saving. With annotations covering the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It front-loads the core action ('Save a message to the Drafts folder without sending it') and adds the user benefit ('so the user can review and send it from any mail client') in a compact clause. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 9 parameters, 100% schema coverage, an output schema, and annotations. The description is brief but sufficient for a draft-saving tool: it clarifies the non-sending behavior and the purpose. It doesn't explain return values, but the output schema exists, so that's not required. It could mention that 'to' may be omitted (schema already says this) or clarify attachment behavior, but the schema covers those. Overall, complete enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters. The description adds no parameter-specific meaning beyond what the schema provides. Baseline 3 is correct when the schema does the heavy lifting. The description's mention of 'review and send from any mail client' adds context but not parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Save a message to the Drafts folder without sending it.' It uses a specific verb ('save') and resource ('message to the Drafts folder'), and distinguishes it from sending by explicitly noting 'without sending it.' This differentiates it from the sibling tool email_send.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when the user wants to save a draft for later review/sending from any mail client. It doesn't explicitly name alternatives or exclusions, but the contrast with 'without sending it' and the sibling email_send provides clear context. It could be improved by explicitly saying 'use email_send to send immediately,' but the current guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_searchEmail SearchARead-onlyIdempotent
Server-side IMAP search. All provided filters are combined with AND; at least one is required.
Returns the same fields as email_list_messages, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Text matched against the message body (slower). | |
| limit | No | Maximum number of messages, newest first. | |
| query | No | Text matched against Subject OR From. | |
| since | No | Only messages on/after this date, YYYY-MM-DD. | |
| before | No | Only messages before this date, YYYY-MM-DD. | |
| folder | No | IMAP folder name. iCloud folders: INBOX, Drafts, 'Sent Messages', 'Deleted Messages', Junk, Archive (see email_list_folders). | INBOX |
| sender | No | Text/address matched against From (full addresses work best). | |
| subject | No | Text matched against Subject. | |
| recipient | No | Text/address matched against To or Cc. | |
| unread_only | No | Only unread messages. | |
| include_body | No | Include body_text (readable text, truncated). Set false for a fast header-only listing. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, open-world, and non-destructive traits. The description adds meaningful behavior: server-side execution, AND combination of filters, requirement of at least one filter, and newest-first ordering. It also references the output format via email_list_messages. This adds value beyond annotations without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with no redundancy. The purpose is front-loaded, followed by the key filter rule, and then the output reference. Every sentence contributes essential information, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the comprehensive schema and the existence of an output schema, the description does not need to repeat parameter or return details. It covers the search-specific behavior and references email_list_messages for output structure. It could mention pagination or default folder behavior, but those are already in the schema. The description is sufficient for correct usage in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema documents all 11 parameters with detailed descriptions (100% coverage). The description adds crucial interaction semantics: 'All provided filters are combined with AND; at least one is required.' This clarifies how multiple filter parameters relate and enforces a constraint not present in the schema. Since schema already covers each parameter individually, the description's added value is the cross-parameter rule, which earns a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Server-side IMAP search' and specifies the key behaviors: filters combined with AND, at least one required, and newest-first ordering. It also references email_list_messages for return fields, which distinguishes it from that sibling. This is a specific, unambiguous purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for filtered searching by stating 'Returns the same fields as email_list_messages' which hints at the similarity to a listing tool. However, it does not explicitly state 'use this when you need to filter' or contrast with alternatives like email_get_message. The guidance is clear enough for an agent to infer when to search, but lacks explicit exclusions or alternative selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
email_sendEmail SendA
Send an email from the account (SMTP) and store a copy in the Sent folder.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | CC address(es), comma-separated. | |
| to | Yes | Recipient address(es), comma-separated. | |
| bcc | No | BCC address(es), comma-separated. | |
| body | Yes | Message body (plain text, or HTML when html=true). | |
| html | No | Treat body as HTML (a plain-text alternative is generated automatically). | |
| subject | Yes | Subject line. | |
| reply_to_folder | No | Folder of reply_to_message_id. | INBOX |
| attachment_paths | No | Local file paths (on the machine running this server) to attach. Only available when local file access is enabled (default for stdio). | |
| reply_to_message_id | No | UID of a message to reply to; sets In-Reply-To/References so mail clients thread the reply. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true. The description adds the important behavioral detail that a copy is stored in the Sent folder, which is beyond the annotations. It also implies side effects (sending) consistent with readOnlyHint=false. No contradiction. It could mention failure modes or rate limits, but the Sent folder detail is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise and front-loaded with the primary action. It includes the key side effect without unnecessary detail. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a rich schema (100% coverage) and an output schema, the description is mostly complete. It could mention that sending is irreversible or that attachments require local file access, but the schema already notes the local file access condition. The Sent folder side effect is disclosed. Overall adequate for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters. The description adds no additional parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate because the schema carries the burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Send an email'), the mechanism ('from the account (SMTP)'), and a key side effect ('store a copy in the Sent folder'). It is specific and distinguishable from siblings like email_save_draft, which would be the closest alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for sending email but does not explicitly state when to use this tool versus email_save_draft or other email tools. It does not mention prerequisites like SMTP configuration or local file access for attachments, though the schema covers attachment_paths availability. No explicit exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
v0.3.0- First observed
calendar_create_event - First observed
calendar_delete_event - First observed
calendar_list_calendars - First observed
calendar_list_events - First observed
calendar_search_events - First observed
calendar_update_event - First observed
contacts_create - First observed
contacts_delete - First observed
contacts_get - First observed
contacts_list - First observed
contacts_search - First observed
contacts_update - First observed
email_delete - First observed
email_get_attachment - First observed
email_get_message - First observed
email_get_messages - First observed
email_list_folders - First observed
email_list_messages - First observed
email_mark_read - First observed
email_mark_unread - First observed
email_move - First observed
email_save_draft - First observed
email_search - First observed
email_send
TDQS
Scored across 24 tools
Each tool maps to a distinct resource-action pair across calendar, contacts, and email, with list/search/get variants clearly separated by purpose. Batch operations and single-item fetches are explicitly distinguished, and no two tools appear interchangeable.
All tool names follow a uniform lowercase snake_case pattern of domain_verb_noun, such as calendar_create_event, contacts_update, and email_list_messages. The domain prefix makes the resource area immediately clear, and there are no mixed conventions or vague verbs.
24 tools is above the typical single-domain range, but the server intentionally covers three distinct iCloud domains—Calendar, Contacts, and Mail—and each tool serves a specific operational need. The count is slightly heavy but well justified by the breadth of functionality exposed.
Calendar events and contacts have full create/read/update/delete coverage, while email covers searching, fetching, sending, drafting, moving, deleting, and read-state management. Recurring events, invitations, and attachments are handled, leaving no obvious dead ends in the core workflows.
Maintenance
Related MCP Connectors
A MCP server that works with Outlook Calendar to manage event listing, reading, and updates.
A MCP server that works with Google Calendar to manage event listing, reading, and updates.
MCP server for Nylas — read email, calendars, events and contacts, and send email or create events.
MCP server for Appcircle mobile CI/CD platform.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceMCP server for iCloud (Apple) Calendar access via CalDAV19-
- AlicenseNot gradedqualityDmaintenanceMCP server for Apple Calendar, Mail, Reminders, and Files on macOS using native frameworks.15 npm21MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for iCloud integration, enabling management of calendars (CalDAV), contacts (CardDAV), and email (IMAP/SMTP) through natural language.MIT
- -licenseNot gradedqualityNot gradedmaintenanceAn MCP server for iCloud Calendar, Contacts, and Mail, usable from any MCP-capable AI client.-