Apple Tools MCP
Provides semantic search and data access across Apple Mail, Messages, Calendar, and Contacts on macOS, enabling natural language queries to find emails, iMessages, calendar events, and contact information stored locally on the device.
Enables semantic search of iMessages and SMS conversations, retrieving recent messages, viewing full conversation histories, and listing message contacts through natural language queries.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Apple Tools MCPFind emails from John about the quarterly report"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
apple-tools-mcp
An MCP (Model Context Protocol) server for Apple Mail, Messages, Calendar, and Contacts on macOS. Search them in plain language, send mail and messages, and create, edit, or remove calendar events and contacts. It talks to any MCP client over stdio. Search embeddings and the write bridge stay on your Mac.
Requirements
macOS
Node.js 18 or later
The permissions below, on the Mac that holds the Mail, Messages, Calendar, and Contacts data
Related MCP server: MCP Apple Notes
Setup
1. Install
npm install -g apple-tools-mcpOr from source:
git clone https://github.com/sfls1397/Apple-Tools-MCP.git
cd Apple-Tools-MCP
npm installA global install uses npx in the client config below. A clone uses node and the full path to index.js.
After you install or upgrade, run the permissions command in the next section. npm install only prints a reminder. It cannot click Allow for you.
2. Permissions
Two grants, and they do different jobs.
Full Disk Access lets the server read the Mail, Messages, Calendar, and Contacts databases.
Automation lets it send mail, send messages, and change calendar events and contacts through those apps. Mail compose also needs Accessibility, because the message body is pasted into Mail's window through System Events.
Full Disk Access
In Terminal, run
which nodeand copy the path it prints.Open System Settings → Privacy & Security → Full Disk Access.
Click +.
Press Cmd+Shift+G, paste the path, and press Enter.
Select
nodeand turn its toggle on.
Automation, Accessibility, and System Events
Run this from Terminal.app, so the prompts attach to node and not to your editor or chat app:
$(which node) $(which apple-tools-mcp) permissionsFrom a clone, npm run permissions does the same thing. npx apple-tools-mcp permissions also works.
The command prints process.execPath. That path is the node binary macOS will ask about. Click Allow for that binary. If it warns that the command you typed started a different node, run it again with the printed path in front:
/absolute/path/to/node $(which apple-tools-mcp) permissionsAllow node to control Mail, Messages, Contacts, Calendar, and System Events. Each app is its own switch. Allowing Contacts does not allow Mail.
Then open System Settings → Privacy & Security → Accessibility and turn on the same node binary. Without that, Mail compose cannot paste the body.
The command does real work so the prompts can appear. It opens a Mail compose window and throws it away (nothing is sent), lists Messages accounts (nothing is sent), creates and deletes a throwaway contact, lists calendars without creating events, and checks System Events. It exits with an error until Mail, Messages, Contacts, Calendar, and System Events are all allowed.
dry_run on mail_send or messages_send does not talk to those apps, so a dry run will not pop the prompts and will not tell you whether Automation is allowed. A real send that hangs after you denied Mail is an Automation denial, not "Mail could not be reached."
If you clicked Don't Allow, macOS will not ask again. Open System Settings → Privacy & Security → Automation, turn the switch on for that node and that app, then run apple-tools-mcp permissions again.
Do not add node with the + button under Privacy & Security → Contacts or Calendars. Those lists are not how these writes are granted, and on current macOS they often have no Add button.
Leave Mail, Messages, and Contacts running. If they are quit, writes to them often fail even when Automation is allowed. Calendar does not need to stay open.
If your chat app cannot hold the grants
macOS treats these actions as coming from the app that launched the server. A chat app that starts a short-lived MCP process is that app. Some of them cannot be granted Contacts or Calendar access, no matter what you allow for node.
Run the indexer in the next section. It is a normal node process, so the Automation grants you gave node apply. Your MCP client sends writes to it through a local socket at ~/.apple-tools-mcp/writer.sock. When a prompt appears, allow node, not the chat app.
3. Connect your MCP client
Any stdio MCP client can run the server. The command is the same; only the client's settings screen changes.
Global install:
{
"mcpServers": {
"apple-tools": {
"command": "npx",
"args": ["-y", "apple-tools-mcp"]
}
}
}From a clone, use node and the absolute path to index.js instead of npx.
Claude Desktop keeps this in ~/Library/Application Support/Claude/claude_desktop_config.json. Quit the client completely and reopen it after you save.
The MCP process is short-lived: it exits when the client disconnects. Ongoing indexing belongs on the indexer, not on a wrapper that holds this process open.
When stdin closes, requests already received still get their replies before the process exits (up to two minutes). A script can pipe a whole session in one go:
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"script","version":"0"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"mail_search","arguments":{"query":"invoice","limit":5}}}' \
| apple-tools-mcp4. Indexer
On first use the server builds a search index of your mail, messages, and calendar events. That can take a while if you have a lot of mail. The index is stored in ~/.apple-tools-mcp/vector-index/.
Run the indexer so the index stays current and so writes go through node:
apple-tools-indexerThat command is node index.js --mode=indexer. A global install provides apple-tools-indexer (bin/apple-tools-indexer.js). From a clone, use npm run indexer.
To keep it running across logins, use a LaunchAgent. LaunchAgents do not see your shell PATH, so put in the absolute paths from which node and npm root -g.
Save this as ~/Library/LaunchAgents/com.apple-tools-mcp.indexer.plist and replace both paths:
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.apple-tools-mcp.indexer</string>
<key>ProgramArguments</key>
<array>
<string>/absolute/path/to/node</string>
<string>/absolute/path/to/node_modules/apple-tools-mcp/index.js</string>
<string>--mode=indexer</string>
</array>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>StandardOutPath</key>
<string>/tmp/apple-tools-indexer.out.log</string>
<key>StandardErrorPath</key>
<string>/tmp/apple-tools-indexer.err.log</string>
</dict>
</plist>launchctl load ~/Library/LaunchAgents/com.apple-tools-mcp.indexer.plistThe refresh interval is chosen once, when the process starts. The default is 5 minutes. Set INDEX_INTERVAL_MS (30s, 1m, 1h, or a number of milliseconds) or add this to ~/.apple-tools-mcp/config.json:
{
"indexInterval": "1m"
}Values are clamped to 15 seconds through 6 hours. A missing config file is fine; the 5-minute default is used.
If you do not run the indexer, the MCP process indexes when it starts, then exits when the client disconnects. Writes then run inside that process, and macOS attributes them to the app that launched it.
5. Another computer (optional)
The same tools are available over HTTP, with a bearer token on every request. stdio on this Mac does not use the token.
apple-tools-http
apple-tools-mcp http-token~/.apple-tools-mcp/config.json (the same file as the index interval):
{
"httpHost": "127.0.0.1",
"httpPort": 8421
}127.0.0.1 accepts connections only on this Mac. Set httpHost to a LAN or Tailscale address when a client on another machine should connect. APPLE_TOOLS_HTTP_HOST and APPLE_TOOLS_HTTP_PORT override the file. Do not send the token over an open network.
An example LaunchAgent is examples/com.apple-tools-http.plist. How to point Claude and other clients at it, including apple-tools-http-proxy, is in examples/mcp-client.md.
Index commands
Stop the indexer before a manual rebuild so two processes are not writing the index at once.
# Rebuild (all email history)
npm run build-index
# Optional shorter email lookback
APPLE_TOOLS_INDEX_DAYS_BACK=30 npm run build-indexYour client can also call rebuild_index.
Force a clean rebuild when the index is corrupt or out of date:
launchctl unload ~/Library/LaunchAgents/com.apple-tools-mcp.indexer.plist
rm -rf ~/.apple-tools-mcp/vector-index
rm -f ~/.apple-tools-mcp/index-meta.json
rm -f ~/.apple-tools-mcp/indexer.lockStart the indexer or your MCP client again so it builds a new index.
Watch progress. With the LaunchAgent above:
tail -f /tmp/apple-tools-indexer.err.logClient logs depend on the app. Claude Desktop:
tail -f ~/Library/Logs/Claude/mcp-server-apple-tools.logAudit coverage against the source databases:
npm run audit
npm run audit -- --reporter=verbose > audit-report.txtYour client can also call audit_index.
Available tools
Universal search
Tool | Description |
| Search across mail, messages, and calendar. Picks the sources from the question. |
| Find communication with one person across Mail, Messages, and Calendar. |
Tool | Description |
| Semantic search, with filters for sender, recipient, attachments, and mailbox. |
| Most recent emails. Can limit to unread. |
| Emails from a date such as "today", "yesterday", or "Nov 13". |
| Full email content by file path. |
| Most frequent senders. |
| Every email in a conversation. |
| Exact lookup in Mail's own database by subject prefix or text, sender, recipient, mailbox, and date. Returns every copy with its Message-ID. Does not wait on the index. |
Messages
Tool | Description |
| Semantic search for iMessage and SMS, with filters. |
| Most recent messages. |
| Full conversation with a contact. |
| Contacts you have messaged. |
Calendar
Tool | Description |
| Semantic search for events, with filters. |
| Events on a date, read live from Calendar rather than the search index. |
| Next upcoming events. |
| Events for this week or a future week. |
| Open time on a date, read live from Calendar rather than the search index. |
| Recurring events. |
Contacts
Tool | Description |
| Search by name, email, phone, or organization. |
| Full details for one contact. |
Index
Tool | Description |
| Rebuild the search index for one source or all of them. |
| Check index health and coverage. |
Write tools
Write tools change your data. They do not use the search index, so they work while the indexer is rebuilding.
Every write tool accepts:
Argument | Meaning |
| Preview only. Nothing is sent or changed. This wins even if |
| Approves an action that is otherwise blocked. |
Deletes need confirm: true: mail_trash, calendar_remove, contacts_remove, and contacts_edit when it clears every email or phone. These take one id per call, except mail_trash, which also takes a list (message_ids).
Sends to more than one person need confirm: true: mail_send and mail_forward when To, Cc, and Bcc add up to more than one address, mail_reply with reply_all: true, and messages_send to several people or to a group chat. A single recipient sends on the first call. mail_draft never sends, so it does not need that confirm.
A dry run or a call that still needs confirm does not deliver anything. The result is ok: false, planned: true, delivered: false, and the client may mark it as an error. The text says DRY RUN or CONFIRMATION REQUIRED. Run it again with dry_run false, and with confirm: true when the tool asked for it.
Responses name the action, ids, and recipients. They do not echo message bodies.
A missing or malformed id or recipient is refused. Calendar times are local datetimes, YYYY-MM-DD HH:MM. Phrases like "next Tuesday" are rejected on writes.
Tool | Arguments | Confirm |
|
| when recipients > 1 |
|
| no (saved to Drafts) |
|
| when |
|
| when recipients > 1 |
|
| no |
|
| no |
|
| yes |
Pass the RFC822 Message-ID, or the file_path from mail_search or mail_recent and the server reads the Message-ID from the file. mail_archive moves the message to that account's Archive (or All Mail). mail_trash moves it to Trash.
mail_trash with message_ids moves every copy of each message, for example the INBOX copy and the Sent copy of a mail you sent yourself, to that account's Trash. It finds each copy through Mail's database (mail_find uses the same lookup) and checks its Message-ID again before it moves it. One call stops starting new moves after about 35 seconds and reports remaining. Run it again for the rest. Copies that are already in Trash are left alone. Nothing is deleted for good. Mail empties Trash on its own schedule.
body_format: "html" pastes the text into Mail's compose window. Mail may still create its own HTML version. New mail is a new message, not a quoted reply. Reply and forward still include the original.
A send is successful only after the message is in Sent or still in Outbox. If the call times out, check Sent before you try again. Retrying a message that already went out sends a second copy. A timeout is not the same thing as a denied Automation prompt.
Compose needs the Accessibility and System Events grants from the permissions section.
Messages
Tool | Arguments | Confirm |
|
| several handles, or a group chat |
tois a phone number in E.164 form (+15551234567) or an Apple ID email.chat_idis an existing conversation id, such asiMessage;-;+15551234567oriMessage;+;chat123456789. It is checked against Messages before anything is sent. That lookup is also how group chats are detected.service:auto(default) tries iMessage and can fall back to SMS;imessageandsmspin the service.attachment_pathis an absolute path to a file that already exists on this Mac. Text and a file can go together.
Calendar
Tool | Arguments | Confirm |
| none | no. Lists calendars so you can pick one. |
|
| no |
|
| no |
|
| yes |
|
| no |
calendar_date and calendar_add report the event id. Pass that id to edit, remove, and RSVP. Run calendar_list_calendars first so a new event lands on the calendar you mean. calendar_add also returns eventkit_id; pass it back on edit and remove.
Recurrence is either structured fields or one raw rule:
Argument | Values |
|
|
| 1–366 |
| 1–1000 occurrences. Do not combine with |
| local datetime |
|
|
| a raw rule such as |
alerts_minutes_before takes up to five values, from 0 to 40320 minutes (four weeks). On edit, supplying alerts replaces the existing ones. replace_alerts: true removes them without adding new ones.
calendar_rsvp sets your response on an invitation. If that macOS version will not write it, the tool says so instead of pretending it worked.
Contacts
Tool | Arguments | Confirm |
|
| no |
|
| yes, when replacing with an empty list |
|
| yes |
Use the contact id from contacts_search, contacts_lookup, or contacts_add. Creating a contact needs at least one of first_name, last_name, or organization. Writes go through Contacts, which keeps iCloud in sync.
Examples
{ "name": "mail_send", "arguments": { "to": ["a@example.com"], "subject": "Status", "body": "All good", "dry_run": true } }
{ "name": "mail_send", "arguments": { "to": ["a@example.com"], "subject": "Status", "body": "All good" } }
{ "name": "mail_send", "arguments": { "to": ["a@example.com", "b@example.com"], "subject": "Status", "body": "All good", "confirm": true } }
{ "name": "calendar_add", "arguments": { "calendar_name": "Work", "title": "Standup", "start": "2026-09-21 09:00", "end": "2026-09-21 09:15", "frequency": "weekly", "by_day": ["MO","TU","WE","TH","FR"], "alerts_minutes_before": [10] } }
{ "name": "calendar_remove", "arguments": { "event_id": "EVT-UID", "confirm": true } }This package does not include Reminders, Notes, FaceTime, or Files.
Example questions
"Find emails from John about the quarterly report"
"What messages did I get from Mom last week?"
"When is my next dentist appointment?"
"Search for emails about the AWS bill from November"
"Find all calendar events with Zoom links"
"What's Sarah's phone number?"
"Show me all communication with David from last month"
Privacy
Search and the index read your mail, messages, calendar, and contacts on this Mac. Write tools send messages and change events and contacts. Deletes and multi-recipient sends wait for confirm.
Embeddings are computed locally. Nothing is uploaded. The write bridge is a socket in your home directory, not a network port. The server stores no passwords; it uses the Apple apps you are already signed into. The index stays in ~/.apple-tools-mcp/.
Troubleshooting
"Authorization denied" or the databases will not open. Full Disk Access is missing or it was granted to a different node than the one that is running. Repeat the Full Disk Access steps with which node.
A write is denied, or sending mail hangs. Automation was not granted for the node that is actually running. Run apple-tools-mcp permissions from Terminal.app and allow the printed process.execPath for Mail, Messages, Contacts, Calendar, and System Events. Allow node, not the chat app. If you previously clicked Don't Allow, turn the switch on in Automation and run the command again. dry_run does not talk to Mail, so it will not catch this.
"CONFIRMATION REQUIRED". The safety check worked. Pass confirm: true, or use dry_run: true if you only wanted a preview.
Search returns nothing. Check ls ~/.apple-tools-mcp/vector-index/. If it is missing or stale, rebuild with the commands above.
The server does not show up in the client. Confirm the config is valid JSON, then quit the client completely and reopen it.
Development
git clone https://github.com/sfls1397/Apple-Tools-MCP.git
cd Apple-Tools-MCP
npm install
npm test
npm run indexer
npm run build-index
npm run audit
npm run smoke:writes
npm run smoke:writes -- --applynpm run smoke:writes checks the write path without creating lasting contacts or events. npm run smoke:writes -- --apply performs real creates, edits, and deletes, then removes the test items. The extra -- is required so npm forwards --apply. --apply fails if Mail, Messages, Contacts, or Calendar Automation is missing.
License
MIT License. See LICENSE.
Acknowledgments
Available Tools
39 toolsaudit_indexB
Audit search index against source data with 0% tolerance. Reports missing items, orphaned entries, and duplicates with detailed file paths and remediation suggestions. Validates 100% of source data.
| Name | Required | Description | Default |
|---|---|---|---|
| sources | No | Data sources to audit (default: all) | |
| max_items | No | Max items to list per category (default: 100, use 0 for unlimited) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses scope and output behavior (0% tolerance, 100% validation, remediation suggestions implying no auto-fix), but never states that the operation is read-only, whether it has side effects, or any cost/runtime characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and scope. Mostly efficient, though 'with 0% tolerance' and 'Validates 100% of source data' restate the same strictness idea.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-optional-parameter audit tool with no output schema and no annotations, the description explains what categories of results it returns (missing, orphaned, duplicates, file paths, remediation). Only the mutation/side-effect question remains open.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both optional parameters (sources with enum values, max_items with its 0-for-unlimited semantics). The description adds no parameter meaning beyond that, which is the correct baseline here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audit) and resource (search index) and enumerates the findings it produces: missing items, orphaned entries, duplicates. It is clearly distinct from index-building siblings, though it never explicitly contrasts itself with rebuild_index.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no named alternative. The natural relationship to the sibling rebuild_index (audit detects issues, rebuild fixes them) is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_addB
Create a calendar event. Times must be explicit local datetimes (YYYY-MM-DD HH:MM). Supports recurrence (FREQ=DAILY/WEEKLY/MONTHLY/YEARLY with INTERVAL, COUNT or UNTIL, BYDAY) and alerts in minutes before the start. Returns the new event id.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | End as YYYY-MM-DD HH:MM local time (default: start + 1 hour) | |
| count | No | Number of occurrences (1-1000); cannot be combined with until | |
| notes | No | Event notes / description | |
| start | Yes | Start as YYYY-MM-DD HH:MM local time, or YYYY-MM-DD for all-day (required) | |
| title | Yes | Event title (required) | |
| until | No | Repeat until this local datetime (YYYY-MM-DD HH:MM) | |
| by_day | No | Weekly by-day list: MO TU WE TH FR SA SU | |
| all_day | No | Create an all-day event | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| interval | No | Recurrence interval, e.g. 2 for every other week (1-366) | |
| location | No | Event location | |
| frequency | No | Recurrence frequency | |
| recurrence | No | Raw RRULE instead of the structured fields, e.g. FREQ=WEEKLY;INTERVAL=1;COUNT=10 | |
| calendar_name | Yes | Target calendar name from calendar_list_calendars (required) | |
| alerts_minutes_before | No | Display alerts in minutes before the start (max 5, up to 40320) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the required time format, recurrence capability, alert behavior, and that the new event id is returned, but says nothing about permissions, side effects, or how the confirm/dry_run preview behavior (documented only in the schema) interacts with creation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core purpose, then format, then capabilities, then return value. Efficient and well-ordered, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter create tool with no annotations and no output schema, the description covers the essentials: purpose, time format, recurrence, alerts, and the return value. The comprehensive schema supplies the rest, so the remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 16 parameters thoroughly, establishing a baseline of 3. The description reinforces the recurrence RRULE syntax and time format but adds little beyond what the schema fields already state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Create a calendar event"), which cleanly separates it from the mutation siblings calendar_edit and calendar_remove. It doesn't explicitly name those alternatives, but the create semantics are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to use this tool versus calendar_edit, calendar_recurring, or calendar_remove, and offers no preconditions or exclusions. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_dateA
Get all events on a specific date. Supports natural language dates like 'today', 'tomorrow', 'next Tuesday', 'Jan 15'.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | Date to check (e.g., 'today', 'tomorrow', 'next Monday', '2025-01-15') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that natural-language dates are accepted, but says nothing about return format, which calendars are covered, permissions, or error behavior on an unparseable date.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the purpose front-loaded and a concrete format hint second; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the definition covers what the agent must know to invoke it correctly. Only the output/scope details (returned fields, calendar coverage) are absent, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'date' parameter is already fully documented with examples. The description's sentence about NL dates overlaps with the schema and adds no syntactic detail beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get all events') scoped to 'a specific date', which naturally distinguishes it from siblings like calendar_week and calendar_upcoming. It stops short of naming those siblings explicitly, so differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope ('on a specific date') implies when to use it versus calendar_week, but there is no explicit when-to-use guidance or named alternative. An agent can infer the routing, but nothing tells it to prefer calendar_search or calendar_free_time in other situations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_editA
Update an existing event by id. Pass only the fields to change. Setting alerts_minutes_before replaces the event's existing alerts.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | New end as YYYY-MM-DD HH:MM local time | |
| count | No | Number of occurrences (1-1000) | |
| notes | No | New notes / description | |
| start | No | New start as YYYY-MM-DD HH:MM local time | |
| title | No | New title | |
| until | No | Repeat until this local datetime | |
| by_day | No | Weekly by-day list: MO TU WE TH FR SA SU | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| event_id | Yes | Event ID from calendar_date or calendar_add (required) | |
| interval | No | Recurrence interval (1-366) | |
| location | No | New location | |
| frequency | No | Recurrence frequency | |
| recurrence | No | Raw RRULE | |
| eventkit_id | No | EventKit eventIdentifier from calendar_add (via EventKit). Prefer this over Calendar.app uid lookup. | |
| replace_alerts | No | Remove existing alerts even when no replacements are given | |
| alerts_minutes_before | No | Replacement alerts in minutes before the start |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden; it usefully discloses that setting alerts_minutes_before replaces existing alerts (a destructive side effect) and that updates are partial. It still omits permission requirements and the fact that deletes/multi-recipient sends only preview without confirm, leaving meaningful mutation behavior undocumented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse sentences, front-loaded with the core operation and each carrying distinct information (operation, patch semantics, alert side effect). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter mutation tool with no annotations and no output schema, the schema does heavy lifting on parameter detail, and the description covers the essential update/replacement contract. It nonetheless leaves the confirm/dry_run behavior and required permissions for the agent to infer from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 17 parameters are already documented in the schema, making baseline 3 appropriate. The description adds one useful nuance (alerts_minutes_before replacement semantics) but nothing else beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update an existing event by id'), which cleanly distinguishes it from the sibling write tools calendar_add and calendar_remove. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Pass only the fields to change' gives the patch/partial-update contract, which is genuine usage guidance. However, it gives no when-to-use vs alternatives (e.g. calendar_add vs calendar_edit) and does not explain the dry_run/confirm preview workflow that governs whether a call actually mutates state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_free_timeB
Find free time slots on a specific date. Analyzes calendar to find available time windows.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | Date to check (e.g., 'today', 'tomorrow', 'next Monday') | |
| end_hour | No | End of working hours (default 17 = 5 PM) | |
| start_hour | No | Start of working hours (default 9 = 9 AM) | |
| calendar_name | No | Only consider events from this calendar |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden. 'Find free time slots' and 'analyzes calendar' implicitly signal a non-mutating read operation, which is meaningful, but it never states read-only behavior, permission requirements, or how working-hour defaults affect results beyond what the schema says.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and scope. The second sentence is close to a restatement of the first but adds the 'available time windows' framing, keeping waste minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter read tool with no output schema and no annotations, the description covers the purpose but omits how results are shaped or constrained (e.g., whether defaults 9-17 bound the output, granularity of slots). Just enough to call it, not enough to predict its behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents date, start_hour, end_hour, and calendar_name; baseline is 3. The description adds no extra meaning about how the working-hours range scopes the returned slots or how calendar_name filters analysis.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find) and resource (free time slots) with a scope constraint (a specific date), and 'Analyzes calendar to find available time windows' clarifies the mechanism. It does not, however, distinguish itself from siblings like calendar_date or calendar_upcoming, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the purpose (use this when you want availability rather than a list of events), but there is no explicit when-to-use/when-not framing and no named alternative among the many calendar siblings. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_calendarsA
List the calendars in Calendar.app with their writability, so events land on the intended calendar instead of the default.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It is a zero-parameter read-only lister, and the description does disclose that the response includes writability information, which is the key behavioral detail. It does not state read-only nature or whether any calendar types (shared, subscribed) are omitted, leaving modest gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action first and the payoff last; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, no-annotation tool with no output schema, the description supplies both the action and the shape of the useful return field (writability). It is sufficient to select and call the tool; only edge cases about what calendars are excluded are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline of 4 applies. There is nothing for the description to clarify beyond the return content it already mentions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the calendars in Calendar.app') and adds what the result contains ('with their writability'), so it is clearly a discovery tool. It does not explicitly name a sibling it is distinct from, but its purpose is unambiguous against calendar_upcoming/calendar_week/calendar_add.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'so events land on the intended calendar instead of the default' gives a concrete motivating context: call this before choosing a calendar for an event. No explicit exclusions or alternatives are named, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_recurringB
List recurring events (events that appear multiple times). Shows upcoming occurrences.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum recurring events to return (default 30) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden, yet it discloses almost nothing beyond what the name implies. It does not state whether this is read-only, whether the list is sorted or paginated, what time window 'upcoming' spans, or how occurrences are represented. The parenthetical definition is the only real addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, front-loaded with the core purpose and followed by a brief scope note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and several similarly named calendar siblings, the description is too thin. It fails to explain the return shape, sort order, how recurrence instances are counted against the limit, or how it differs from calendar_upcoming, leaving real gaps an agent would want closed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single limit parameter, so the schema already documents default (30) and meaning. The description adds nothing about the parameter, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recurring events') with a clarifying parenthetical defining the resource. It does not, however, explicitly differentiate itself from the closely named siblings calendar_upcoming and calendar_week, which also deal with upcoming occurrences.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied. 'Shows upcoming occurrences' hints at a time-scoped view, but the description never says when to choose this over calendar_upcoming or calendar_week, nor whether it is a filtered variant. An agent must infer the selection condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_removeA
Delete a calendar event by id. Destructive: requires confirm=true, otherwise the call only reports what would be deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| event_id | Yes | Event ID from calendar_date or calendar_add (required) | |
| eventkit_id | No | Optional EventKit eventIdentifier returned by calendar_add (via EventKit). Needed so writeOnly EventKit can delete the event it created. | |
| calendar_name | No | Optional calendar from calendar_list_calendars; narrows the AppleScript delete so Calendar does not scan every calendar |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key behavior: this is destructive and gated behind confirm=true, otherwise it degrades to a report-only preview. It omits permissions/auth requirements and the irreversibility/effect on recurring events, which would round it out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the destructive nature and confirm gate front-loaded immediately after the purpose statement. Nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the essential safety behavior (confirm gate, preview fallback) and the schema fully documents parameters. Auth/permission expectations and side effects (recurrence handling, notification of attendees) are the notable remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters including confirm, dry_run, and eventkit_id. The description only restates event_id ('by id') and confirm, adding no syntax or meaning beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a calendar event by id'), which naturally separates it from calendar_edit, calendar_add, and calendar_search. It does not explicitly name a sibling to differentiate against, but the destructive verb is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit call-time rule: confirm=true is required, and without it the call only previews. That tells the agent how to invoke it safely, though it does not point to alternatives (e.g., calendar_edit) for modifying rather than removing an event.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_rsvpB
Respond to a calendar invitation: accept, decline, or tentative.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| event_id | Yes | Event ID of the invitation (required) | |
| response | Yes | accept, decline, or tentative (required) | |
| attendee_email | No | Your invited address, when the event lists several attendees |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, yet it says nothing about side effects (this mutates the user's RSVP status), permission requirements, or that confirm/dry_run control previewing versus execution. The schema mentions preview semantics, but the description itself adds no behavioral context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste and the action options listed inline. It is efficient, though slightly under-specified rather than elegantly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The rich schema (5 params, 100% coverage, enum) and absence of an output schema reduce the burden, but with no annotations the description should at least note that this changes state and how preview works. It is minimally adequate but leaves a mutation-behavior gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description only restates the response enum values that the schema already documents, adding no new meaning beyond what the structured fields provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (respond) and resource (calendar invitation) and enumerates the three response actions, so the agent can identify it as the RSVP tool distinct from calendar_edit or calendar_remove. It stops short of explicitly naming a sibling it is not, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Respond to a calendar invitation' implies the usage context, but no when-to-use vs alternatives guidance is given and no exclusions or prerequisites (e.g., needing an event_id from calendar_search) are stated. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_searchB
Semantic search for calendar events using AI embeddings. Finds events by meaning. Supports filtering by calendar name and all-day events.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (default 30) | |
| query | Yes | Natural language search (e.g., 'meetings', 'doctor appointments', 'lunch') | |
| sort_by | No | Sort by relevance (default) or date (chronological) | |
| days_back | No | Include events from last N days (0 = none) | |
| days_ahead | No | Include events in next N days (0 = none). Use for 'today', 'this week', etc. | |
| all_day_only | No | Only show all-day events | |
| calendar_name | No | Filter to specific calendar (e.g., 'Work', 'Personal') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the semantic/embedding mechanism, which is genuine behavioral context beyond the schema, but says nothing about permissions, read-only nature, performance characteristics, or result semantics for a tool that returns AI-ranked matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with the core purpose first and capability notes after. Minor redundancy between 'semantic search... using AI embeddings' and 'finds events by meaning', but nothing is bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter search tool with no output schema and no annotations, the description is minimally adequate. It omits how results are returned or ranked, whether the search spans all calendars by default, and any coverage of the date-window parameters that meaningfully shape results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters including examples, defaults, and enum values. The description's mention of calendar name and all-day filtering merely restates a subset of what the schema provides, adding no syntax or behavioral detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ('Semantic search for calendar events') and clarifies the mechanism ('using AI embeddings', 'finds events by meaning'). This distinguishes it from keyword/date siblings like calendar_date and calendar_upcoming, though it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it ('by meaning' rather than by date), but never states explicit conditions, exclusions, or which sibling to prefer for date-based lookups (calendar_date, calendar_upcoming, calendar_week). Usage is inferable rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_upcomingA
Get next N upcoming events across all calendars. Simpler than calendar_search for quick schedule overview.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum events to return (default 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden. It does disclose behavior beyond the schema: events are aggregated across ALL calendars and ordered as the 'next' upcoming ones. However it says nothing about permissions, timezone handling, whether recurring/all-day events are included, or pagination, which are meaningful gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler, and the core scoping fact ('across all calendars') is front-loaded before the sibling comparison. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only list tool with no output schema, the description covers what it returns and how it differs from alternatives. It omits ordering/timezone specifics, but the schema carries the only parameter and no return structure needs explaining.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the single 'limit' parameter already documents itself with its default of 10. The description only echoes this as 'N' without adding format or constraint detail, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get next N upcoming events') plus a clear scope ('across all calendars'), and explicitly differentiates from a sibling by calling itself 'simpler than calendar_search'. An agent can distinguish this from search and the other calendar tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (calendar_search) and gives a condition for preferring this tool ('quick schedule overview'). It stops short of stating explicit exclusions or prerequisites, so it is clear context without full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_weekA
Get all events for the current week or a future week. Shows events grouped by day.
| Name | Required | Description | Default |
|---|---|---|---|
| week_offset | No | 0 = this week, 1 = next week, 2 = week after, etc. (default 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it discloses only that results are grouped by day. It does not state whether all calendars are included, how the week boundary/timezone is defined, or that the operation is read-only and side-effect free.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the output shape; nothing is wasted. It is near-minimal, though it leaves room for a brief routing hint at no real length cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter read tool with a fully documented schema and no output schema, the description plus schema covers what an agent needs to invoke it. Remaining gaps (calendar selection, week-boundary definition) are minor for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter's offset semantics are fully documented in the schema ('0 = this week, 1 = next week'). The description's 'current week or a future week' corroborates that negative offsets are not intended but adds no syntax or format detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get all events') with a clear scope ('current week or a future week'), which lets an agent separate it from date-range siblings like calendar_date and calendar_search. It does not explicitly name the sibling it competes with, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'current week or a future week' implicitly tells the agent this is the week-granularity view rather than a date-range or upcoming-N-events view, but no alternative tool is named and no exclusions are given. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_addC
Create a contact in Contacts.app. Returns the new contact id.
| Name | Required | Description | Default |
|---|---|---|---|
| emails | No | Email addresses (max 10) | |
| phones | No | Phone numbers (max 10) | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| job_title | No | Job title | |
| last_name | No | Last name | |
| first_name | No | First name | |
| email_label | No | Label for the emails, e.g. work or home (default work) | |
| phone_label | No | Label for the phones, e.g. mobile or home (default mobile) | |
| organization | No | Company / organization |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers almost nothing: it does not mention duplicate handling, what happens on partial failure, auth/permission needs, or how the dry_run/confirm semantics in the schema apply to a create. Saying only that it returns the new contact id leaves the mutation profile undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, operation front-loaded before the return value, with zero filler. Nothing in the text is redundant with itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 10-parameter mutation tool with no annotations and no output schema needs more than a single-line statement. Nothing explains that zero parameters are required or what minimum identity (name/email/phone) makes a useful contact, and the write semantics are left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is documented in the schema itself (including the max-10 email/phone limits and the confirm/dry_run behavior). The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a contact in Contacts.app'), which cleanly separates it from the sibling contacts_edit and contacts_remove. It does not explicitly name those alternatives, so it falls short of a 5, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no pointer to alternatives such as contacts_edit for updating an existing contact. The agent must infer that this is the creation path from the verb alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_editB
Update a contact by id. Emails and phones are added unless replace_emails / replace_phones is set; replacing with an empty list removes them and requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| emails | No | Email addresses to add (or to replace with) | |
| phones | No | Phone numbers to add (or to replace with) | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| job_title | No | New job title | |
| last_name | No | New last name | |
| contact_id | Yes | Contact ID from contacts_search or contacts_lookup (required) | |
| first_name | No | New first name | |
| email_label | No | Label for those emails (default work) | |
| phone_label | No | Label for those phones (default mobile) | |
| organization | No | New organization | |
| replace_emails | No | Remove existing emails before adding | |
| replace_phones | No | Remove existing phones before adding |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose genuinely useful traits absent from the schema: that emails/phones are additive by default, that an empty list plus replace removes them, and that confirm=true is required for that removal. However it omits permissions, reversibility of other field changes, and partial-update behavior for unmentioned fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action ('Update a contact by id') front-loaded before the caveats. Every clause earns its place, though the confirm/preview interaction could be surfaced slightly more cleanly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation tool with no annotations and no output schema, the description is on the thin side. It covers the tricky delete/preview path but says nothing about partial-update behavior for unmentioned fields, permissions, or failure modes that an agent would need before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema lacks: the add-vs-replace interaction between emails/phones and their replace_* flags, plus the empty-list-removal rule requiring confirm. That is real meaning beyond the individual parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (a contact by id), which cleanly separates it from the sibling mutation tools contacts_add and contacts_remove. It doesn't explicitly name those siblings, but the edit/add/remove triad makes the intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to prefer this over contacts_add or contacts_lookup, nor does it state prerequisites beyond the implicit 'need an id'. It describes mutation semantics (add vs replace) rather than offering any usage routing or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_lookupB
Look up a specific contact by email, phone number, or name. Returns full contact details including all emails and phone numbers.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | Email address, phone number, or name to look up |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the return contents ('full contact details including all emails and phone numbers'), which is useful, but says nothing about behavior when the identifier is not found, when a name matches multiple contacts, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the action and accepted inputs are front-loaded and the return behavior follows immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup with no output schema, the description does helpfully state the return shape, compensating for the missing output schema. However it omits not-found behavior and ambiguity handling for name matches, which are the main risks for an agent calling this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single identifier parameter, and the description merely restates the same three accepted forms (email, phone number, name). Baseline 3 is correct when the schema already fully documents the parameter and the description adds no format or precedence detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up) and resource (a contact) plus the three accepted key types. It is clear what the tool does, but it never distinguishes itself from the overlapping siblings contacts_search and person_search, so an agent must guess which lookup path to take.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a specific contact' faintly implies singular, exact retrieval versus a search, but there is no explicit when-to-use guidance, no exclusions, and no routing to contacts_search or person_search despite strong sibling overlap. Nothing tells the agent when this should be preferred over the alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_removeA
Delete a contact by id. Destructive: requires confirm=true, otherwise the call only reports what would be deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| contact_id | Yes | Contact ID from contacts_search or contacts_lookup (required) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does well on the critical point: it flags the operation as destructive and explains the confirm-gate that converts a real delete into a read-only preview. It omits whether deletion is permanent/reversible and what permissions are needed, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the destructive nature and the gate condition front-loaded; there is no filler and every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter-required mutation tool with no output schema and no annotations, the description covers the essential risk (destructive) and the safety mechanism (confirm gate / preview). It leaves minor gaps around permanence of the delete and what the preview output contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters (contact_id, confirm, dry_run) are already documented in the schema. The description restates the confirm semantics rather than adding syntax, formats, or interactions the schema lacks, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ("Delete") plus resource ("a contact") and identifies the key ("by id"), which cleanly separates it from siblings like contacts_edit, contacts_add, or calendar_remove. An agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the operating condition explicitly: confirm=true is required to actually delete, otherwise the call previews. That is clear when-to-use context. It does not name an alternative tool (e.g., contacts_edit when the goal is modification rather than removal), so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_searchB
Search your contacts by name, email, phone, or organization. Returns matching contacts with all their details.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (default 30) | |
| query | Yes | Search query (e.g., 'John', 'Acme Corp', 'john@example.com') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read-only search and notes that results include all contact details, which is useful, but it omits permissions, pagination behavior tied to limit, and any other operational traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and scope, then the return behavior. Efficient with no filler, though it is not rich enough to merit a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description is adequate: it covers what can be searched and that full details are returned. It does not address paging/limit semantics or sibling disambiguation, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (query and limit) are already fully documented in the schema. The description only restates the searchable fields, adding essentially nothing beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (contacts) plus the searchable fields (name, email, phone, organization), so the core function is unmistakable. However, it does not distinguish itself from siblings like contacts_lookup or person_search, which likely overlap in intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when not to use it, and no reference to the closely related contacts_lookup or person_search siblings. The agent must infer the selection criteria with no help from the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_archiveC
Move an email to its account's Archive mailbox.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| file_path | No | Alternative to message_id: the .emlx file_path from mail_search, resolved to its Message-ID | |
| message_id | No | RFC822 Message-ID of the email (from mail_search / mail_read results) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says 'Move', implying mutation, but never states whether the email is recoverable from Archive, whether the original folder is recorded, or that the schema's confirm/dry_run flags gate a preview-vs-commit flow. That preview behavior is material and only discoverable by reading parameter docs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though its brevity contributes to the gaps elsewhere rather than compensating for them.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the terse description is the only source of behavioral framing, and it omits the preview/confirm mechanics, result reporting, and error behavior. Notably, the confirm parameter's schema text mentions deletes and multi-recipient sends, which reads as generic boilerplate unrelated to archiving and is not reconciled in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, including the file_path alternative to message_id. The description adds nothing about parameter usage, which is acceptable at this coverage level but earns no bonus.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Move an email to its account's Archive mailbox'), which is clearly distinguishable from mass operations. It does not explicitly distinguish itself from the closest sibling, mail_trash, leaving the agent to infer that Archive is a non-destructive relocation rather than a delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusion of alternatives, and no mention of prerequisites such as first resolving a message_id via mail_search or mail_read. The agent must infer the entire selection context from sibling names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_dateB
Get all emails from a specific date. Supports natural language like 'today', 'yesterday', 'November 13', 'last Friday'. Use this when the user asks for emails on a specific date.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | Date to retrieve emails (e.g., 'today', 'yesterday', 'Nov 13', '2025-01-15') | |
| include_junk | No | Include emails from Junk/Trash folders (excluded by default) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It notes natural-language date support, but that same capability is already spelled out in the schema's date description, so it adds little; nothing is said about return format, ordering, volume, or which mailboxes are scanned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, with the usage cue last. Slightly redundant since 'emails from a specific date' is effectively restated in the usage sentence, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only retrieval tool with no output schema and no annotations, the description is adequate for selection but silent on result structure, sorting, and pagination limits. It is minimally complete rather than thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including the natural-language examples repeated in the description. Baseline 3 is appropriate since the description adds no syntax or semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') plus resource ('emails') and the scoping condition ('from a specific date'), making the purpose unambiguous. It does not explicitly name the many sibling list/search tools (mail_search, mail_recent, mail_find) it overlaps with, so differentiation requires inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this when the user asks for emails on a specific date' gives an implied usage context, which is more than nothing. However it names no alternatives and gives no exclusion guidance against mail_search or mail_recent for overlapping date-related queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_draftA
Save an email to Drafts without sending it. Drafts are never delivered, so no recipient confirmation is required.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | CC email addresses | |
| to | Yes | Recipient email addresses (required, at least one) | |
| bcc | No | BCC email addresses | |
| body | No | Message body | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| subject | No | Subject line | |
| body_format | No | Plain text (default) or HTML. HTML is tag-stripped; the body is pasted into Mail's compose window (System Events focus on the message body, then Cmd-A and clipboard paste). AppleScript content/html content is not set because it quote-wraps the Sent body |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It usefully clarifies that drafts are never delivered and that recipient confirmation is unnecessary, but says nothing about where the draft is stored, whether it persists across sessions, or what permissions are needed for a mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the save-without-send behavior and no filler. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutating tool with no annotations and no output schema, the description is thin: it covers the send/no-send distinction but omits return behavior, draft location, and how confirm/dry_run interact with this operation. The rich schema compensates partially, leaving a clear gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters including enum and the confirm/dry_run semantics are documented in the schema itself. The description adds no parameter-level detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save an email to Drafts') plus the key contrast ('without sending it'), which cleanly separates it from the mail_send sibling without requiring either schema to be opened.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (compose-but-don't-send, no recipient confirmation needed) but the description never names mail_send as the alternative or states when a caller should draft versus send. The guidance is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_findA
Exact email lookup in Mail's own database (not the semantic index, so it never waits on indexing and matches filters exactly). Returns JSON: one row per copy (the same Message-ID in INBOX and Sent is two rows) with message_id, subject, from, to, received, mailbox, account_id. Needs at least one of subject_prefix, subject_contains, from, to, message_ids. Trash is excluded unless include_trash or mailboxes asks for it. Pair with mail_trash message_ids to clear what it finds.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Exact recipient address (To or Cc) | |
| from | No | Exact sender address | |
| limit | No | Maximum rows (default 50, max 500); 'truncated' says when more matched | |
| from_me | No | Only messages sent from one of your own addresses (any address that appears as a sender in a Sent mailbox) | |
| to_only | No | With to: that address is the only recipient | |
| mailboxes | No | Only these mailbox kinds (default: every mailbox except trash) | |
| message_ids | No | RFC822 Message-IDs (finds every copy) | |
| include_trash | No | Also return copies already in Trash (default false) | |
| received_after | No | ISO 8601 date/time; only messages received at or after it | |
| subject_prefix | No | Subject starts with this text (case-insensitive), e.g. '[Network] ' | |
| received_before | No | ISO 8601 date/time; only messages received before it | |
| older_than_hours | No | Only messages received more than this many hours ago | |
| subject_contains | No | Subject contains this text (case-insensitive) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that it never waits on indexing, that results are one row per copy with a named field list, and that Trash is excluded unless include_trash or mailboxes asks for it. It stops short of permission/safety context, but for a read tool that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the scope distinction, then packs return shape, constraints, defaults, and workflow into a few dense sentences with no filler. Slightly bracket-heavy, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter read tool with no output schema and no annotations, the description supplies the return format, default exclusions, and the minimum-filter requirement, which is most of what an agent needs. It omits any mention of pagination/truncation defaults beyond what the schema covers, keeping it just short of full.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a cross-parameter constraint the schema cannot express ('needs at least one of ...') and explains the dedup semantics of message_ids ('finds every copy'). That pushes it above the schema-only baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource (exact email lookup) and immediately differentiates itself from the semantic-index sibling by scope ('Mail's own database ... not the semantic index'). An agent can tell this apart from mail_search without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions ('Needs at least one of subject_prefix, subject_contains, from, to, message_ids') and a follow-up workflow ('Pair with mail_trash message_ids to clear what it finds'). It implies the semantic alternative but never names mail_search directly, so routing is left slightly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_forwardA
Forward an existing email to new recipients. An optional note is pasted above the forwarded original in a visible Mail window (same System Events prepend as mail_reply; AppleScript content is not set). More than one recipient requires confirm=true. A real send is verified in Sent/Outbox before success. A send hang is a timeout, not TCC, unless Mail reports -1743/-10004; hang recover matches intended To plus Fwd: subject or original Message-ID. Clients must Sent-check before retrying a timed-out send.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Forward recipients (required, at least one) | |
| body | No | Optional note pasted above the forwarded message (System Events prepend; omitted note skips paste) | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| file_path | No | Alternative to message_id: the .emlx file_path from mail_search, resolved to its Message-ID | |
| message_id | No | RFC822 Message-ID of the email (from mail_search / mail_read results) | |
| save_as_draft | No | Save the forward to Drafts instead of sending (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: the note is pasted above the original in a visible Mail window, AppleScript content is not set, sends are verified in Sent/Outbox before success, and it distinguishes timeout vs TCC failures including error codes -1743/-10004 and hang-recovery matching. This is genuinely rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by confirm and verification behavior. The trailing sentences on TCC codes and hang recovery are dense but each carries actionable retry guidance, so little is wasted even if the density is high.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no annotations and no output schema, the description covers send verification, timeout semantics, retry rules, and the optional note behavior. It is largely self-sufficient, with only minor gaps such as what the success/preview response actually contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter is already documented in the schema, and the description mostly restates these (note mechanism, confirm requirement). It adds little parameter-level meaning beyond what the structured fields provide, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Forward an existing email to new recipients') and implicitly separates itself from mail_reply by noting it shares the same System Events prepend mechanism. An agent can identify this as the forward action rather than reply or send.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating conditions: multi-recipient sends require confirm=true, a preview-only path exists, and clients must Sent-check before retrying a timed-out send. It stops short of an explicit 'use this instead of mail_reply/mail_send when X' routing statement, so it is clear context without full alternatives mapping.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_markC
Mark an email as read or unread.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | New read state (default read) | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| file_path | No | Alternative to message_id: the .emlx file_path from mail_search, resolved to its Message-ID | |
| message_id | No | RFC822 Message-ID of the email (from mail_search / mail_read results) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether the change is reversible, what permissions are needed, whether it affects only the local index or the server, or that the call can preview rather than apply. The confirm/dry_run preview semantics exist only in the schema, not in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with zero filler, appropriately sized for a narrowly scoped state-toggle tool. Nothing in it is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations and no output schema, the description is too thin. It never explains the target-identification choice between message_id and file_path, nor the preview/confirm behavior, both of which an agent needs to call it correctly without opening the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (status, confirm, dry_run, file_path, message_id) is documented in the schema itself, so the baseline of 3 applies. The description adds no meaning beyond the schema – it does not explain the status default or the message_id/file_path alternative.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Mark') and resource ('an email') plus the effect dimension ('as read or unread'), so the operation is unambiguous. It does not, however, distinguish itself from siblings like mail_archive or mail_trash, which also mutate message state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus mail_archive, mail_trash, or mail_read, and no mention of prerequisites or when marking is preferred over leaving a message unread. The single sentence is pure purpose, leaving selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_readA
Read full email content. Use the file_path from mail_search or mail_recent results.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | File path from mail_search results |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. 'Read' implies a non-mutating operation, so the safety profile is inferable, but the description says nothing about error behavior for invalid/stale paths, authentication, or whether attachments and HTML are included in 'full content.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the purpose is front-loaded and the prerequisite follows immediately. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool with no annotations or output schema, the description covers what it does and where the argument comes from. It could be slightly richer about what 'full content' includes (body, attachments, headers), but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is self-describing, so the baseline would be 3. The description adds real value by naming mail_recent as an additional valid source for file_path, while the schema mentions only mail_search.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read full email content,' which is clearly distinct from search/list operations. It names the tools that produce its input (mail_search, mail_recent), giving implicit differentiation, but does not contrast against the closely-related siblings mail_thread or messages_conversation that also surface message bodies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear chaining instruction — use the file_path from mail_search or mail_recent results — so the agent knows the prerequisite and the workflow position. There is no explicit 'do not use this for X' statement and no guidance on when to prefer mail_thread instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_recentA
Get most recent emails without semantic search. Use this when the user asks for 'recent emails', 'latest emails', 'what emails did I get', or 'unread emails'.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (default 30) | |
| days_back | No | Only emails from last N days (default 7) | |
| unread_only | No | Only show unread emails (queries Mail.app for read status) | |
| include_junk | No | Include emails from Junk/Trash folders (excluded by default) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not state that this is a safe read-only operation, the sort order beyond 'most recent', whether results are capped only by the limit, or anything about pagination. It is adequate for a simple list tool but adds no behavioral context beyond the purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the core purpose front-loaded before the routing triggers. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-required-param list tool with full schema coverage and no output schema, the description covers what an agent needs to select and invoke it. It only omits minor operational detail such as ordering guarantees, which is a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each of the four parameters (limit, days_back, unread_only, include_junk) documented including defaults and the Mail.app read-status nuance. The description adds nothing beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get most recent emails') and immediately draws the boundary against semantic-search siblings such as mail_search and smart_search. An agent can distinguish it from mail_search without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates the user utterances that should route here ('recent emails', 'latest emails', 'what emails did I get', 'unread emails') and names the alternative modality (semantic search) it is not. This is exactly the when-to-use signal an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_replyA
Reply to an existing email. The reply text is pasted above the quoted original in a visible Mail window (System Events focus on the message body, then Command-Up and clipboard paste — not Cmd-A, which would replace the quote; AppleScript content is not set because Mail 16 leaves it empty). reply_all=true fans out to every original recipient and requires confirm=true. A real send is verified in Sent/Outbox before success. A send hang is a timeout, not TCC, unless Mail reports -1743/-10004; hang recover uses In-Reply-To. Clients must Sent-check before retrying a timed-out send.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Reply text, prepended above the quoted original via System Events paste (required). AppleScript content is not set. | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| file_path | No | Alternative to message_id: the .emlx file_path from mail_search, resolved to its Message-ID | |
| reply_all | No | Reply to all original recipients (default false) | |
| message_id | No | RFC822 Message-ID of the email (from mail_search / mail_read results) | |
| save_as_draft | No | Save the reply to Drafts instead of sending (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and it does so richly: paste-above-quote mechanics, reply_all fan-out requiring confirm, real-send verification in Sent/Outbox, the timeout-vs-TCC distinction, error codes -1743/-10004, and hang recovery via In-Reply-To. This is unusually deep behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and every sentence carries behavioral content, though the parenthetical paste mechanics (System Events focus, Command-Up vs Cmd-A, Mail 16 detail) are deeply implementation-specific and verge on over-specification for a tool-selection context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must carry behavior and result semantics, and it does - explaining that success is verified in Sent/Outbox and how timeouts should be handled. Remaining detail (draft vs send distinctions) is left to the fully-covered schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by stating the coupling that reply_all=true requires confirm=true and that body is pasted rather than set via AppleScript, adding meaning the schema alone doesn't convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Reply to an existing email'), clearly distinct from mail_send, mail_draft, and mail_forward by the reply semantic. It does not explicitly name those siblings, so differentiation is inferred from the verb rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'reply to an existing email' and by the confirm/reply_all coupling, but the description never states when to pick this over mail_send or mail_draft, nor any precondition for a valid reply (e.g. must have a resolvable Message-ID). Adequate but with clear gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_searchB
Semantic search for emails using AI embeddings. Finds emails by meaning, not just keywords. Supports filtering by sender, recipient, attachments, mailbox, sent/received, and flagged.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (default 30) | |
| query | Yes | Natural language search (e.g., 'invoices', 'meeting notes', 'from John about project') | |
| sender | No | Filter by sender name or email address | |
| mailbox | No | Filter by mailbox name (e.g., 'INBOX', 'Archive', 'Sent Messages') | |
| sort_by | No | Sort by relevance (default) or date (newest first) | |
| days_back | No | Only emails from last N days (0 = all time) | |
| recipient | No | Filter by recipient name or email address | |
| sent_only | No | true = only sent emails, false = only received emails, omit for all | |
| flagged_only | No | Only show flagged/starred emails | |
| include_junk | No | Include emails from Junk/Trash folders (excluded by default) | |
| has_attachment | No | Filter to only emails with attachments (true) or without (false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a meaningful behavioral trait: ranking is semantic-embedding based rather than keyword based, which explains result relevance. It omits other behaviors such as default limits, pagination, return format, and any permission or rate constraints, so gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short front-loaded sentences with the core purpose stated first. Efficient overall, though the final sentence enumerating filters overlaps with structured schema fields and mildly duplicates information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, no-annotation, no-output-schema tool, the description covers purpose and filter scope but says nothing about what results look like, default/vs-max limits, or how relevance scores are returned. Adequate but leaves real gaps an agent would want filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema and the baseline is 3. The description's list of filter categories (sender, recipient, attachments, mailbox, sent/received, flagged) largely restates what the schema already provides and adds no syntax or semantic detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (emails) plus the distinguishing mechanism (semantic AI embeddings, 'by meaning, not just keywords'). This implicitly separates it from keyword-based siblings, but it never names the near-duplicate tools it competes with (mail_find, smart_search, messages_search), so an agent must still infer which one to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but gives no when-to-use or when-not-to-use guidance, and never references the alternative search tools in the sibling list. There is no stated condition that would route an agent here rather than to mail_find or smart_search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_sendA
Send an email through Mail.app. Success means the message was verified in Sent or is still in Outbox (still sending) — AppleScript send returning is not enough. Sending to more than one recipient in total (to + cc + bcc) requires confirm=true. dry_run and unconfirmed multi-recipient calls are not deliveries (MCP isError). A send hang is a timeout, not TCC, unless Mail reports -1743/-10004; hang recover matches To+subject (Message-ID when available), never subject alone. Clients must Sent-check before retrying a timed-out send.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | CC email addresses | |
| to | Yes | Recipient email addresses (required, at least one) | |
| bcc | No | BCC email addresses | |
| body | Yes | Message body (required) | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| subject | Yes | Subject line (required) | |
| body_format | No | Plain text (default) or HTML. HTML is tag-stripped; the body is pasted into Mail's compose window (System Events focus on the message body, then Cmd-A and clipboard paste). AppleScript content/html content is not set because it quote-wraps the Sent body |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses success verification criteria (Sent or Outbox, not merely AppleScript return), multi-recipient confirm requirements, dry_run and unconfirmed non-delivery semantics, timeout vs. TCC error codes (-1743/-10004), hang recovery matching rules, and the need for a Sent-check before retry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and contains no filler. It packs many distinct operational rules into a dense paragraph, which slightly reduces scannability, but every sentence adds necessary information for correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of sending mail, the absence of annotations, and the lack of an output schema, the description is complete for an agent to call the tool correctly. It covers success, failure, required flags, timeout recovery, and retry prerequisites, so little critical context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by specifying that multi-recipient means more than one recipient in total (to + cc + bcc) and that dry_run and unconfirmed multi-recipient calls are non-deliveries returned as MCP isError. These are useful semantic extensions for confirm and dry_run.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Send an email through Mail.app.' The operational constraints are immediate and unambiguous. It does not explicitly differentiate itself from sibling send tools such as messages_send, mail_draft, or mail_reply, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear rules for confirm=true on multi-recipient sends and describes dry_run and unconfirmed calls as non-deliveries. It also warns about retrying timed-out sends. However, it never says when to choose mail_send over sibling tools like mail_draft, messages_send, or mail_reply, so usage guidance for tool selection is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_sendersB
List most frequent email senders. Helps identify who you communicate with most.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum senders to return (default 30) | |
| days_back | No | Only count emails from last N days (0 = all time) | |
| include_junk | No | Include senders from Junk/Trash folders (excluded by default) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read/aggregation but doesn't disclose return shape (senders plus counts?), whether Junk is excluded by default, or any ordering semantics beyond 'most frequent'. Only the ranking behavior is hinted at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the primary action and followed by a single clarifying purpose line. No filler or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Parameters are fully covered by the schema, but with no output schema the description omits what the result looks like (e.g., sender + frequency, ordering). For a ranking tool this leaves a modest gap an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (limit, days_back, include_junk) are already fully documented in the schema. The description adds no parameter-level detail beyond what the schema provides, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List most frequent email senders', which is a distinct aggregation over the mail corpus rather than a message-level sibling like mail_search or mail_recent. The purpose is clear without needing the schema, though it does not explicitly call out how it differs from related mail tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Helps identify who you communicate with most' implies the use case (relationship/frequency analysis) but gives no explicit when-to-use versus alternatives and no exclusions or prerequisites. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_threadB
Get all emails in a conversation thread. Finds related emails by matching subject lines.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum emails to return (default 30) | |
| file_path | Yes | File path to any email in the thread |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that thread grouping works by subject-line matching, which can cause false groupings or omissions. Beyond that, it says nothing about ordering, permissions, or whether bodies/attachments are included, so it is only moderately transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the purpose and then the mechanism. It is efficient and wastes little space. A slight overstatement—'all emails' when a default limit of 30 exists—keeps it from a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter retrieval tool with no output schema, the description covers the core action and a key behavior. It still omits likely relevant context such as result ordering, failure modes when no thread is found, and the relationship to sibling search tools, so it is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both file_path and limit. The description adds no parameter-level detail, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Get all emails in a conversation thread.' It also adds a mechanism ('matching subject lines'), making the scope concrete. However, it does not differentiate from siblings like messages_conversation or mail_search, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no conditions for choosing this tool over alternatives, and no exclusions. The implied use case is retrievable from the name and first sentence, but the description does not help an agent route between this and similar thread/search tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_trashA
Move email to Trash. Destructive: requires confirm=true, otherwise the call only reports what would be trashed. One message: message_id or file_path. Many: message_ids (up to 500, e.g. from mail_find) moves every copy of each (INBOX, Sent, Archive...) to that account's Trash; each copy's Message-ID is re-checked before it moves. A bulk call stops starting new moves after ~35s and reports 'remaining'; run it again for the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| file_path | No | Alternative to message_id: the .emlx file_path from mail_search, resolved to its Message-ID | |
| message_id | No | RFC822 Message-ID of the email (from mail_search / mail_read results) | |
| message_ids | No | Bulk: RFC822 Message-IDs (from mail_find). Every copy outside Trash is moved. Do not combine with message_id / file_path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it flags the operation as destructive, explains the confirm/preview gate, discloses that EVERY copy across folders (INBOX, Sent, Archive) is moved, notes each Message-ID is re-checked before moving, and documents the ~35s bulk cap with a 'remaining' signal plus the re-run remedy. This is exactly the behavioral context an agent needs for a destructive bulk tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the destructive gate, then the mode semantics and limits. Dense and mostly waste-free, though the streaming clauses make it slightly run-on and some folder examples (INBOX, Sent, Archive) are illustrative rather than essential.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, and the description compensates by describing the preview report, the 'remaining' field on bulk timeouts, and the per-copy identity re-check. It does not address recovery/undo or permission requirements, but for a move-to-Trash tool with a preview gate the coverage is close to sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics beyond the schema: it groups message_id/file_path as 'one message' vs message_ids as 'many', states the 500-item bulk limit and the mail_find source, and notes copies span multiple folders. This meaningfully enriches the parameter contract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Move email to Trash') and immediately scopes the operation as destructive. An agent can distinguish this from siblings like mail_archive and mail_mark without inspecting schemas, and the single-vs-bulk modes are stated up front.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states the confirm=true prerequisite, that omitting it yields a preview, and when to use the bulk message_ids path vs a single message_id/file_path. It also points to mail_find as the source for IDs. It stops short of naming a sibling alternative (e.g. mail_archive) for non-destructive moves, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_contactsB
List all contacts you've messaged, sorted by most recent. Shows message count and last message date.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum contacts to return (default 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose the sort order (most recent) and the returned fields (message count, last message date). However, it says nothing about auth requirements, pagination beyond the limit, or behavior on an empty mailbox, so it only partially fills the gap left by missing annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the core purpose and append the return-value detail. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter read tool with no output schema, the description covers the operation, ordering, and returned fields adequately. Minor gaps around pagination behavior and prerequisites keep it just short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single limit parameter is fully documented in the schema, so no description compensation is needed. The description adds no extra meaning about the parameter, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and scoped resource (contacts you've messaged, sorted by most recent), which clearly distinguishes it from contacts_search, contacts_lookup, and messages_recent. It does not explicitly name any sibling, but the messaging-scoped contact list is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no alternatives named. The purpose implies a use case (finding people you've corresponded with), but an agent gets no routing help versus contacts_search or person_search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_conversationB
Get full conversation history with a specific contact. Shows messages in chronological order.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum messages to return (default 50) | |
| contact | Yes | Contact name or phone number |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden; it usefully discloses ordering ('chronological order') and completeness ('full'), which are real behavioral facts. However it omits pagination behavior, whether the limit truncates the 'full' history, permission/auth requirements, and what happens for an unknown contact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the scope (specific contact) leads and the ordering detail follows. Slightly generic but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema and no annotations, the description covers the core purpose but leaves the limit/truncation relationship and error behavior unaddressed, so an agent cannot fully predict the result set size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both contact and limit, making 3 the baseline. The description adds nothing about parameter meaning, and its claim of 'full conversation history' sits awkwardly against a default limit of 50 messages.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (full conversation history) scoped to a specific contact. It is clearly distinct from broadcast siblings like messages_recent or messages_search, but never names an alternative to route the agent explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'with a specific contact' implicitly signals this is the per-thread lookup versus a general search, but the description never states when to prefer it over messages_search or messages_recent, nor any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_recentA
Get most recent messages without semantic search. Use this when the user asks for 'recent messages', 'latest texts', or 'what messages did I get'.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (default 30) | |
| days_back | No | Only messages from last N days (default 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It correctly frames this as a plain recency retrieval with no semantic ranking, but it does not disclose result ordering guarantees beyond "most recent", pagination behavior, or whether defaults are applied server-side. Adequate but with clear gaps for a zero-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core purpose front-loaded and the routing cues immediately after. There is no filler and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-output-schema read tool with fully documented parameters, the description covers what an agent needs in order to call it correctly. It could be slightly more complete by pointing to the sibling to use when filtering by contact or thread is needed, which is the only real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the two parameters (limit, days_back) and their defaults are already fully documented in the schema. The description adds the recency framing but no additional syntax or format detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ("Get most recent messages") and explicitly scopes it against the semantic-search alternative ("without semantic search"), which distinguishes it from siblings like messages_search and smart_search. An agent can tell what this does and how it differs from the search tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete trigger conditions with quoted user phrasings ("recent messages", "latest texts", "what messages did I get"), which is strong when-to-use guidance. It stops short of naming the sibling to use instead for filtered/lookup cases, so it is clear context but not a full use-this-not-that rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_searchA
Semantic search for iMessages/SMS using AI embeddings. Finds messages by meaning. Supports filtering by contact, group chats, specific group name, and attachments.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (default 30) | |
| query | Yes | Natural language search (e.g., 'dinner plans', 'about the trip', 'address') | |
| contact | No | Filter by contact name or phone number | |
| sort_by | No | Sort by relevance (default) or date (newest first) | |
| days_back | No | Only messages from last N days (0 = all time) | |
| has_attachment | No | Filter to messages with attachments (photos, files) | |
| group_chat_name | No | Filter by specific group chat name | |
| group_chat_only | No | Only show messages from group chats |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the key behavioral trait that this is embedding-based semantic matching rather than literal keyword matching, but it omits return format, pagination/precedence, and any auth or scope constraints for an 8-parameter read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose and followed by scope. 'Finds messages by meaning' is mildly redundant with 'semantic search,' but otherwise the text is tight and earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema and no annotations, the description covers purpose and filter scope adequately, and the schema supplies full parameter detail. It would be stronger if it hinted at result shape or how many results to expect, but nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters are already documented in the schema. The description only restates the filter categories (contact, group chats, group name, attachments) without adding syntax, defaults, or interaction rules beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and method ('Semantic search for iMessages/SMS using AI embeddings') plus the resource it operates on, which distinguishes meaning-based search from keyword search. It does not explicitly name siblings like messages_recent or messages_conversation, so an agent must infer the boundary, but the resource and mechanism are otherwise unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Filtering capability ('by contact, group chats, specific group name, and attachments') implies when the tool is useful, but there is no explicit when-to-use versus messages_recent, messages_conversation, or smart_search, and no exclusions or prerequisites. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_sendA
Send an iMessage or SMS. Address it with 'to' (E.164 phone numbers like +15551234567, or Apple ID emails) or with 'chat_id' (a chat GUID from chat.db, e.g. 'iMessage;-;+15551234567' or 'iMessage;+;chat123456'). Group chats and multiple recipients require confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Phone numbers in E.164 form or Apple ID emails | |
| text | No | Message text (required unless attachment_path is given) | |
| chat_id | No | Existing chat GUID; verified against the Messages database before sending | |
| confirm | No | Required for deletes and multi-recipient sends; without it the call only previews | |
| dry_run | No | Preview only: report what would happen and change nothing (default false) | |
| service | No | Delivery service (default auto: iMessage, then SMS relay) | |
| attachment_path | No | Absolute path to an existing file on this Mac to attach |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the confirm gating (preview without it), that chat_id is verified against the Messages database, and that multi-recipient sends need confirmation — but it omits permissions/account requirements, irreversibility of sending, and any return behavior. Partial coverage for an irreversible mutating action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then the addressing options, then the confirm caveat. No filler, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description plus fully-documented schema cover addressing, confirmation, service selection and dry-run adequately for a send tool. Missing only peripheral details such as required permissions and confirmation of what a successful send returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all seven parameters including the E.164 and chat GUID formats. The description reinforces the addressing formats with examples but adds little beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (send an iMessage or SMS) and names the delivery medium, which cleanly separates it from sibling senders like mail_send. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear routing guidance for the two addressing modes ('to' vs 'chat_id') and states the condition that requires confirm=true for group/multi-recipient sends. It does not name alternatives or say when not to use it, but the context is sufficient to invoke correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
person_searchA
Search ALL communication with a specific person across Mail, Messages, and Calendar. Automatically finds their emails and phone numbers from Contacts to search all sources.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Person's name to search for (will resolve to all their email addresses and phone numbers) | |
| limit | No | Maximum results per source (default 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose real behavior: it auto-resolves a name to all emails and phone numbers via Contacts and fans out across three sources. It omits other behavioral facts such as result ordering, deduplication across sources, permission needs, or pagination behavior beyond the limit param.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler, and the defining scope ('ALL communication ... across Mail, Messages, and Calendar') is front-loaded in the first clause. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param search tool with no output schema and no annotations, the description covers what it does and the non-obvious contact-resolution mechanism. It is slightly thin on return-value shape and result ordering, but nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema and the baseline is 3. The description reinforces the name-to-contacts resolution semantics but adds no syntax, format, or matching-rule detail beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (ALL communication with a specific person) plus the scope (Mail, Messages, and Calendar). By declaring cross-source, person-centric aggregation it implicitly and cleanly distinguishes itself from the single-source siblings like mail_search, messages_search, and calendar_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The cross-source scope implies when this is preferable to the per-source search tools, and the mention of automatic contact resolution hints at the setup it handles for you. However, it never explicitly names an alternative (e.g. smart_search or contacts_search) or states when NOT to use it, so guidance stays implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rebuild_indexA
Rebuild the search index for one or more data sources. This clears the existing index and re-indexes all content from scratch. Use this if search results are stale, missing, or if the index is corrupted. Can rebuild emails, messages, calendar, or all sources at once.
| Name | Required | Description | Default |
|---|---|---|---|
| sources | No | Which sources to rebuild. Defaults to all sources if not specified. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the important destructive trait: it "clears the existing index and re-indexes all content from scratch." It does not cover permissions, expected duration, or whether search is unavailable during the rebuild, so it falls short of full behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and destructive effect are front-loaded, and the sentences are mostly free of waste. The final sentence recapping the sources is largely redundant with the schema enum, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description covers purpose, trigger conditions, and the destructive nature of the operation. It leaves open operational details such as required permissions or availability during the rebuild, but nothing that would prevent a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the enum values and default are already documented in the schema. The description repeats the accepted sources ("emails, messages, calendar, or all sources at once") but adds no syntax or behavioral detail beyond the schema, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ("Rebuild the search index") and scopes it to "one or more data sources." It is clear what the tool does, but it does not differentiate itself from the closest sibling, audit_index, which an agent might confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states an explicit trigger condition: "Use this if search results are stale, missing, or if the index is corrupted." That is clear context for when to reach for it, though it names no alternative tool or when-not-to-use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
smart_searchA
Intelligent search across Mail, Messages, and Calendar. Automatically determines which sources to search based on your query. Returns results grouped by time when multiple sources match. Use this for complex queries that might span multiple data sources.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results per source (default 5) | |
| query | Yes | Natural language search query (e.g., 'meeting with John', 'budget discussion', 'what happened yesterday') | |
| synthesize | No | Group results by time proximity (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that source selection is automatic and that results are grouped by time when multiple sources match, which is real behavioral context. However it says nothing about permissions, how sources are chosen, ordering, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with purpose, then mechanism, then routing guidance. No filler or repetition, though "Intelligent" is slightly marketing-flavored rather than informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers what comes back (results grouped by time across matched sources) and when to reach for it. Adequate for a read-only search tool; missing only edge behavior such as empty results or source-selection failure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents query, limit and synthesize with examples, giving a baseline of 3. The description's time-grouping note loosely reinforces synthesize but adds no syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource set (Mail, Messages, Calendar) with the notable differentiator that it auto-selects sources. Distinguishes itself from the many single-source siblings like mail_search/messages_search/calendar_search by scope, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this for complex queries that might span multiple data sources" gives an explicit when-to-use condition that implicitly excludes simple single-source lookups. It stops short of naming the alternative single-source tools for the non-matching case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
39 tool updates
v3.0.8- First observed
audit_index - First observed
calendar_add - First observed
calendar_date - First observed
calendar_edit - First observed
calendar_free_time - First observed
calendar_list_calendars - First observed
calendar_recurring - First observed
calendar_remove - First observed
calendar_rsvp - First observed
calendar_search - First observed
calendar_upcoming - First observed
calendar_week - First observed
contacts_add - First observed
contacts_edit - First observed
contacts_lookup - First observed
contacts_remove - First observed
contacts_search - First observed
mail_archive - First observed
mail_date - First observed
mail_draft - First observed
mail_find - First observed
mail_forward - First observed
mail_mark - First observed
mail_read - First observed
mail_recent - First observed
mail_reply - First observed
mail_search - First observed
mail_send - First observed
mail_senders - First observed
mail_thread - First observed
mail_trash - First observed
messages_contacts - First observed
messages_conversation - First observed
messages_recent - First observed
messages_search - First observed
messages_send - First observed
person_search - First observed
rebuild_index - First observed
smart_search
TDQS
Scored across 39 tools
Most tools have clearly distinct purposes, but some pairs overlap: contacts_search vs contacts_lookup, calendar_upcoming vs calendar_week, and the numerous mail search variants (mail_search, mail_find, smart_search, mail_recent, mail_date). Descriptions help, but an agent could still misselect among the semantically similar options.
All tools use snake_case with clear domain prefixes (calendar_, mail_, messages_, contacts_) and a consistent naming style. The few cross-cutting tools (person_search, smart_search, rebuild_index, audit_index) are readable and fit the overall pattern.
39 tools is excessive for the scope; many read/search operations overlap across Mail and Calendar. While the multi-app breadth justifies some count, significant consolidation could reduce the surface without losing functionality.
Core CRUD and search operations are covered for Calendar, Mail, Messages, and Contacts, plus index management tools. Minor gaps exist (e.g., message deletion, mail flagging, contact groups), but the surface is largely complete for common workflows.
Maintenance
Related MCP Connectors
Mac & Windows: let ChatGPT, Claude & Cursor use your email, calendar, iMessage, Teams, files. Free.
One semantic search across your sites, Drive, Notion, email, files and Basecamp.
Personal memory layer: files, notes, messages, tasks and email, semantically indexed.
Your own AI reads, searches and drafts in your mailbox, on your Windows computer.
Related MCP Servers
- AlicenseAqualityAmaintenancePrivacy-first local document search using semantic search. Runs entirely on your machine with no cloud services, supporting PDF, DOCX, TXT, and Markdown files.97,180 npm405MIT
- FlicenseNot gradedqualityNot gradedmaintenanceEnables AI assistants like Claude to search and reference your Apple Notes using semantic search and RAG capabilities, with fully local execution and no API keys required.15,901 npm-
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to search and retrieve memories from your Mac, including screen captures, meeting transcripts, and browsing history, all locally and privately.1-
- AlicenseNot gradedqualityCmaintenanceEnables semantic search, connection discovery, and grounded synthesis across your Apple Notes using on-device embeddings and optional LLM synthesis, with direct SQLite access for fast indexing.13 npm8MIT