ofw-mcp
This server connects an AI assistant to OurFamilyWizard so it can read, sync, compose, and manage co-parenting messages, drafts, attachments, calendar events, expenses, and journal entries.
Check credential health and upstream reachability.
Read parent/co-parent profiles and dashboard notifications.
List message folders and cached messages with filters, sorting, date ranges, and pagination.
Read individual messages or drafts, including attachment metadata.
Send messages or drafts with confirmation safeguards, reply threading, and attachment support.
Create, list, replace, and delete drafts with revision conflict protection.
List sent messages that recipients have not read.
Upload local files as private or shared OFW attachments.
Download attachments inline or to disk, with content extraction for PDFs, spreadsheets, documents, and text.
Sync messages into a local cache, including resumable/deep syncs.
Check whether cached data matches OFW live state.
Get a live status snapshot of drafts and message lifecycle states.
List, create, update, and delete calendar events with co-parent visibility confirmations.
Read expense totals and list expenses with paging.
Create expenses, with confirmation because they appear in the shared ledger.
List and create journal entries.
Connects to Apple's web services, including Apple Music, iCloud Calendar/Contacts/Mail, Apple Maps, WeatherKit, and the iTunes Search API, enabling access to music, calendar, contacts, mail, maps, weather, and iTunes catalog data.
Provides tools for Apple Music catalog search, charts, library and playlist management (create/add tracks/folders), listening history, recommendations, Replay, ratings, and favorites, with optional web-player mode for playlist editing, track reordering, removal, and deduplication.
Provides access to iCloud Calendar, Contacts, and Mail over CalDAV, CardDAV, and IMAP/SMTP, including listing/searching/managing calendar events, searching/creating/updating/deleting contacts and groups, and searching/reading/sending/replying to/flagging/moving mail.
Provides tools for the iTunes Search API and top charts, enabling search and lookup of songs, albums, podcasts with episodes, apps, books, and audiobooks without credentials.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ofw-mcpWhat's on the kids' calendar next week?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
apple-icloud-mcp
A Model Context Protocol server that connects Claude to Apple's web services: Apple Music (your playlists, library and the catalog), iCloud Calendar, Contacts and Mail, Apple Maps, WeatherKit, and the iTunes Search API.
It talks to Apple over the network through Apple's own web APIs and standard protocols (CalDAV, CardDAV,
IMAP/SMTP), so — unlike Mac-only Apple integrations such as
apple-swift-mcp — it runs anywhere: Linux, Windows, a
container, or hosted on mcp-host as a claude.ai connector. The two complement each other:
apple-swift-mcp drives the Mac's own apps (Calendar, Reminders, Contacts, Mail, Messages, Notes, Photos, Maps),
while apple-icloud-mcp reaches Apple's web services from anywhere and adds Apple Music, WeatherKit and iTunes.
Unofficial. This project is not affiliated with, endorsed by or sponsored by Apple Inc. Apple, iCloud, Apple Music, Apple Maps and WeatherKit are trademarks of Apple Inc., used here only to name the services it connects to.
AI-developed project. This codebase was entirely built and is actively maintained by Claude Code. No human has audited the implementation. Review all code and tool permissions before use.
Not yet verified against a live Apple account. Every request shape comes from Apple's documentation or
Apple's own web client and is covered by tests against faithful fakes, but the build environment had no
Apple credentials. The first real run is the real verification — run apple_healthcheck first, and see
docs/APPLE-API.md for what each call relies on.
What you can do
"Make a playlist called Road Trip 2026 with twenty upbeat 80s songs"
"Add the top 5 songs from Taylor Swift's latest album to my Workout playlist"
"Remove duplicates from my Chill playlist and sort it by artist" (web-player mode)
"What have I been listening to on heavy rotation?" / "Show my Apple Music Replay for this year"
"What's on my calendar next week?" / "Find me a free hour on Thursday afternoon"
"Move my dentist appointment to Friday at 3pm"
"What's Jane's phone number?" / "Add Bob Lee from Acme with bob@acme.com to my contacts"
"Any unread email from my landlord?" / "Reply to the last message from Sam saying I'll be there"
"How long will it take to drive to the airport at 5pm?" / "Coffee shops near Union Square"
"Will it rain in Chicago tomorrow?"
"Find the podcast Hard Fork and list its latest episodes" / "What are the top albums in the UK right now?"
Related MCP server: OurFamilyWizard MCP
Services and what each needs
Every service is optional: each one switches on when its credentials are present, and
apple_healthcheck tells you which are working.
Service | What you get | Credential | Cost |
Apple Music (official API) | Catalog search, charts, your library, playlists (create, add tracks, folders), history, recommendations, Replay, ratings, favorites | Apple Developer key + a Music User Token | Apple Developer Program ($99/yr) + Apple Music subscription |
Apple Music (web-player mode, opt-in, unofficial) | Everything above plus rename/delete playlists, remove/reorder/sort/dedupe tracks, move to folders, remove from library | The | Apple Music subscription only |
iCloud Calendar (CalDAV) | List/search events (recurring ones expanded), create/update/delete with correct time zones and recurrence, free-time finder | Apple ID + app-specific password | Free |
iCloud Contacts (CardDAV) | Search, read, create, update (per-email/phone edits), delete, groups | Apple ID + app-specific password | Free |
iCloud Mail (IMAP/SMTP) | List mailboxes, search, read (without marking read), send and reply, flag, move | Apple ID + app-specific password | Free |
Apple Maps (Maps Server API) | Geocode, reverse geocode, place search, directions, ETAs, signed map-image URLs | Apple Developer key | Free up to 25,000 calls/day |
WeatherKit | Current conditions, hourly/daily forecast, next-hour precipitation, severe-weather alerts | Apple Developer key + Services ID | Free up to 500,000 calls/month |
iTunes Search + charts | Search/lookup songs, albums, podcasts (with episodes), apps, books, audiobooks; top charts | None | Free |
Requirements
Node.js 22 or later (the hosted runner uses Node 26).
Whatever the services you want need (table above).
Acknowledgement of Terms
By using this MCP server, you acknowledge and agree to the following:
1. It accesses your own Apple accounts only, with credentials you supply. It cannot reach anyone else's data, and on a shared deployment each person's Apple ID stays in their own process.
2. Apple's terms govern your use, just as they govern your direct use of Apple's services — in particular the Apple Media Services Terms, the iCloud Terms and Conditions and, for the developer-key services, the Apple Developer Program License Agreement. The iCloud Terms say you may not "interfere with or disrupt the Service (including accessing the Service through any automated means, like scripts or web crawlers)". The iCloud paths here use the mechanism Apple provides for third-party apps — app-specific passwords over CalDAV, CardDAV and IMAP/SMTP — at the pace of a person using an assistant. Apple Music web-player mode is different: it uses Apple's private web-player API with your browser session, which Apple has said may be blocked at any time. It is off unless you turn it on. Terms read by the maintainer 2026-09-26.
3. Personal use only. This project is not affiliated with, endorsed by, or sponsored by Apple Inc. "Apple", "Apple Music", "iCloud" and related marks are Apple's. Do not share your credentials, and do not use this to bulk-extract data.
4. Writes are real. Creating a calendar event with attendees makes iCloud email them invitations; sending
mail sends it; deleting a playlist or contact deletes it. Those actions ask for confirmation first (see
Confirmation), and APPLE_WRITE_MODE
can remove write tools entirely.
5. You accept full responsibility for the consequences, technical (rate limits, a locked account after repeated failed sign-ins) and otherwise.
This is the maintainer's good-faith summary, not legal advice, and it does not modify or supersede Apple's actual terms.
Installation
Claude Code
claude mcp add apple -- npx -y apple-icloud-mcpthen set the variables for the services you want (see Setting up credentials) in
the server's env, for example in .mcp.json:
{
"mcpServers": {
"apple": {
"command": "npx",
"args": ["-y", "apple-icloud-mcp"],
"env": {
"ICLOUD_USERNAME": "you@icloud.com",
"ICLOUD_APP_PASSWORD": "abcd-efgh-ijkl-mnop",
"DISPLAY_TZ": "America/New_York"
}
}
}
}Claude Desktop
Install the .mcpb bundle from the latest release and fill
in the settings it asks for, or add the same npx entry to claude_desktop_config.json.
From source
git clone https://github.com/chrischall/apple-icloud-mcp.git && cd apple-icloud-mcp
npm install && npm run build
cp .env.example .env # fill in what you need
npm run devSetting up credentials
iCloud Calendar, Contacts and Mail — an app-specific password
Sign in at appleid.apple.com → Sign-In and Security → App-Specific Passwords → generate one (two-factor authentication must be on).
Set
ICLOUD_USERNAME(your Apple ID email) andICLOUD_APP_PASSWORD(thexxxx-xxxx-xxxx-xxxxpassword). Your normal Apple ID password does not work here.If your Apple ID is not an
@icloud.com/@me.comaddress, also setICLOUD_MAIL_ADDRESSto your iCloud Mail address (Mail only).
Changing your Apple ID password revokes every app-specific password — the usual reason a working setup suddenly returns "credentials rejected". After a definitive rejection the server stops sending that password — for 24 hours or until you change it, across restarts too (it records a digest of the rejected pair, never the password) — because repeated failed sign-ins can lock an Apple ID.
Reminders are not available: since iOS 13, iCloud Reminders no longer sync over CalDAV (see Not supported).
Apple Developer key — Apple Music (official), Apple Maps, WeatherKit
These three need a paid Apple Developer Program membership.
In Certificates, Identifiers & Profiles, create the identifiers for the services you want: a Media ID (Apple Music), a Maps ID (Apple Maps), and a Services ID for WeatherKit.
Under Keys, create a key and tick Media Services (MusicKit), MapKit JS and/or WeatherKit. One key can carry all three. Download the
.p8(you only get one chance).Set
APPLE_TEAM_ID,APPLE_KEY_IDandAPPLE_PRIVATE_KEY(the.p8contents — multi-line PEM, a one-line value with\nescapes, or base64 all work). For WeatherKit also setAPPLE_WEATHERKIT_SERVICE_ID. Separate keys per service are supported withAPPLE_MUSIC_KEY_ID/APPLE_MUSIC_PRIVATE_KEY,APPLE_MAPS_…andAPPLE_WEATHERKIT_….
The server signs short-lived tokens itself (12 hours for Apple Music, 1 hour for Maps and WeatherKit).
Apple Music — your library (official API): a Music User Token
Catalog search works with the developer key alone. Your library and playlists also need a Music User Token, which Apple only issues through an interactive MusicKit sign-in in a browser. With the developer key set, run:
npx apple-icloud-mcp music-authIt opens a page on 127.0.0.1, you click Sign in with Apple Music, and it prints
APPLE_MUSIC_USER_TOKEN=…. Put that value in APPLE_MUSIC_USER_TOKEN. It lasts about six months, is tied to
the developer key that minted it, and changing your Apple ID password revokes it.
Someone without the developer key (another person on a shared deployment) never needs the .p8. The key's
owner runs npx apple-icloud-mcp music-auth --print-developer-token --days 7 and sends them the short-lived
token it prints; they run APPLE_MUSIC_DEVELOPER_TOKEN=<that token> npx apple-icloud-mcp music-auth on their
own machine, sign in, and keep the user token it prints. The developer token expires on its own; the user token
keeps working with the server's key.
Apple Music — web-player mode (opt-in, unofficial)
Apple's official API cannot rename or delete playlists, remove or reorder tracks, or remove anything from your library. Apple's own web player can. Web-player mode uses that same private API with your browser session:
Sign in at music.apple.com in a desktop browser.
Open the developer tools → Application/Storage → Cookies →
https://music.apple.comand copy the value of themedia-user-tokencookie.Set it as
APPLE_MUSIC_WEB_USER_TOKEN.
The web player's own developer token is read from music.apple.com automatically (and cached). This mode also
works with no Apple Developer account. It is unsupported by Apple and may stop working without notice; the
tools that need it say so in their descriptions, and their answers report which backend (official or web)
served them.
Running on mcp-host
The package ships a mint.yaml that mcp-host reads when you register apple-icloud-mcp:
Owner settings (the developer key,
APPLE_WRITE_MODE,APPLE_SERVICES,MCP_CONFIRM_SECRET, …) go in the registration's environment — store the private key and confirm secret as secrets.Everything personal (Apple ID, app-specific password, Music tokens, iCloud Mail address, time zone, default calendar, Apple Music storefront) is declared as
auth.fields, so mcp-host asks each person when they connect and remembers the answers per person; each caller gets their own process and data directory.State:
dataDirholds a few small files (see Security) that keep a scale-to-zero cold start cheap and stop a revoked password or a spent confirmation token from being reused after a restart.Egress: only the Apple hosts this server calls (
api.music.apple.com,amp-api.music.apple.com,music.apple.com,caldav.icloud.com,contacts.icloud.com,*.icloud.com,imap.mail.me.com,smtp.mail.me.com,maps-api.apple.com,weatherkit.apple.com,itunes.apple.com,rss.marketingtools.apple.com). HTTPS goes through the runner's proxy via Node's built-in fetch; IMAP and SMTP are tunnelled through the same proxy with HTTP CONNECT.
Give it your time zone (DISPLAY_TZ): the runner is on UTC, and times you give without an offset ("3pm") are
read in DISPLAY_TZ. Set MCP_CONFIRM_SECRET too: without it, a confirmation token issued just before the
child restarts (a redeploy, a machine move) stops working and you have to preview again. Spent tokens are
recorded on disk, so a shared secret does not let one be replayed.
Tools
58 tools. Mode is the lowest APPLE_WRITE_MODE that registers the tool; Confirm marks tools that ask first (details); 🌐 marks Apple Music tools that need web-player mode.
Health
Tool | What it does | Mode | Confirm |
| Check which Apple services (Apple Music, iCloud Calendar, Contacts, Mail, Apple Maps, WeatherKit, iTunes) are configured and reachable. | read |
Apple Music
Tool | What it does | Mode | Confirm |
| Search Apple Music's catalog by text for songs, albums, artists, playlists, music videos or stations. | read | |
| Look up Apple Music catalog songs, albums, artists, playlists, music videos or stations by id — or songs/music videos by ISRC, albums by UPC. | read | |
| Get Apple Music's top charts (most played songs, albums, playlists, music videos) for a storefront, optionally for one genre. | read | |
| List the playlists in your Apple Music library (alphabetical), or the contents of one playlist folder (folders and playlists). | read | |
| Read one playlist — a library playlist (p.…) or a catalog playlist (pl.…) — with its tracks in order. | read | |
| List the playlist folders in your Apple Music library: the top level, or the sub-folders of one folder. | read | |
| Search YOUR Apple Music library (not the whole catalog) by text for songs, albums, artists, playlists or music videos. | read | |
| List what is in your Apple Music library: all songs, albums, artists or music videos (alphabetical, paged), or recently-added items. | read | |
| Your recent Apple Music listening: recently-played (albums, playlists, stations), recently-played-tracks (songs), recent-stations, or heavy-rotation. | read | |
| Your personal Apple Music recommendations ("Made for You", "Recently Played" and similar groups), each with its title and the albums, playlists or stations in it (catalog ids). | read | |
| Apple Music Replay: your top songs, albums and artists for the latest Replay year, or for a given year (with play counts where Apple provides them). | read | |
| Whether you have loved or disliked songs, albums, playlists, music videos or stations — catalog or library ids — returning love, dislike or none per id. | read | |
| Create a new playlist in your Apple Music library, optionally with tracks (up to 500 catalog or library song ids; added 100 at a time), a description, a folder and public visibility. | additive | |
| Append songs (up to 500 catalog or library ids) to the end of one of your library playlists. | additive | |
| Create a playlist folder in your Apple Music library, at the top level or inside another folder. | additive | |
| Add catalog songs, albums, playlists or music videos to your Apple Music library by catalog id (up to 100 per type). | additive | |
| Mark catalog songs, albums, playlists, artists or music videos as favorites (the star in Apple Music; favorite songs go to your Favorite Songs playlist), by catalog id, up to 100 per type. | additive | |
| Set your rating on a song, album, playlist, music video or station (catalog or library id): love, dislike, or none to clear it. | all | |
| Rename one of your Apple Music library playlists, change its description, or make it public/private. | all | |
| Remove tracks from one of your library playlists by library track id and/or 1-based position (from apple_music_get_playlist). | all | yes |
| Reorder one of your library playlists: move tracks, sort (name, artist, album, release date, duration, date added), reverse, dedupe (keep the first copy of each song), or replace with a complete new order of its track ids (can drop tracks). | all | yes |
| Move one of your library playlists into a playlist folder, or back to the top level ("root"). | all | |
| Delete one of your library playlists (songs stay in your library). | all | yes |
| Remove songs, albums, music videos or playlists from your Apple Music library by LIBRARY id (i.…, l.…, p.… from apple_music_list_library / apple_music_search_library), up to 50 at a time; the preview names each item. | all | yes |
| Remove the favorite (star) from catalog songs, albums, playlists, artists or music videos, by catalog id, up to 100 per type. | all |
iCloud Calendar
Tool | What it does | Mode | Confirm |
| List your iCloud calendars (event calendars only): id, name, color, whether you can add events to it, whether it is shared (shared: with you by someone else; sharedByYou: by you with others), and which one new events go into by default. | read | |
| List iCloud Calendar events (appointments, meetings) in a date window, recurring events expanded into occurrences, sorted by start. | read | |
| Search iCloud Calendar events by text (case-insensitive match in title, location or notes) within a date window: fromDate (default today; may be in the past) + toDate or daysAhead (default 30), max 366 days. | read | |
| Get one iCloud Calendar event in full (notes untruncated, attendees, alerts, recurrence rule in plain English) by the id list/search returned. | read | |
| Create an iCloud Calendar event: title, startDate/endDate (timed default 1 hour; all-day endDate = last day), location, notes, url, alarms, recurrence, attendees. | additive | with attendees |
| Change an iCloud Calendar event: title, startDate/endDate, isAllDay, location, notes, url ("" clears), alarms, attendees (full new list), calendar (moves it). | all | with attendees |
| Delete an iCloud Calendar event. Recurring: span thisEvent (default; the one occurrence an "#occ=" id names), futureEvents (it and all later ones) or allEvents (the whole series). | all | yes |
| Find free time in your iCloud calendars: open slots per day within working hours (workdayStart/workdayEnd, default 09:00–17:00, weekdays only by default) at least minDurationMinutes long (default 30). | read |
iCloud Contacts
Tool | What it does | Mode | Confirm |
| Search the user's iCloud Contacts (address book) by name, nickname, company, job title, email, or phone number digits — optionally only within one contact group. | read | |
| Get one iCloud contact in full by id (from apple_contacts_search): names, organization, job title, emails, phones, postal addresses and URLs — each with its label and an entryId that apple_contacts_update can target — birthday, note, the… | read | |
| List the contact groups in iCloud Contacts (e.g. Family, Work): each group's id, name and member count. | read | |
| Create a new contact in iCloud Contacts. Needs a givenName, familyName or organization; optional middleName, nickname, department, jobTitle, note, birthday (YYYY-MM-DD, or --MM-DD without a year), and lists of emails, phones, urls and po… | additive | |
| Edit an existing iCloud contact in place. Scalar fields (givenName, familyName, middleName, nickname, organization, department, jobTitle, note, birthday) replace the current value; "" clears it. | all | |
| Permanently delete one contact from iCloud Contacts (on every device) by id. | all | yes |
iCloud Mail
Tool | What it does | Mode | Confirm |
| List the iCloud Mail mailboxes (folders): path, name, special use (inbox, sent, drafts, trash, junk, archive) and, by default, message and unread counts. | read | |
| Search emails in one iCloud Mail mailbox (default INBOX) by sender, recipient, subject, full text, received date range, unread and flagged state. | read | |
| Read one iCloud Mail message by uid (from apple_mail_search): headers (from, to, cc, reply-to, date, subject, message-id), the body as plain text (HTML converted to readable text when there is no text part), a truncated flag, and attachm… | read | |
| Send a plain-text email from your iCloud Mail address (to/cc/bcc, subject, body; no attachments). | all | yes |
| Mark iCloud Mail messages read or unread, and flag or unflag them, by uid (1–100 uids from apple_mail_search, one mailbox). | all | |
| Move iCloud Mail messages (1–100 uids from apple_mail_search, one mailbox) to another mailbox: a path from apple_mail_list_mailboxes or an alias (inbox, archive, trash, junk, sent, drafts). | all |
Apple Maps
Tool | What it does | Mode | Confirm |
| Turn an address or place name into coordinates with Apple Maps (geocoding). | read | |
| Find the street address at a latitude/longitude with Apple Maps (reverse geocoding) — e.g. "where is 37.33,-122.01?". | read | |
| Search Apple Maps for places: businesses, points of interest, addresses, landmarks (e.g. "coffee", "EV charger", "Golden Gate Bridge"). | read | |
| Driving, walking or cycling directions between two places with Apple Maps (addresses or "lat,lng"). | read | |
| Travel time and distance from one point to up to 10 destinations at once with Apple Maps — driving with live traffic, transit, walking or cycling (e.g. "which of these stores is closest by car?"). | read | |
| Look up Apple Maps places by place id — the id field from apple_maps_search, apple_maps_geocode or apple_maps_reverse_geocode results — 1 to 50 at once. | read | |
| Make a signed link to a static Apple Maps image (PNG): centred on an address or "lat,lng", and/or with pins (each with an optional label, colour and one-character glyph; the map fits the pins when no center is given). | read |
WeatherKit
Tool | What it does | Mode | Confirm |
| Weather forecast for a place from Apple Weather (WeatherKit): current conditions, hourly (up to 240 h), daily (up to 10 days), next-hour rain, severe-weather alerts (need countryCode). | read | |
| Get one severe-weather alert's full official text from Apple Weather (WeatherKit), unmodified, by its id (the alerts[].id from apple_weather_get called with countryCode). | read |
iTunes Search and charts
Tool | What it does | Mode | Confirm |
| Search Apple's iTunes Store catalog — songs, albums, artists, podcasts and podcast episodes, audiobooks, apps and ebooks — with no Apple account or key. | read | |
| Look up iTunes Store items (no Apple account or key) by ids — 1–200 trackId/collectionId/artistId values, e.g. from apple_itunes_search or an Apple Music link — or by one UPC/EAN (album), ISBN (book) or bundleId (app). | read | |
| Apple's current top charts (no Apple account or key): most-played songs, albums, music videos and playlists on Apple Music; top podcasts, trending podcast episodes and top subscriber channels; top free/paid apps and books; top audiobooks. | read |
Write protection (APPLE_WRITE_MODE)
Value | What is registered |
| Read tools only. |
| Reads, plus writes that only add to your own account: create a playlist or folder, append tracks, add to library/favorites, create an event (without attendees, and not in a calendar shared with other people) or a contact. Nothing existing is modified or removed and nothing is sent to anyone. |
| Everything. |
Gated tools are not registered at all below their mode, so no prompt or injected instruction can call them.
An unrecognised value fails closed to none. APPLE_SERVICES (comma-separated: music, calendar,
contacts, mail, maps, weather, itunes) narrows which services register tools at all.
Confirmation (MCP_CONFIRM_MODE)
Irreversible actions and anything that reaches another person ask for confirmation first: sending mail,
deleting an event, contact, playlist or library item, removing or reordering playlist tracks, and
creating or changing an event with attendees (iCloud emails them). A client that can show a prompt
(Claude Code) asks you directly. One that cannot (claude.ai) gets a two-step flow: the first call does nothing
and returns a preview plus a confirmToken; only a repeat call with that token acts, and it is refused if the
target changed in between.
Variable | Meaning |
|
|
| How long a token is valid (default 600). |
| Token signing key; random per process by default. Set it so a token survives a restart; spent tokens are recorded on disk ( |
Environment variables
All optional; each service activates when its credentials are present. Values that are blank, undefined, null or an unexpanded ${VAR} count as unset.
Apple Developer key
Variable | Meaning |
| Your Apple Developer Team ID (10 characters). Needed for Apple Music (official API), Apple Maps and WeatherKit. |
| Key ID of a private key created in Certificates, Identifiers & Profiles → Keys with Media Services (MusicKit), MapKit JS and/or WeatherKit enabled. |
| Contents of that key's .p8 file (PEM; one-line values with \n escapes and base64 are accepted). |
| Local installs only: a path to the .p8 file instead of APPLE_PRIVATE_KEY. |
| Optional per-service override of APPLE_KEY_ID for Apple Music (pair with APPLE_MUSIC_PRIVATE_KEY). |
| Optional per-service override of APPLE_PRIVATE_KEY for Apple Music. |
| Optional per-service override of APPLE_KEY_ID for Apple Maps. |
| Optional per-service override of APPLE_PRIVATE_KEY for Apple Maps. |
| Optional per-service override of APPLE_KEY_ID for WeatherKit. |
| Optional per-service override of APPLE_PRIVATE_KEY for WeatherKit. |
| WeatherKit only: the Services ID registered for WeatherKit (e.g. com.example.weather). |
Apple Music
Variable | Meaning |
| Optional: a pre-minted Apple Music developer token (JWT) instead of signing one from the key above. |
| Music User Token for your library (official API), from a one-time MusicKit sign-in: |
| Opt-in web-player mode (no developer account needed; unlocks rename/delete/remove/reorder): the media-user-token cookie from a signed-in music.apple.com tab. |
| Optional override for the web-player developer token (normally read automatically from music.apple.com). |
| Two-letter Apple Music storefront (e.g. us, gb). Default: your account's storefront, else us. |
iCloud
Variable | Meaning |
| Your Apple ID email, for iCloud Calendar, Contacts and Mail. |
| An app-specific password from appleid.apple.com → Sign-In and Security → App-Specific Passwords (NOT your Apple ID password). |
| Your @icloud.com address, only if your Apple ID email is not an iCloud address (needed for Mail). |
| Calendar new events go to when none is named (default: the first writable event calendar). |
Behaviour
Variable | Meaning |
| "none" = read-only tools; "additive" = also create/append, never modify, delete or send; "all" = everything (default). Unrecognized values fail closed to "none". |
| Comma-separated services to enable (music, calendar, contacts, mail, maps, weather, itunes). Default: all. |
| IANA time zone (e.g. America/New_York) for displayed times and for dates you give without an offset. Set this on a hosted server, which runs in UTC. |
| "metric" (default) or "imperial" units for weather (Maps distances always show both). |
| Set to false to write nothing under $MCP_DATA_DIR/.apple-icloud-mcp: no web-player token or iCloud discovery cache, and the rejected-password latch and spent confirmation tokens then last only as long as the process. |
| Per-request timeout in milliseconds (default 30000). |
| Set to 1 to log every upstream request line to stderr (credentials redacted). |
Confirmation
Variable | Meaning |
| How confirm-gated writes (send mail, deletes, removing tracks, invitations) behave on a client with no prompt, like claude.ai: "ask-user" (default: preview + confirmToken, the model must get your OK), "auto", or "refuse". Unknown values mean refuse. |
| Lifetime of a confirmToken in seconds (default 600). |
| Signing key for confirmTokens. Random per process by default; set it so a token issued just before a restart or redeploy still works (spent tokens are recorded on disk, so none can be replayed). |
🔒 = a secret: store it as one.
Not supported (and why)
Reminders, Notes, Find My, iCloud Drive, Photos, Hide My Email. None has a public API or a standard protocol that still works: iCloud Reminders left CalDAV with iOS 13. The only route is iCloud.com's private API, which needs your full Apple ID password, an SRP sign-in and a two-factor code, has broken four times in about two years, is blocked by Advanced Data Protection, and is exactly the "automated means" the iCloud Terms forbid. On a Mac, apple-swift-mcp covers Reminders, Notes and Photos natively.
Playing music. Apple's web APIs manage your library; playback happens in an Apple Music app.
Creating or deleting calendars — iCloud's CalDAV server does not reliably support it.
Mail attachments — sending is plain text; reading lists attachments but does not download them.
Troubleshooting
Run apple_healthcheck first. For each service it reports whether credentials are configured (and which
variables to set if not), whether Apple accepted them just now, the active write mode and the display time zone.
Symptom | Likely cause |
iCloud "credentials rejected" | The app-specific password was revoked (Apple ID password changed) or the normal password was used. Generate a new app-specific password. |
Apple Music 401 | Developer key problem: wrong Team/Key ID, or the key lacks MusicKit. In web mode: the |
Apple Music 403 | The Music User Token expired (≈6 months), was minted with a different key, or the account has no Apple Music subscription. Run |
| Apple has not caught up with the previous playlist change yet (its reads lag writes by seconds), or the playlist was edited elsewhere. Re-read with |
Times are off by hours | Set |
Mail times out on a hosted deployment | Check |
APPLE_DEBUG_LOG=1 logs every upstream request line to stderr, with credentials redacted.
Security
Credentials are read from the environment at call time and never written to disk, logged, or put in a URL; every error and log line is scrubbed of every credential the process has used.
Only Apple's hosts are contacted; a redirect or DAV href pointing anywhere else is refused before any credential travels.
Reading mail never marks it read (
apple_mail_update_flagsdoes that when asked). HTML mail is converted to text with hidden content dropped, and mail and calendar text is labelled as content from its sender.Local data: small files under
$MCP_DATA_DIR/.apple-icloud-mcp/(or~/.apple-icloud-mcp/), all mode 0600 and none holding your password or tokens:music-web-token.json— Apple Music web player's own public developer token (web mode only);dav-calendar.json,dav-contacts.json— iCloud discovery URLs, bound to the credential they came from;mail-login.json— which IMAP login form iCloud accepted;icloud-rejected.json— salted digests of rejected Apple ID/password pairs and when, so a revoked password is not re-sent for 24 hours even after a restart;confirm-spent.json— digests of used confirmation tokens, so none can be replayed after a restart.
Delete the directory to remove them, or set
APPLE_STATE_CACHE=falseto never write them (the latch and the spent-token record then last only as long as the process, and a warning says so).
Development
npm run build # tsc → dist/, then esbuild bundle → dist/bundle.js
npm test # typecheck + vitest
npm run test:coverage # 100% line/branch/function/statement coverage is enforcedSee CLAUDE.md for the architecture and docs/APPLE-API.md for the provenance of every Apple API detail.
License
MIT
Available Tools
25 toolsofw_check_freshnessARead-only
Cheaply confirm whether the local cache still matches OurFamilyWizard, WITHOUT running a full sync. Use this before asserting anything about current state — especially "draft X is still sitting unsent". Costs one OFW request for the folder check plus one per messageId. For each folder it returns the live server count next to the cached count. For each id it returns a LIVE lifecycle state — "draft" | "sent" | "received" | "deleted" | "unknown" — alongside folder, sentAt, existsOnServer and a content comparison. state is the field that answers "is this still a draft?": a draft that has been SENT still exists on the server, so existsOnServer:true never distinguished the two. A cached draft whose state is no longer "draft" reports inSync:false even when its text is byte-identical. Content is compared by revision hash, because OFW draft timestamps do NOT change when a draft is edited in the web app. Does not fetch bodies into the cache, does not touch attachments, and does not depend on sync state. For draftKeys, or a full live draft inventory, use ofw_status.
| Name | Required | Description | Default |
|---|---|---|---|
| folders | No | Folders to compare cached vs live counts for. Defaults to all three when messageIds is not given. Must be non-empty if given. | |
| messageIds | No | Specific ids to verify against OFW (max 25). Ids cached as drafts, as sent messages, or as already-read inbox messages are probed freely — none of those can stamp the record. Anything else is skipped — see allowMarkRead. | |
| allowMarkRead | No | Default false. Probing an id whose cached state cannot rule out an unread INBOX message requires fetching its detail, which marks it READ on OurFamilyWizard and stamps a co-parent-visible "First Viewed" time — irreversible. Such ids are skipped (reason:"WOULD_MARK_READ") unless you set this to true. The server-wide OFW_ALLOW_MARK_READ=false is a ceiling this cannot raise. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation declares readOnlyHint: true, implying no side effects. However, the description explicitly states that setting allowMarkRead: true 'marks it READ on OurFamilyWizard and stamps a co-parent-visible "First Viewed" time — irreversible.' This is a write operation that directly contradicts the readOnlyHint. The contradiction is serious and undermines trust in the tool's safety classification.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with necessary context: costs, use cases, state semantics, side effects, and alternatives. It's front-loaded with the core purpose and avoids fluff. It could be trimmed slightly, but every detail serves the agent's decision-making.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description fully describes what is returned: live server counts per folder, and for each id: state, folder, sentAt, existsOnServer, and content comparison. It also explains the meaning of state and why existsOnServer is insufficient. Costs, side effects, and exclusions are covered. No missing information is evident for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter has a baseline description. The tool description goes well beyond this: it explains the default behavior of folders when messageIds is absent, the max of 25 ids, why certain ids are probed freely, and the full implications of allowMarkRead including the irreversible mark-read side effect. It also clarifies that content comparison uses revision hash because timestamps don't change on edit. This is substantial added meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Cheaply confirm whether the local cache still matches OurFamilyWizard, WITHOUT running a full sync.' It clearly distinguishes the tool from a full sync and from ofw_status, which it names as the alternative for draft inventories. The purpose is unambiguous and immediately actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Use this before asserting anything about current state — especially "draft X is still sitting unsent".' It also states when not to use it: 'For draftKeys, or a full live draft inventory, use ofw_status.' This is a clear routing decision with no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_create_eventA
Create a calendar event in OurFamilyWizard. Unless privateEvent is true, the event is immediately visible to the co-parent — there is no draft stage — so a shared event is confirmed first (a private one is not). If the request fails without a definitive answer the result is EVENT_UNCONFIRMED: the event may already exist, so do NOT retry until ofw_list_events shows it did not land. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call performs NO write and returns a preview of exactly what would happen plus a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| title | Yes | ||
| allDay | No | ||
| endDate | No | End date YYYY-MM-DD (default: startDate) | |
| endTime | No | End time HH:mm, 24-hour (required unless allDay) | |
| children | No | Child userIds to tag (see ofw_get_profile) | |
| location | No | ||
| startDate | Yes | Start date YYYY-MM-DD | |
| startTime | No | Start time HH:mm, 24-hour (required unless allDay) | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| privateEvent | No | true = visible only to you; default false = shared with co-parent | |
| eventParentId | No | userId of the parent the event is 'for' | |
| pickUpParentId | No | userId of the pick-up parent | |
| dropOffParentId | No | userId of the drop-off parent | |
| reminderMinutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false), the description discloses the no-draft publication behavior, the confirmation requirement, the two-step fallback with confirmToken, and the 'may already exist' uncertainty. This significantly extends what the annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause carries operational value: confirmation, visibility, failure ambiguity, and retry prohibition. It is somewhat long and could be split into shorter sentences, but none of the content is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter write tool with no output schema, the description covers the non-obvious execution model thoroughly: preview-first behavior, confirmToken reuse rule, and the EVENT_UNCONFIRMED edge case. Required/optional fields are left to the schema, which is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 67% schema coverage, the schema already handles most parameters. The description adds behavioral meaning for privateEvent (visibility to co-parent) and confirmToken (fallback-only, never first call), which are the parameters most likely to be misinterpreted by an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Create a calendar event in OurFamilyWizard.' It immediately distinguishes itself from update/delete/list siblings by its create semantics and further clarifies the sharing model with co-parent visibility.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong situational guidance: shared events need confirmation, private events do not, and on an EVENT_UNCONFIRMED failure the agent must check ofw_list_events before retrying. It does not, however, explicitly contrast this tool with ofw_update_event or state when update should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_create_expenseADestructive
Log a new expense in OurFamilyWizard. The expense is a money claim that appears in the shared ledger in front of the co-parent immediately, and this server cannot delete it. If the request fails without a definitive answer the result is EXPENSE_UNCONFIRMED: the expense may already exist, so do NOT retry until ofw_list_expenses shows it did not land. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call performs NO write and returns a preview of exactly what would happen plus a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| amount | Yes | Expense amount | |
| description | Yes | Expense description | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond annotations by disclosing that the expense appears immediately to the co-parent, that the server cannot delete it, that an ambiguous failure means EXPENSE_UNCONFIRMED, and that the first call may perform NO write in fallback mode. These are critical behavioral traits not inferable from readOnlyHint, openWorldHint, or destructiveHint alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover purpose, side effects, failure handling, confirmation behavior, and retry guidance without redundancy. The most important operational warning, the no-retry rule, is front-loaded near the beginning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description covers the full call lifecycle: preview mode, confirmToken, repeat call, ambiguous failure, and verification via ofw_list_expenses. For a mutation tool with confirmation semantics and destructive implications, this is complete enough for an agent to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by clarifying that a repeat call must use the same arguments and that the confirmToken must only be passed after explicit user approval, which reinforces and extends the confirmToken schema semantics. It does not need to restate amount or description since those are fully documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Log a new expense in OurFamilyWizard,' and clarifies that it is a money claim appearing in the shared ledger. This clearly distinguishes it from sibling tools like ofw_list_expenses and ofw_get_expense_totals.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use and when-not-to-use guidance: it instructs not to retry after an ambiguous failure until ofw_list_expenses confirms the expense did not land, and it explains the confirmation flow with a two-step fallback. This is concrete operational guidance beyond a generic description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_create_journal_entryA
Create a new journal entry in OurFamilyWizard
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Entry text content | |
| title | Yes | Entry title |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description essentially restates the tool name and adds only the product context, so it discloses no behavioral traits beyond what the annotations already provide. The destructiveHint:false annotation covers non-destructiveness, but the description does not add anything about persistence, visibility, authentication, or expected effects of creating the entry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short, front-loaded sentence with no filler or redundant clauses. It names the operation and the resource efficiently, which is appropriate for a simple two-parameter creation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple creation tool with two self-explanatory parameters, the description is largely usable, but it omits any indication of what response the agent should expect after creation. Since there is no output schema, a brief note about the result or a pointer to ofw_list_journal_entries for verification would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with 'title' and 'body' already described as 'Entry title' and 'Entry text content'. The description adds no additional parameter nuance, but because the schema fully documents the parameters, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Create') and a specific resource ('journal entry') within a named product ('OurFamilyWizard'). This clearly distinguishes it from sibling create tools like ofw_create_expense and ofw_create_event, so an agent knows which operation this is without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is only implied by the resource name and description; nothing explicitly says when to choose this over ofw_list_journal_entries, ofw_create_expense, or other journal-related tools. It is adequate for a simple create operation but provides no explicit alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_delete_draftADestructive
Delete a draft message from OurFamilyWizard. Also removes the draft from the local cache. Before deleting, the draft is re-read from OFW and the delete is REFUSED if it changed since you last read it (the current server body is returned so nothing is lost) — pass expectedRevision to assert which version you mean, or force:true to delete regardless.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Default false. Delete even if the draft changed on OurFamilyWizard since you read it. The discarded server version is echoed back in the response. | |
| messageId | Yes | Draft message ID to delete | |
| expectedRevision | No | The `revision` you got from ofw_list_drafts/ofw_get_message. Asserts you are deleting THAT version; if the draft changed on OFW since, the delete is refused and the current server body returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag destructiveHint:true, but the description goes far beyond: it details the re-read before delete, the refusal on change, the local cache removal, and the behavior of force and expectedRevision. It fully discloses the destructive nature and safety mechanisms, exceeding the annotation's minimal coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but each sentence contributes: it front-loads the primary action, then explains the cache side-effect, the safety re-read, and the two ways to handle revisions. No fluff, though it could be tightened slightly without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with 3 parameters and no output schema, the description covers all critical aspects: the action, side-effects, conflict resolution, and what is returned on refusal or force. It gives an agent everything needed to invoke it correctly and understand outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds valuable nuance: it explains that expectedRevision asserts a specific version and that force bypasses the check, with the discarded server body echoed back. This clarifies the parameters' roles beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Delete a draft message from OurFamilyWizard') and the specific resource (draft message). It distinguishes itself from siblings like ofw_save_draft by focusing on deletion and mentions cache removal, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the version-checking mechanism and how to handle conflicts via expectedRevision or force:true, giving clear usage context. It doesn't explicitly state when not to use it or mention alternatives, but the draft-specific scope and safety checks imply appropriate use. Could be improved by contrasting with ofw_save_draft for editing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_delete_eventADestructive
Delete an OurFamilyWizard calendar event. Reads the event first; deleting one the co-parent can see is confirmed first, with a preview of exactly which event (title, date, time) is removed, and is refused if the event changed on OFW after that preview. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call performs NO write and returns a preview of exactly what would happen plus a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| eventId | Yes | Event id — the `id` from ofw_list_events / eventRecurrenceId from ofw_create_event | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| includeFuture | No | For repeating events: also delete future occurrences (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, but the description goes far beyond: it discloses the read-first behavior, co-parent visibility confirmation, preview of the exact event, refusal if the event changed, and the two-step confirmToken fallback. This is exceptional transparency for a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and each subsequent sentence adds essential safety information. It is slightly long but every sentence earns its place; it could be tightened but remains effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex destructive tool with no output schema, the description covers the full confirmation flow, preview, and refusal conditions. The only minor gap is that it never explicitly states what a successful deletion returns, but the overall flow is well-covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds context about confirmToken usage in the confirmation flow, but does not introduce new parameter semantics beyond what the schema already documents. Thus a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Delete an OurFamilyWizard calendar event.' This unambiguously identifies the operation and differentiates it from sibling tools like ofw_update_event or ofw_create_event by the action verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear operational context—reads first, requires confirmation, handles repeating events with includeFuture—but it does not explicitly name alternatives or state when not to use this tool. That's a clear context with no exclusions, matching a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_download_attachmentA
Download an OFW message attachment by fileId and return content you can actually read. Inline delivery walks a ladder and returns the first rung that works: (1) host-renderable images (PNG/JPEG/GIF/WEBP) come back as ImageContent; (2) .xlsx/.csv/.tsv, .pdf, .docx, .pptx and text files come back as EXTRACTED CONTENT — per-sheet CSV, per-page/slide text, document text — in the response JSON under extracted; (3) anything else comes back as an EmbeddedResource blob of the raw bytes. The meta block names the rung as deliveredVia and, when it falls through to bytes, lists what was tried in deliveryAttempts. Reported mime types are always normalized to a bare media type (no charset/name parameters). In disk mode the bytes are saved to ~/Downloads/ofw-mcp/ and the response carries the absolute path; pass extract:true to ALSO get the extracted content in that response. The default for inline can be flipped server-side via the OFW_INLINE_ATTACHMENTS env var. On a hosted deployment with no filesystem, disk mode is unavailable, so inline is forced (forcedInline:true) rather than failing — a saveTo path never costs you the content. fileId comes from attachments[].fileId on ofw_get_message. Override disk destination with OFW_ATTACHMENTS_DIR or saveTo; saveTo must stay inside the attachments directory, and an existing file is never overwritten unless force:true. Re-downloading to the same path is a no-op (disk mode only).
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Re-download even if already on disk, replacing any existing file at the destination. Default false. Ignored when inline:true (inline always fetches fresh bytes, or reuses an on-disk copy if present). | |
| parts | No | Which sheets / slides / pages to extract, e.g. "1-3,5" (1-based positions) or a sheet name like "2026". A bare number matches either a position or a name. Omit for everything. Unselected parts are listed in `extracted.omitted`. | |
| fileId | Yes | Attachment file id (from ofw_get_message → attachments[].fileId) | |
| inline | No | If true, return content inline as MCP content blocks and skip the disk write. If false, write to disk and return the path — except on a hosted deployment with no filesystem, where inline is forced (forcedInline:true) so the content is still returned. If omitted, falls back to the OFW_INLINE_ATTACHMENTS env var (default: false = disk). | |
| saveTo | No | Path or directory to write to, INSIDE the attachments directory (OFW_ATTACHMENTS_DIR, default ~/Downloads/ofw-mcp); a relative path is resolved against it and anything outside it is refused. If a directory (trailing /), the OFW filename is used. Default: <attachments dir>/<fileId>-<filename>. An existing file is not overwritten unless force:true. Ignored when inline is in effect. | |
| extract | No | Whether to extract readable content from the file. Default: on for inline delivery of any non-image type, off in disk mode. Set false to get the raw bytes inline instead of extracted text (e.g. to hash or re-upload the file); set true in disk mode to get both the saved path and the extracted content. | |
| maxChars | No | Ceiling on extracted characters (default 50000). Over it, content is clipped on a row/line boundary, `truncated` is set, and anything dropped whole is listed in `extracted.omitted`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations only say readOnlyHint=false and destructiveHint=false, which are minimal. The description carries the full behavioral burden and does so excellently: it discloses the delivery ladder (images vs extracted content vs raw bytes), the fallback to forced inline on hosted deployments, the disk-write behavior with path and overwrite rules, mime-type normalization, and the no-op re-download behavior. Nothing contradicts the annotations, and the description adds far more context than the annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with essential details. It front-loads the core purpose and then systematically explains the delivery ladder, modes, defaults, and constraints. The numbered list and logical flow make it scannable. Every sentence contributes to understanding, with no filler. It is not overly verbose given the tool's complexity, so a 4 is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with seven parameters, one required, and no output schema, the description is remarkably complete. It covers all delivery modes, return formats (ImageContent, extracted content, EmbeddedResource blob, saved path), the meta block fields (deliveredVia, deliveryAttempts), environment variable overrides, filesystem constraints, overwrite rules, and extraction limits. It even tells the agent how to obtain fileId from a sibling tool. No critical operational detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% parameter coverage with descriptive text for all seven parameters, so the baseline is 3. The description adds meaningful value beyond the schema: it explains the interaction between extract and inline/disk mode, the default values (e.g., extract default varies by mode), the saveTo path constraints and relative resolution, and the effect of maxChars on clipping. This elevates it to a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource pair ('Download an OFW message attachment by fileId') and immediately states the output's value ('return content you can actually read'). It then enumerates three distinct delivery modes, which clearly differentiates it from siblings like ofw_upload_attachment (which does the opposite) and ofw_get_message (which provides the fileId). The purpose is unambiguous and distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong contextual guidance: it tells the agent that fileId comes from ofw_get_message's attachments[].fileId, explains when inline vs disk mode applies, and documents the default behavior and env-var override. It stops short of explicitly stating when not to use the tool or naming alternative tools beyond the fileId source, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_get_expense_totalsARead-only
Get OurFamilyWizard expense summary totals (owed/paid)
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true. The description adds context about output type (owed/paid totals) but does not disclose additional traits like time period scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no wasted words. Front-loaded with verb and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and no output schema, description sufficiently conveys purpose. Could specify time period scope but not critical for basic functionality.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist; schema coverage is 100%. Description adds value by clarifying output content. Baseline 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get' and resource 'expense summary totals (owed/paid)', distinguishing it from sibling tool 'ofw_list_expenses' which likely lists individual expenses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'ofw_list_expenses'. No explicit context for usage or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_get_messageADestructive
Get a single OurFamilyWizard message OR draft by ID. Reads from local cache when available; otherwise fetches from OFW — and for an UNREAD INBOX message that fetch marks it read and stamps a "First Viewed" time the co-parent can see, which is part of the record and cannot be undone. Pass allowMarkRead:false to refuse such a fetch instead (cached bodies, sent messages and already-read messages are unaffected, because none of them stamp anything). For ids that match a draft (in the drafts cache), the response carries folder="drafts" and the body/subject/recipients reflect the drafts cache (which ofw_sync_messages keeps fresh) — drafts have no fromUser, and sentAt/fetchedBodyAt mirror the draft's modifiedAt. For inbox/sent messages, folder is "inbox" or "sent" as before.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | Response shape: "compact" (default) drops fields the response already carries elsewhere; "full" returns every field this server understands. compact omits OurFamilyWizard's raw `listData` echo, which duplicates this record's own id, subject, sentAt, recipients and read flag; the sender is promoted to `from`, and `files`/`replied` are kept. Pass "full" for the echo. | |
| messageId | Yes | Message ID (also accepts draft IDs — drafts are routed via the drafts cache) | |
| allowMarkRead | No | Default true (the long-standing behaviour). Set false to refuse a fetch that would mark an unread INBOX message as READ on OurFamilyWizard — an irreversible, co-parent-visible change to the record. Reads that cannot stamp anything (a cached body, a sent message, an already-read message) still succeed. The server-wide OFW_ALLOW_MARK_READ=false is a ceiling this argument cannot raise. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses the precise side effect: fetching an unread inbox message marks it read, stamps a co-parent-visible 'First Viewed' time, and this cannot be undone. It also explains which reads are side-effect-free and how the allowMarkRead parameter interacts with the server-wide setting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries distinct operational information: cache behavior, side-effect warning, opt-out path, draft routing, and folder semantics. It is front-loaded with the core purpose and side-effect warning before the detailed draft behavior, and it is appropriately sized for a tool with this much nuance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and it mostly delivers: it specifies folder values, draft-specific fields, mirroring of timestamps, and the compact/full distinction is covered in the schema. Minor gaps include no explicit failure/not-found behavior, but overall the agent has enough context to call this tool and interpret its main response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is already met. The description adds real meaning beyond the schema, especially for allowMarkRead (irreversible co-parent-visible stamp, refusal behavior, unaffected cases) and for draft IDs (folder="drafts", no fromUser, sentAt/fetchedBodyAt mirroring modifiedAt). The view parameter is only covered by the schema, but the added parameter semantics still push this above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Get a single OurFamilyWizard message OR draft by ID." It clearly distinguishes this single-record retrieval tool from the list-oriented siblings and immediately communicates the draft/message dual behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the core usage context clear: reads cache when available, otherwise fetches, and tells the agent exactly when to pass allowMarkRead:false. It does not explicitly name list-oriented alternatives, but the single-by-ID semantics are unambiguous enough that an agent can select this tool correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_get_notificationsARead-only
Get OurFamilyWizard dashboard summary: unread message count, upcoming events, outstanding expenses. Note: updates your last-seen status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a critical side effect: 'updates your last-seen status.' This is valuable behavioral context beyond the readOnlyHint annotation, which might otherwise lead an agent to assume no state changes. The annotation says readOnlyHint=true, but the description correctly warns of a state mutation, so there is no contradiction—the description adds important nuance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the tool's purpose and then adds the critical side-effect warning. Every word earns its place; no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description adequately covers what the tool does and its side effect. It could be slightly more complete by noting whether the returned summary includes counts only or also details, but the core information an agent needs is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description doesn't need to explain parameter meaning. The baseline for 0 params is 4, and the description appropriately focuses on what the tool returns rather than parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves a dashboard summary with specific content types (unread message count, upcoming events, outstanding expenses). It distinguishes itself from siblings like ofw_list_messages or ofw_list_events by being a summary rather than a detailed list, though it doesn't explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the tool to use when you need an overview of multiple data types at once, rather than detailed lists. However, it doesn't explicitly state when to use this over alternatives like ofw_list_messages or ofw_list_events, nor does it mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_get_profileARead-only
Get current user and co-parent profile information from OurFamilyWizard
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description correctly indicates a read operation, consistent with the readOnlyHint annotation. It adds no extra behavioral details beyond what the annotation already provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that directly conveys the tool's purpose without unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple nature of the tool (no parameters, readOnlyHint annotation), the description is complete. No output schema is needed for such a straightforward retrieval.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the baseline is 4 as per guidelines. The description does not need to elaborate on parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (get) and the resource (current user and co-parent profile information). It implicitly distinguishes from sibling tools that handle notifications, messages, events, expenses, and journal entries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when profile information is needed but provides no explicit guidance on when not to use it or alternatives. With zero parameters and no sibling profile tools, this is adequate but minimal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_get_unread_sentARead-only
List sent messages that have not been read by one or more recipients. Reads from local cache. Returns complete describing whether every sent message was scanned. An empty SENT cache that is not verified-fresh is REFUSED (result:"UNVERIFIED_EMPTY") rather than reported as "nothing sent"; pass autoRefresh:true to sync and answer instead.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page (default 1) | |
| size | No | Per page (default 50) | |
| autoRefresh | No | If the result comes back EMPTY from a cache that is not verified-fresh, sync the backing folders first and answer from the refreshed cache instead of refusing. Defaults to the OFW_AUTO_REFRESH env var (false unless set), in which case the call refuses with result:"UNVERIFIED_EMPTY" and names the remedy. Costs OFW requests when it fires. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description exposes important behavior: it reads from local cache, includes a `complete` flag, and can refuse with UNVERIFIED_EMPTY if the cache is not fresh. The autoRefresh remedy is also disclosed, adding substantial value beyond annotations. No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured. It leads with the main action, then covers cache behaviorкну, return indicator, and the key edge case in a few efficient sentences with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with fully documented optional parameters, the description is largely complete: it explains purpose, cache behavior, a key refusal case, and the remedy. The only notable omission is the shape of the returned message list itself, which is more noticeable because no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters are fully documented in the schema with descriptions, including autoRefresh's behavior and defaults. The description reinforces autoRefresh's effect but adds no new page/size semantics, so the baseline 3 for high schema coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('sent messages'), and a precise filter ('not been read by one or more recipients'). This semantically distinguishes it from sibling tools like ofw_list_messages without requiring the reader to open schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context for using the tool is clear: it lists unread sent messages and can refresh on empty results. However, it does not explicitly name sibling alternatives or state when not to use this tool, leaving usage guidance somewhat implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_healthcheckVerify credentials and upstream reachabilityARead-onlyIdempotent
Resolves the credential the way real tools do, then makes one authenticated request to ourfamilywizard.com. Reports which source supplied the credential, whether ourfamilywizard.com accepted it, the round-trip time, and a plain-English hint distinguishing 'no credential' from 'credential rejected' from 'a ourfamilywizard.com-side problem'. Read-only; never returns the credential itself. Call this when a real tool fails and you want to know which hop broke.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds valuable behavioral context: it explicitly states 'Read-only; never returns the credential itself' and describes the exact outputs (source, acceptance, round-trip time, and a plain-English hint distinguishing failure modes). It also mentions that it makes one authenticated request, which is a side effect not covered by annotations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is approximately 70 words and front-loads the core action ('Resolves the credential... makes one authenticated request') before listing outputs and usage. Every sentence adds value: the credential resolution, the request, the reported metrics, the read-only/never-returns-credential note, and the when-to-call guidance. There is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic tool with no output schema, the description adequately explains what it does, what it reports, and when to use it. It specifies the three failure categories and mentions round-trip time, giving the agent a clear expectation of the result. No critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (vacuously). The description does not need to explain parameters because there are none. The baseline for 0 params is 4, and no additional parameter semantics are required or provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: resolving the credential as real tools do and making an authenticated request to test upstream reachability. It explicitly distinguishes this healthcheck from the sibling tools (messages, expenses, etc.) by naming it as a diagnostic that reports credential source, acceptance, and round-trip time. The verb 'resolves' and 'makes' are specific, and the resource is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to call: 'Call this when a real tool fails and you want to know which hop broke.' This gives a clear trigger condition and implies it's for troubleshooting, not routine use. It does not mention alternatives, but the context is sufficient because this tool is unique among siblings for diagnostics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_list_draftsA
List draft messages, verified against OurFamilyWizard in ONE call: when the local drafts cache is not verified-fresh, a cheap drafts sync runs first by default (verify:true), so the answer is server-confirmed without a second call. Pass verify:false to answer purely from the cache (no OFW requests). Returns an explicit complete boolean describing the RESULT SET: true means "these are ALL the drafts on OurFamilyWizard as of freshness.asOf" — check it before saying "you have N drafts". Each draft carries its draftKey (stable across the create-then-delete churn of editing) when one is known. An empty result from a cache that is not verified-fresh is REFUSED (result:"UNVERIFIED_EMPTY"); pass autoRefresh:true to sync and answer instead.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number (default 1) | |
| size | No | Drafts per page (default 50) | |
| view | No | Response shape: "compact" (default) drops fields the response already carries elsewhere; "full" returns every field this server understands. compact omits OurFamilyWizard's raw `listData` echo, which duplicates this draft's own id, subject, modifiedAt and recipients. `revision`, `draftKey` and `cacheStatus` are kept on both rungs. | |
| verify | No | Default true: when the drafts cache is not verified-fresh, run a drafts sync first (cheap — one list page plus one detail per draft) so the response is server-confirmed in one call. Set false to serve straight from the local cache with no OFW requests. | |
| autoRefresh | No | If the result comes back EMPTY from a cache that is not verified-fresh, sync the backing folders first and answer from the refreshed cache instead of refusing. Defaults to the OFW_AUTO_REFRESH env var (false unless set), in which case the call refuses with result:"UNVERIFIED_EMPTY" and names the remedy. Costs OFW requests when it fires. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses behavior far beyond the annotations: default sync behavior, what 'complete' means semantically, the stable draftKey across draft edit churn, and the UNVERIFIED_EMPTY refusal plus remedy. This is especially valuable because there is no output schema and readOnlyHint=false is otherwise under-explanatory. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries a unique, decision-relevant fact: one-call verification, cache-only mode, complete semantics, draftKey stability, empty-refusal and autoRefresh remedy. It is front-loaded with the core one-call promise and is as tight as the complexity allows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining return semantics, and it does: complete boolean, result UNVERIFIED_EMPTY, draftKey, cacheStatus, view shapes, and failure remedies. An agent has enough to call it correctly, interpret the response, and handle edge cases without further investigation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents every parameter. The description adds meaningful semantics for verify, autoRefresh, and view beyond their schema text, including defaults, cost implications, and result-set consequences. However, page and size receive no additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'List draft messages', then sharpens it with 'verified against OurFamilyWizard in ONE call'. It clearly distinguishes this from draft-mutating siblings like ofw_save_draft and ofw_delete_draft, and from ofw_list_messages, by focusing on drafts and their freshness/verification semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit conditions for parameter choices: 'Pass verify:false to answer purely from the cache' and 'pass autoRefresh:true to sync and answer instead' when UNVERIFIED_EMPTY occurs. It does not, however, explicitly contrast this tool with sibling list tools (e.g., when to use list_drafts vs list_messages), so it stops short of full alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_list_eventsARead-only
List OurFamilyWizard calendar events in a date range
| Name | Required | Description | Default |
|---|---|---|---|
| endDate | Yes | End date YYYY-MM-DD | |
| detailed | No | Return full event details (default false) | |
| startDate | Yes | Start date YYYY-MM-DD |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds no further behavioral information such as default summary vs detailed output or ordering, but it does not contradict the annotations. With annotations carrying the safety profile, the description's limited addition earns a baseline 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the verb and resource, then states the date-range filter. It contains no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This simple list tool has only three parameters and an annotation for read-only safety, so the description covers the core purpose. However, with no output schema, the description does not explain the shape or fields of the returned events, and the detailed parameter's effect is only documented in the schema. This leaves a modest gap in what the agent can expect as a result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides full descriptions for all three parameters at 100% coverage, including formats for startDate/endDate and the effect of the detailed flag. The description does not add semantic detail beyond the schema, so it remains at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the verb 'List' with a specific resource ('OurFamilyWizard calendar events') and a clear scope ('in a date range'), which distinguishes it from mutation siblings like ofw_create_event and ofw_delete_event. No ambiguity exists about the tool's high-level function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool is for retrieving events within a date range, aligned with the required startDate and endDate parameters. It does not explicitly name alternatives or exclusions, but the sibling tool names make the read-only usage context obvious, providing clear context without explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_list_expensesARead-only
List OurFamilyWizard expenses. Offset-paged via start/max. The response leads with its paging state — hasMore and nextStart (null when the list is exhausted) — BEFORE the records, so a truncated or partially-read response still says whether more remain. Never state an expense total or an absence from one page.
| Name | Required | Description | Default |
|---|---|---|---|
| max | No | Max results (default 20) | |
| start | No | Start offset, 0-based (default 0). To continue a listing, pass the `nextStart` from the previous response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses a non-obvious response behavior – paging state (`hasMore`, `nextStart`) appears BEFORE the records, and `nextStart` is null when exhausted – which annotations (only readOnlyHint) do not cover. The warning about not stating totals or absence adds safety-critical context beyond the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose, pagination mechanism, response behavior, and a caution. The verb and resource are front-loaded, with zero filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with full schema coverage and a readOnly annotation, the description covers purpose, pagination, and a key limitation. It does not enumerate the fields inside each expense record, which an agent might need if it must process the data, but the tool's name and warning reduce the risk of misuse. Minor gap only.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: both parameters have descriptive text in the input schema (`max` with default 20, `start` with offset semantics and nextStart guidance). The description's mention of 'Offset-paged via start/max' and the nextStart instruction essentially echoes schema content, adding no new parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'List OurFamilyWizard expenses' – a specific verb and resource that unambiguously identifies the tool's function. It distinguishes this from siblings like ofw_get_expense_totals and ofw_list_messages by naming the exact resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear pagination usage ('Offset-paged via start/max', 'pass the nextStart from the previous response') and a critical caution ('Never state an expense total or an absence from one page') that implies when not to rely on this tool. However, it does not explicitly name alternative sibling tools (e.g., ofw_get_expense_totals) for those cases, leaving some routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_list_journal_entriesARead-only
List OurFamilyWizard journal entries. Offset-paged via start/max (1-based). The response leads with its paging state — hasMore and nextStart (null when the list is exhausted) — BEFORE the records, so a truncated or partially-read response still says whether more remain. Never state an entry count or an absence from one page.
| Name | Required | Description | Default |
|---|---|---|---|
| max | No | Max results (default 10) | |
| start | No | Start offset, 1-based (default 1). To continue a listing, pass the `nextStart` from the previous response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even with readOnlyHint=true already present, the description adds substantive behavioral detail: the response leads with paging state, nextStart is null when exhausted, and a truncated/partial response still indicates whether more records remain. It also warns against overclaiming counts or absence from a single page, which is valuable beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded with the tool's purpose. The pagination and caution sentences earn their place, though they are slightly dense; the explicit 'BEFORE' emphasis and 'Never state...' instruction could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by explaining the paging state placement and exhaustion semantics, which are the main risks when invoking a paged read tool. It does not enumerate journal entry fields or sorting, but those are not required for correct invocation and selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters, including defaults and the nextStart continuation behavior, so the description does not need to compensate. The description adds the 'offset-paged' framing and 1-based emphasis, but this largely repeats what the schema provides; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List OurFamilyWizard journal entries.' This cleanly distinguishes it from sibling list tools like ofw_list_messages, ofw_list_events, and ofw_list_expenses. The added pagination details reinforce the intended operation without obscuring the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear invocation context, including the 1-based offset paging via start/max and how to continue from nextStart. It does not explicitly name alternative tools or exclusion conditions, but the purpose is specific enough that an agent knows when this listing tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_list_message_foldersARead-only
List OurFamilyWizard message folders (inbox, sent, etc.) and their unread counts. Fetched LIVE from OFW, so the counts are current. Returns folder IDs needed to call ofw_list_messages. Does NOT return message content.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, so the description does not need to repeat that. It adds valuable behavioral context: 'Fetched LIVE from OFW, so the counts are current.' This goes beyond the annotation by explaining the data freshness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences, each adding value. The first sentence states the core purpose, the second adds behavioral detail, and the third clarifies what is not returned. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema), the description is complete. It explains the action, the output (folder IDs and counts), the distinction from message content, and the live nature of the data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the baseline is 4. The description does not need to add parameter details, and it correctly focuses on the tool's purpose and output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List', the resource 'OurFamilyWizard message folders', and specifies that it returns folder IDs and unread counts. It distinguishes itself from sibling tools like ofw_list_messages by explicitly noting that it does not return message content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage hint: 'Returns folder IDs needed to call ofw_list_messages', implying the tool is a prerequisite. It also warns that it does not return message content, preventing misuse. However, it does not explicitly list when not to use or mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_list_messagesARead-only
List messages from the local OurFamilyWizard cache. Supports filtering by folder, date range, and a substring query on subject+body. Pagination is offset-based (1-based page) but if you know what you want (a date range, a topic), prefer the filters over walking pages — the cache may have 1000+ messages. Results are newest-first by default; sort:"oldest" starts at the old end of a range instead of paging to it. Returns an explicit complete boolean describing the RESULT SET: true means "this is every message on OurFamilyWizard matching these filters as of freshness.asOf" — check it before asserting a count. An empty result from a cache that is not verified-fresh is REFUSED (result:"UNVERIFIED_EMPTY") rather than reported as an absence; pass autoRefresh:true to sync and answer instead.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Substring match on subject AND body (case-insensitive). Use to find messages on a specific topic. | |
| page | No | Page number (default 1) | |
| size | No | Messages per page (default 50) | |
| sort | No | Result order: "newest" (default, newest first) or "oldest" (oldest first). This decides which end a truncated page keeps — with "newest" page 1 of a wide date range holds its most RECENT slice, with "oldest" its earliest. Use "oldest" to start at the old end of a range instead of paging to it. | |
| view | No | Response shape: "compact" (default) drops fields the response already carries elsewhere; "full" returns every field this server understands. compact omits OurFamilyWizard's raw `listData` echo, which duplicates this record's own id, subject, sentAt, recipients and read flag; the sender is promoted to `from`, and `files`/`replied` are kept. Pass "full" for the echo. | |
| since | No | ISO date or datetime — only messages with sent_at >= since (inclusive). A value with an offset or Z is compared as that instant; a naive value is read as the account's local time (DISPLAY_TZ) | |
| until | No | ISO date or datetime — only messages with sent_at < until (exclusive). A value with an offset or Z is compared as that instant; a naive value is read as the account's local time (DISPLAY_TZ) | |
| folderId | No | Folder name: "inbox", "sent", or "both" (default "both") | |
| autoRefresh | No | If the result comes back EMPTY from a cache that is not verified-fresh, sync the backing folders first and answer from the refreshed cache instead of refusing. Defaults to the OFW_AUTO_REFRESH env var (false unless set), in which case the call refuses with result:"UNVERIFIED_EMPTY" and names the remedy. Costs OFW requests when it fires. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint:true, and the description adds substantial behavioral context: cache-based results, the complete boolean describing result-set completeness, the UNVERIFIED_EMPTY refusal for unverified caches, autoRefresh sync behavior, sort semantics affecting which end of a range a page keeps, and view field shaping. No contradiction with annotations; this exceeds the baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence adds value. It is front-loaded with the core purpose, then flows into filtering, pagination/sort, result-set semantics, and edge cases. There is no fluff or redundancy; the structure is logical and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter list tool with no output schema and only readOnlyHint annotation, the description covers all critical aspects: caching, freshness, pagination, sort behavior, response shape via view, and the complete boolean. It also explains the UNVERIFIED_EMPTY refusal and autoRefresh remedy. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all 9 parameters are documented, but the description adds significant meaning beyond the schema: it explains sort's effect on paging (start at old end vs paging to it), the compact vs full view trade-offs (dropping duplicate fields, promoting sender to from), and autoRefresh's cost and default behavior. These enrich the parameter semantics well beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List messages from the local OurFamilyWizard cache' with a specific verb and resource. It enumerates filtering options (folder, date range, substring) and pagination, distinguishing itself from sibling tools like ofw_get_message (single message) and ofw_list_drafts. The scope is precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance: prefer filters over walking pages when a specific date range or topic is known, and pass autoRefresh:true when the cache might be stale. It implies when to use this list tool vs alternatives, though it does not name specific sibling tools or explicit exclusions. The context is clear enough for an agent to decide when to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_save_draftA
Save a message as a draft in OurFamilyWizard. RECIPIENTS: OurFamilyWizard does NOT persist recipients on drafts — recipientIds are accepted but the saved draft comes back with none (documented OFW behavior, noted once in the response, not warned about; supply recipientIds at send time instead). IDENTITY: the response leads with draftKey, the stable identity that survives editing — key off it, because the id changes on EVERY edit (replacing a draft creates a NEW draft and deletes the old one; OFW's update-in-place endpoint silently no-ops, so we never use it). Pass messageId to replace an existing draft; the response.id will be the NEW id, and a transparency NOTE documents the swap and which fields were carried over. THREADING: if replyToId is provided, the cache may rewrite it to the latest reply in the thread (note included). The threading verdict is read from OFW's full echo (replyToId/inReplyTo/showContext) — a warning appears ONLY when the reply linkage was genuinely dropped or re-targeted, and the response's top-level replyToId/inReplyTo always agree with its listData. Attach files via myFileIDs (from ofw_upload_attachment). After saving, the tool re-fetches the draft from OFW, and the returned revision reflects that authoritative state (so it will match on your next edit). SAFETY: because replacing DESTROYS the old draft rather than merging, passing messageId first re-reads that draft from OFW and REFUSES the write if its subject/body/recipients changed since you read it (drafts edited in the OFW web app do not bump any timestamp, so the local cache can be silently behind). A pure replyToId normalization by OFW is NOT treated as a conflict. The refusal returns the current server body under serverBody — merge your edit into it and retry with expectedRevision.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Message body text | |
| force | No | Default false. Overwrite even when the draft changed on OurFamilyWizard since you read it. The discarded server version is echoed back in the response. Only use after showing the user the conflict. | |
| subject | Yes | Message subject | |
| messageId | No | ID of an existing draft to replace (the new draft will have a new id; the old is deleted) | |
| myFileIDs | No | Attachment file ids (from ofw_upload_attachment) | |
| replyToId | No | ID of the message this draft replies to | |
| recipientIds | No | Array of recipient user IDs (optional for drafts) | |
| expectedRevision | No | With messageId: the `revision` you got from ofw_list_drafts/ofw_get_message for that draft. Asserts you are replacing THAT version. If the draft changed on OFW since, the write is refused and the current server body is returned. Omit and the tool compares the server against the local cache instead — omitting never means "overwrite anyway". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnlyHint=false and destructiveHint=false, but the description goes far beyond: it warns that replacing a draft DESTROYS the old one, explains that the id changes on every edit while draftKey is stable, describes the conflict-check refusal, and details the re-fetch behavior returning authoritative revision. It also covers threading normalization and the fact that recipientIds are not persisted. All of this is behavioral context not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place. It is structured into labeled sections (RECIPIENTS, IDENTITY, THREADING, SAFETY) that are easy to scan. The core purpose is front-loaded, and each paragraph covers a distinct behavior with concrete examples. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters, destructive side effects, identity quirks, and conflict safety, this description is exhaustive. It covers the return values (revision, serverBody, draftKey, response.id), edge cases (thread re-targeting, silent cache lag), and the exact workflow for safe replacement. No output schema exists, so the description must carry that burden, and it does completely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already described. However, the description adds substantial semantic depth: it explains that recipientIds are accepted but not persisted, that messageId triggers a delete-and-replace rather than an in-place update, that expectedRevision asserts a specific version and the fallback to local cache, and that force overrides the conflict check with the server body echoed. This goes well beyond the schema's basic type/description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource statement: 'Save a message as a draft in OurFamilyWizard.' It immediately clarifies the core action and contrasts with siblings like ofw_send_message and ofw_delete_draft. The subsequent paragraphs detail the identity and safety nuances, leaving no ambiguity about what this tool does versus others.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool and when not to: 'supply recipientIds at send time instead' for recipients, 'Pass messageId to replace an existing draft' with an explanation of how to replace, and it even notes that OFW's update-in-place endpoint is never used. It provides clear guidance on expectedRevision and force parameters, including the condition 'Only use after showing the user the conflict.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_send_messageADestructive
Send a message via OurFamilyWizard — the ONE irreversible operation here, so it carries the strongest guard. TO SEND AN EXISTING DRAFT (the safe default): pass draftId (or messageId — same thing). The tool re-reads the draft from OFW and sends the SERVER'S version, so what goes out is what is on OurFamilyWizard, not what this session remembers — subject/body act only as explicit overrides. It is guarded exactly like ofw_save_draft: pass expectedRevision to assert which version you are sending; if the draft changed on OFW since you read it — or no longer exists (it may already have been SENT) — the send is REFUSED with the current server content echoed back, and nothing goes out. RECIPIENTS: OurFamilyWizard does not persist recipients on drafts, so recipientIds is usually still required at send time (ids from ofw_get_profile). After the send is CONFIRMED (OFW returned the new message id and the re-fetched sent record matches what was posted), the source draft is deleted automatically; pass deleteDraftOnSuccess:false to keep it. On ANY failure or ambiguity the draft is never deleted — the response carries draftRetained:true with the reason. If the send request times out or drops without a definitive answer, the result is SEND_UNCONFIRMED: the message may already have been delivered, so do NOT retry until a sent-folder sync (or ourfamilywizard.com) shows it did not go out. TO COMPOSE FROM SCRATCH: supply subject/body/recipientIds with no draftId. If replyToId is provided (or inherited from the draft), the cache may rewrite it to the latest reply in the same thread (a note is included when this happens). ATTACHMENTS: when sending by draftId, the server draft's own attachments carry over automatically; myFileIDs (from ofw_upload_attachment) overrides or attaches files on a fresh compose. The response leads with sentMessageId and the stable draftKey, and reports threaded (whether OFW actually linked the reply) and draftDeleted. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call performs NO write and returns a preview of exactly what would happen plus a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Message body text. Required unless draftId/messageId is given (then it overrides the server draft's body — omit it to send exactly what is on OurFamilyWizard). | |
| force | No | Default false. Send even when the draft changed on OurFamilyWizard since you read it, or its current state could not be read. Only use after showing the user the conflict. | |
| draftId | No | ID of an existing draft to send. The draft is re-read from OurFamilyWizard and its SERVER content is sent; missing subject/body default from it. Guarded: a draft that changed since you read it, or that was already sent/deleted, refuses rather than sending blind. | |
| subject | No | Message subject. Required unless draftId/messageId is given (then it overrides the server draft's subject). | |
| messageId | No | Synonym for draftId (if both are passed they must be equal). | |
| myFileIDs | No | Attachment file ids (from ofw_upload_attachment) to attach to the message. When sending by draftId, omit it to carry the server draft's own attachments over; passing it overrides them. | |
| replyToId | No | ID of the message being replied to. Defaults to the draft's stored reply target when sending by draftId. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| recipientIds | No | Array of recipient user IDs (get from ofw_get_profile). Usually required even when sending a draft: OurFamilyWizard does not persist recipients on drafts. | |
| expectedRevision | No | With draftId: the `revision` from ofw_list_drafts / ofw_get_message / ofw_check_freshness for that draft. Asserts you are sending THAT version; if the draft changed on OFW since, the send is refused and the current server content returned. Omit and the tool compares the server against the local cache instead — omitting never means "send whatever is there now". | |
| deleteDraftOnSuccess | No | Default true. Delete the source draft after — and ONLY after — the send is confirmed (new message id returned and the re-fetched sent record checks out). Set false to keep the draft. On a failed or unverifiable send the draft is ALWAYS kept, regardless of this flag. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate a destructive write (readOnlyHint=false, destructiveHint=true), but the description discloses a wealth of behavioral detail: the irreversible deletion of the source draft, guard/refusal on revision mismatch, SEND_UNCONFIRMED uncertainty, automatic deletion only after confirmation, cache rewriting of replyToId, and the two-step confirmation fallback. It even explains what happens on timeout or failure, leaving little to the agent's imagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is an extensive multi-paragraph monologue, likely over a thousand words, and repeats information already present in the input schema (e.g., deleteDraftOnSuccess, expectedRevision, force). While the key warning is front-loaded, the density of parentheticals and redundant explanations makes it hard to scan, and much could be condensed into a structured summary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—11 parameters, no output schema, destructive side effects, confirmation flow—the description is exceptionally complete. It covers the response shape (sentMessageId, draftKey, threaded, draftDeleted), the SEND_UNCONFIRMED edge case, attachment inheritance, and the confirmation fallback, so an agent has all necessary context to invoke the tool correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds crucial semantic interconnections: draftId/messageId refer to the server's version, subject/body act as explicit overrides, expectedRevision asserts a specific version, recipientIds are still required because OFW does not persist draft recipients, myFileIDs override server attachments, and confirmToken has a strictly defined lifecycle. These nuances go far beyond the per-parameter schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool sends a message via OurFamilyWizard and explicitly labels it 'the ONE irreversible operation here', distinguishing it from siblings like ofw_save_draft and ofw_delete_draft. It further breaks down two distinct use modes (sending an existing draft vs composing from scratch), leaving no ambiguity about the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: the safe default of sending a draft via draftId, when to compose fresh, when force may be used after user confirmation, and a concrete directive to avoid retry on SEND_UNCONFIRMED until a sync verifies the message did not send. It also references ofw_save_draft's guarding mechanism and warns that recipientIds are usually still required, making alternative selection clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_statusARead-only
ONE live call that answers "where does everything stand?". This is the call that should back any status summary about drafts or specific messages — never session memory, and never a cached read alone. With no arguments it returns the FULL current draft inventory, verified against OurFamilyWizard. Pass ids and/or draftKeys to get each one's live lifecycle state ("draft" | "sent" | "received" | "deleted" | "unknown") with sentAt and viewedAt. A draftKey is the stable identity ofw_save_draft returns: editing a draft mints a new OFW id every time (create-then-delete), so the key is the only way to ask "what happened to the thing I was working on?" — it resolves to the chain's current id and keeps resolving after the draft is SENT (state:"sent" with sentMessageId). The top-level complete is true ONLY when every part of this snapshot was verified live; if it is false, do not state a draft count or a lifecycle claim from this payload.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | No | Message/draft ids to resolve to a live state (combined with draftKeys, max 25 probes per call). | |
| draftKeys | No | Stable draft keys (from ofw_save_draft / ofw_list_drafts) to resolve to their CURRENT id and state. | |
| allowMarkRead | No | Default false. An id whose cached state cannot rule out an unread INBOX message can only be probed by fetching its detail, which marks it READ on OurFamilyWizard — irreversible and co-parent-visible. Those are skipped unless this is true. Cached drafts, sent messages and already-read messages are always probed. Capped by OFW_ALLOW_MARK_READ. | |
| includeDraftInventory | No | Return the full current draft list, verified against OurFamilyWizard first. Defaults to TRUE when neither ids nor draftKeys is given (so a bare ofw_status() is a complete status snapshot), otherwise false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint, it discloses the freshness contract ('complete is true ONLY when every part... verified live'), the create-then-delete id churn behind draftKeys, and the irreversible co-parent-visible mark-read side effect gated behind allowMarkRead. This is exactly the kind of behavioral context annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense; the main purpose is front-loaded and every subsequent sentence covers a needed behavioral or parameter nuance. Nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, it still tells the agent what comes back: full inventory, lifecycle states, sentAt/viewedAt, and the meaning of top-level complete. It also covers defaults and safety caps, so an agent has enough to call it correctly in any mode.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds meaning the schema lacks: max 25 probes, draftKey stability semantics across edits and sends, the default of includeDraftInventory based on other args, and the OFW_ALLOW_MARK_READ cap. These clarify how to invoke the tool rather than just what each field is.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a concrete job ('ONE live call that answers where does everything stand?') and specifies exactly what it returns: full verified draft inventory and per-message lifecycle state. It distinguishes itself from cached reads and session memory, so an agent can tell it apart from list/get siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the intended use explicitly ('should back any status summary about drafts or specific messages') and gives negative guidance ('never session memory, and never a cached read alone'). It also defines when each parameter mode applies: bare call for a full snapshot, ids/draftKeys for targeted probes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_sync_messagesA
Sync messages from OurFamilyWizard into the local cache. Returns counts per folder and a list of unread inbox messages whose bodies were NOT fetched (to avoid mark-as-read on OFW). Call ofw_get_message(id) on those to read them. EVERY call re-checks the newest page first, so new messages are picked up promptly even while an old-history backfill is still running; only then does it spend what is left of its budget advancing that backfill. Pass deep:true to walk all OFW pages instead of stopping at the first all-cached page (use to backfill suspected gaps). Sync is BOUNDED and RESUMABLE: on hosted deployments a per-call OFW-request budget (env OFW_SYNC_MAX_REQUESTS, or the maxRequests argument) caps how far one call walks; when the budget is hit the response reports done:false with a note — call again with the SAME arguments to resume. done:false means older history is still being backfilled; it does NOT mean recent messages are missing. Local installs are unbounded by default (done is always true).
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | If true, walk every OFW page until empty regardless of cache state. Use to backfill gaps. Default false. | |
| folders | No | Folders to sync (default: all three). Must be non-empty if given — an empty list would sync nothing while reporting success. | |
| maxRequests | No | Maximum OFW requests this single call may make before pausing. When hit, the response reports done:false — call again with the same arguments to continue. Omit to use the server default (OFW_SYNC_MAX_REQUESTS, or unbounded on local installs). | |
| fetchUnreadBodies | No | If true, also fetch bodies for unread inbox messages — which marks each one READ on OurFamilyWizard and stamps a co-parent-visible "First Viewed" time that cannot be undone. Defaults to the OFW_FETCH_UNREAD_BODIES env var (false unless set), and is forced off entirely when OFW_ALLOW_MARK_READ=false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations. It discloses that the sync is bounded and resumable, that every call re-checks the newest page first, that done:false means older history is still backfilling (not that recent messages are missing), and that fetchUnreadBodies marks messages read with a co-parent-visible timestamp. It also explains the budget mechanism and local vs. hosted behavior. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized, front-loading the core purpose and return shape before diving into behavioral details. Every sentence earns its place, though the length is substantial. The structure is logical: purpose, return shape, follow-up action, sync behavior, deep flag, budget/resume semantics, and local vs. hosted distinction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex sync tool with 4 parameters, no output schema, and no annotations covering safety, the description is remarkably complete. It explains the return value shape, the resume mechanism, the meaning of done:false, the deep flag's purpose, and the mark-as-read side effect. An agent has everything needed to call this tool correctly and interpret its response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all four parameters. The description adds value by explaining the deep:true use case (backfill suspected gaps), the budget/resume semantics for maxRequests, and the mark-as-read consequence of fetchUnreadBodies. It doesn't add much beyond the schema for folders, but the added context for the other parameters justifies a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Sync messages from OurFamilyWizard into the local cache') and immediately distinguishes itself from siblings by explaining the return shape (counts per folder, unread inbox messages without bodies) and the follow-up call (ofw_get_message). It clearly identifies what this tool does and how it differs from related message tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: it explains when to pass deep:true (backfill suspected gaps), when to call again (when done:false), and what done:false means vs. does not mean. It also names the alternative for fetching bodies (ofw_get_message) and explains the mark-as-read tradeoff. This is comprehensive usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_update_eventADestructive
Update an existing OurFamilyWizard calendar event. Fetches the event, applies the given changes, and writes the merged result back (OFW has no partial update). A change to an event the co-parent can see (shared before or after the change) is confirmed first; the confirmation is bound to the event exactly as read, so if it changes on OFW in between (say the co-parent edited it) the update is refused instead of overwriting their edit. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call performs NO write and returns a preview of exactly what would happen plus a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| title | No | ||
| allDay | No | ||
| endDate | No | End date YYYY-MM-DD (default: startDate) | |
| endTime | No | End time HH:mm, 24-hour (required unless allDay) | |
| eventId | Yes | Event id — the `id` from ofw_list_events / eventRecurrenceId from ofw_create_event | |
| children | No | Child userIds to tag; pass [] to remove all child tags (omit to keep current tags) | |
| location | No | ||
| startDate | No | Start date YYYY-MM-DD | |
| startTime | No | Start time HH:mm, 24-hour (required unless allDay) | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| privateEvent | No | true = visible only to you; default false = shared with co-parent | |
| eventParentId | No | userId of the parent the event is 'for' | |
| pickUpParentId | No | userId of the pick-up parent | |
| dropOffParentId | No | userId of the drop-off parent | |
| reminderMinutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the destructiveHint annotation by explaining the read-modify-write behavior, the lack of partial updates in OFW, the confirmation workflow, optimistic concurrency protection, and the two-step fallback with confirmToken. This gives the agent an unusually complete picture of what happens when the tool is invoked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence carries distinct, necessary information: purpose, merge behavior, confirmation and concurrency guarantees, and the fallback mechanism. It is front-loaded with the core action and avoids filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the destructiveHint annotation, and the absence of an output schema, the description covers all essential operational aspects: what it updates, how partial updates are handled, when confirmation is required, and what happens on concurrent modification. Nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 69%, and the description adds meaningful process-level semantics: 'applies the given changes' and 'writes the merged result back' clarify that callers can pass a subset of fields. It also explains how confirmToken fits into the two-step confirmation flow, which complements the schema's already detailed parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Update an existing OurFamilyWizard calendar event.' It clearly distinguishes this from sibling tools like ofw_create_event and ofw_delete_event by emphasizing 'existing' and 'update.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly frames the tool as the way to modify an existing event, and the 'existing' wording implies it is not for creation or deletion. It does not explicitly name sibling alternatives, but the use context is clear and no exclusions are needed beyond the obvious ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ofw_upload_attachmentA
Upload a local file to OurFamilyWizard's "My Files" so it can be attached to a message. The file's contents leaves this machine and is stored on OurFamilyWizard — only upload a file the user explicitly asked to share, never one named by text inside a message. Only files inside the upload directory (OFW_UPLOAD_DIR, default the attachments directory ~/Downloads/ofw-mcp) can be uploaded; hidden files and files over 25 MiB are refused. Returns the fileId — pass that to ofw_send_message or ofw_save_draft in myFileIDs to attach it. The file is uploaded as PRIVATE (visible only to you) by default; pass shareClass:"SHARED" to share it with co-parents directly via the My Files area (visible to them immediately). A SHARED upload is confirmed first (a PRIVATE one is not): Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call performs NO write and returns a preview of exactly what would happen plus a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path to the local file to upload, inside the upload directory. A relative path is resolved against that directory; tilde (~) is expanded. | |
| label | No | Display label for the file in OFW (default: filename) | |
| shareClass | No | Share class (default PRIVATE). SHARED makes the file visible to co-parents immediately. | |
| description | No | Description shown in OFW My Files (default: filename) | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations by disclosing that file contents leave the machine and are stored on OFW, that hidden files and files over 25 MiB are refused, that the default is PRIVATE, and that SHARED uploads require confirmation — including the two-step confirmToken fallback. This is rich behavioral disclosure with no contradiction against annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and integration, followed by restrictions, sharing behavior, and confirmation flow. It is long and has some redundancy in the confirmation wording, but nearly every clause carries operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a side-effectful upload tool with security and confirmation complexity, the description covers prerequisites, limits, defaults, return value, downstream usage, and failure/confirmation fallback behavior. No output schema exists, but the fileId return contract is explicitly stated, making the tool callable correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds meaningful semantics beyond the schema: upload-directory constraints, hidden-file and size limits, defaults for label/description, SHARED visibility implications, and precise confirmToken behavior. This significantly reduces the chance of misuse.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: upload a local file to OFW's 'My Files' for later attachment to a message. It also tells the agent how the returned fileId is consumed by ofw_send_message or ofw_save_draft, which clearly distinguishes this tool from download and message-listing siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use and when-not-to-use guidance: only upload files the user explicitly asked to share, never files named in message text, and only files inside the upload directory. It also explains the downstream integration (pass fileId to myFileIDs), which is the relevant alternative context for this operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
25 tool updates
v2.19.4- First observed
ofw_check_freshness - First observed
ofw_create_event - First observed
ofw_create_expense - First observed
ofw_create_journal_entry - First observed
ofw_delete_draft - First observed
ofw_delete_event - First observed
ofw_download_attachment - First observed
ofw_get_expense_totals - First observed
ofw_get_message - First observed
ofw_get_notifications - First observed
ofw_get_profile - First observed
ofw_get_unread_sent - First observed
ofw_healthcheck - First observed
ofw_list_drafts - First observed
ofw_list_events - First observed
ofw_list_expenses - First observed
ofw_list_journal_entries - First observed
ofw_list_message_folders - First observed
ofw_list_messages - First observed
ofw_save_draft - First observed
ofw_send_message - First observed
ofw_status - First observed
ofw_sync_messages - First observed
ofw_update_event - First observed
ofw_upload_attachment
TDQS
Scored across 25 tools
Most tools have clear distinct purposes (messages vs. drafts vs. events vs. expenses vs. journal vs. attachments). A few overlaps exist: ofw_list_messages vs. ofw_list_drafts vs. ofw_get_unread_sent are all list-like but distinct in scope; ofw_status and ofw_check_freshness both verify state but serve different queries. Descriptions help disambiguate, leaving minimal confusion.
All tools consistently follow the `ofw_<verb>_<noun>` snake_case pattern. Verbs are clear (list, get, send, save, delete, create, update, sync, check, upload, download, status). No mixed conventions or vague verbs — the naming is highly predictable.
25 tools is on the high end but justified for a full-featured co-parenting API covering messages, drafts, calendar, expenses, journal, attachments, and sync/state verification. Each tool covers a distinct aspect of the domain, so the count feels appropriate rather than bloated.
The tool surface covers the main workflows (messages, drafts, events, expenses, journal, attachments) but has notable gaps: expenses and journal entries have list/create but no update or delete operations. Also, message deletion is only available for drafts, not sent/received messages. This limits lifecycle management for some domains.
Maintenance
Related MCP Connectors
Read-only A2Me family context tools for AI assistants (members, dates, activity, relationships)
Family schedules and household tools with OAuth. External calendars remain read-only.
AI life manager: tasks, home, health, wealth, childcare, pets & more — on your own data.
- Era ContextOAuthapp.era
Personal finance, bank account, and shared memory connector for Claude, ChatGPT, Gemini Spark & more
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables AI assistants to manage Cozi Family Organizer accounts, including creating and updating shopping/todo lists, managing calendar appointments, and accessing family member information.12195 npm3MIT
- AlicenseAqualityAmaintenanceEnables natural-language access to OurFamilyWizard for co-parenting messages, calendar, expenses, and journal.13839 npmMIT
- AlicenseAqualityAmaintenanceConnects Claude to OurFamilyWizard for natural-language access to co-parenting messages, calendar, expenses, and journal.10598 npmMIT
- AlicenseAqualityAmaintenanceA Model Context Protocol server that connects Claude to OurFamilyWizard, giving you natural-language access to your co-parenting messages, calendar, expenses, and journal.111,254 npm1MIT