gmail-multi-mcp
Connect an AI assistant to multiple Google accounts for cross-account Gmail, Drive, Calendar, and Contacts work.
Manage accounts: list configured accounts with scopes, aliases, and reauthorization/probe status.
Gmail: search all or one mailbox; read threads/messages; create drafts; send/reply; manage labels; trash or restore email (no permanent delete).
Drive: search across accounts and shared drives; read/export files; list folders; list shared-drive IDs; upload files; create Google Docs; share files.
Calendar: list, read, create, update, and delete events; find free time across accounts with incomplete-result protection.
Contacts: search, list, read, create, and update contacts across address books, with replacement semantics and etag-guarded writes.
Cross-account defaults: omit account to fan out across every connected account in parallel, with failures surfaced instead of silently dropped.
Allows searching, reading, drafting, sending, and labeling emails across multiple Gmail accounts.
Allows listing events and finding free time across multiple Google Calendars.
Allows searching and accessing files across multiple Google Drive accounts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gmail-multi-mcpSearch all my accounts for emails with invoices sent this week."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
gmail-multi-mcp
An MCP server that connects your AI assistant to several Google accounts at once — Gmail, Drive, Calendar and Contacts.
Why this exists
The official connectors handle one account. If your work lives in you@company.com,
your invoices arrive at billing@company.com and your life happens in you@gmail.com, you
end up switching connectors — or simply cannot ask a question that spans two of them.
This server makes the account a first-class parameter, and then goes one step further:
"Any unread invoices this week, and do I have a free hour to deal with them?"
→ search_emails searches work@, billing@ and personal@ in parallel
→ calendar_find_free_time gaps that are free on ALL THREE calendars
→ one answerOmit the account and it searches every mailbox, every Drive, every calendar and every address book you have connected. One Google Cloud project, as many accounts as you want.
Related MCP server: multi-mail-mcp
Features
Multi-account by design. Add as many Google accounts as you like; each keeps its own credentials.
Cross-account by default.
search_emails,drive_search,calendar_list_events,calendar_find_free_timeandcontacts_searchall work across every connected account when you omitaccount.Shared drives included. Drive returns only My Drive unless a request opts in, so every call here does.
drive_searchlooks across My Drive and every shared drive the account can reach;drive_list,drive_read,drive_uploadanddrive_shareall work on shared-drive items.A shared drive is not a file, so it has its own tool. A drive never shows up in
drive_search— not even searched by its exact name, with full access.drive_shared_drivesis the only way to turn a drive's name into the id every other tool wants.A partial file list says so. Searching every drive lets Drive quietly omit documents. When it admits that, the affected accounts come back in
incompleteAccountsinstead of a count that looks complete and is not.Degrades instead of failing. If one account's token is revoked, the others still return — and the broken one is reported in a
failuresfield rather than silently dropped.Free time across accounts.
calendar_find_free_timeintersects the busy blocks of several calendars. If it cannot read one of them it returns no slots at all rather than a confident answer built on half the picture.Aliases. Call an account
workorpersonalinstead of typing the full address.Standard OAuth, no password ever. Consent happens in your browser, on Google's own screen, through a CSRF-protected loopback callback.
Tokens stay on your machine, in your home directory,
0600on Unix — never in the repository, never in the MCP client configuration.Automatic token refresh, written back to disk so a restart does not re-refresh.
Per-service scope gate. An account authorised before Drive existed keeps doing Gmail and gets a clear
MISSING_SCOPEfor Drive — not a baffling Google 403.Twenty-nine tools across mail, files, calendar and contacts.
Contacts that cannot be clobbered.
contacts_updatere-reads the contact for itsetagbefore writing, so a change made on a phone thirty seconds earlier is not silently overwritten.Cannot permanently delete email.
trash_emailmoves a message to the bin anduntrash_emailbrings it back. There is no permanent-delete tool: it would need a scope this server does not request, and an unrecoverable action does not belong behind an agent.Strict TypeScript, no
any, and a dependency list you can read in one glance.
Prerequisites
Node.js | ≥ 20.12 — see the note below |
Google Cloud | A free account. No billing needed for personal use. |
An MCP client | Claude Code, Claude Desktop, or anything that speaks MCP over stdio. |
Why 20.12 and not 18. The server loads a local
.envwith Node's built-inprocess.loadEnvFile(), which landed in 20.12 (and 21.7). On Node 18 that file would be ignored silently, and you would get a confusing "no credentials" error with a perfectly good.envsitting right there. The floor is real, not cautious.
Installation
git clone https://github.com/blackwings-dev/gmail-multi-mcp.git
cd gmail-multi-mcp
npm install
npm run buildThat produces dist/index.js, which is what your MCP client will run. Check it:
node dist/index.js --versionGoogle Cloud setup
You do this once, no matter how many accounts you connect. About five minutes.
1. Create a project
Open the Google Cloud Console and create a project —
call it whatever you like, e.g. google-mcp.
2. Enable four APIs
With your new project selected, enable each of these. Missing one produces a 403 that says nothing useful, and only for the tools that need it:
API | Direct link |
Gmail API | https://console.cloud.google.com/apis/library/gmail.googleapis.com |
Google Drive API | https://console.cloud.google.com/apis/library/drive.googleapis.com |
Google Calendar API | https://console.cloud.google.com/apis/library/calendar-json.googleapis.com |
People API (contacts) | https://console.cloud.google.com/apis/library/people.googleapis.com |
Contacts live behind the People API, not a "Contacts API". Looking for the latter in the library and finding the long-dead Contacts API v3 is the usual wrong turn.
Press Enable on each. Nothing else on those pages matters.
3. Configure the consent screen
In the sidebar this now lives under Google Auth Platform; older projects show it as APIs & Services → OAuth consent screen.
Field | What to put |
App name | Anything, e.g. |
User support email | Your own address. |
Audience / User type | External |
Developer contact | Your own address again. |
Then, under Audience → Test users, press Add users and add every Google address you intend to connect.
⚠️ This step is not optional. While the app is in Testing, an account that is not listed as a test user cannot grant consent — Google shows an "access blocked" screen that does not explain why.
4. Create the OAuth client
APIs & Services → Credentials → Create Credentials → OAuth client ID
Field | Value |
Application type | Desktop app ← this exact type |
Name | Anything, e.g. |
Copy the Client ID and Client secret from the dialog that appears.
Why "Desktop app" specifically. That client type accepts any loopback port as a redirect URI. The setup flow opens a throwaway server on a random free port and hands that URL to Google, so you never register redirect URIs by hand — which is exactly what lets a single project serve any number of accounts.
Scopes, and what they cost you
This server asks for four scopes:
Scope | What it buys | Google's tier |
| Read, draft, send, label. Not permanent deletion. | Restricted |
| Full access to the user's files. | Restricted |
| Read and write events, and free/busy. | Sensitive |
| Read and write the account’s own contacts. | Sensitive |
Be aware of what full drive costs, because it changes your options later. Google
classes it as a restricted scope, the same tier as gmail.modify. The narrower
drive.file — access limited to files this app created or that the user explicitly picked
in a Google file picker — is not restricted and would keep the verification story
simpler. But drive.file cannot see files you already have, which means drive_search
over your existing Drive would return nothing. That is the whole trade-off: searchable
Drive, or an easier path through Google's review.
Full drive is the deliberate choice here. Switching is a one-line change in
src/auth/oauth.ts (DRIVE_SCOPES) followed by re-authorising your
accounts.
contacts is read and write because contacts_create and contacts_update exist;
contacts.readonly would do if you only ever look. It does not cover Google’s "other
contacts" — the addresses Gmail collects by itself — which sit behind a separate scope this
server does not ask for.
Configuring the server
1. Credentials
cp .env.example .envGOOGLE_CLIENT_ID=1234567890-abcdefg.apps.googleusercontent.com
GOOGLE_CLIENT_SECRET=GOCSPX-your-secret-here.env is git-ignored. You can also skip this file entirely — npm run setup will ask for
the credentials and store them in ~/.gmail-multi-mcp/config.json. If both exist, the
environment wins.
2. Connect your accounts
npm run setup gmail-multi-mcp
Gmail, Drive and Calendar. One Google Cloud project, as many accounts as you need.
config C:\Users\you\.gmail-multi-mcp\config.json
tokens C:\Users\you\.gmail-multi-mcp\tokens
✔ OAuth credentials found in the environment.
What now?
─────────
1 Add a Google account
2 List accounts
3 Remove an account
4 Show MCP client configuration
q QuitChoose 1. A browser opens and you sign in. Google’s consent screen lists Gmail, Drive,
Calendar and Contacts; grant them, and the tab says "Account connected". The CLI then asks for an optional alias —
work, personal, billing — which you can use instead of the full address from then on.
⚠️ Adding a second account
Sign out of Google first, or use a private/incognito window.
Otherwise the browser reuses the session you already have, Google skips the account chooser, and you connect the same account twice without noticing.
Repeat for each account, then check the result with option 2.
3. Register the server with Claude Code
claude mcp add gmail-multi-mcp \
--scope user \
-e GOOGLE_CLIENT_ID=your-client-id.apps.googleusercontent.com \
-e GOOGLE_CLIENT_SECRET=GOCSPX-your-secret \
-- node /absolute/path/to/gmail-multi-mcp/dist/index.jsOn Windows the path looks like D:\Dev\gmail-multi-mcp\dist\index.js. Verify with
claude mcp list, then restart Claude Code so it picks the server up.
{
"mcpServers": {
"gmail-multi-mcp": {
"command": "node",
"args": ["/absolute/path/to/gmail-multi-mcp/dist/index.js"],
"env": {
"GOOGLE_CLIENT_ID": "your-client-id.apps.googleusercontent.com",
"GOOGLE_CLIENT_SECRET": "GOCSPX-your-secret"
}
}
}
}The env block is optional if the credentials are already in
~/.gmail-multi-mcp/config.json. Account tokens are always read from there — they never go
in the client configuration.
Option 4 in the setup CLI prints this snippet with the correct absolute path already filled in.
⚠️ Upgrading from a Gmail-only version? Re-authorise
If you connected your accounts when this server only did Gmail, they will keep working for Gmail and fail for everything else. Drive and Calendar tools will answer:
MISSING_SCOPE — <account> is authorised, but not for Drive.A refresh will not fix this. Google binds a refresh token to the exact set of scopes granted when it was issued, so no amount of restarting, re-refreshing or waiting will add Drive and Calendar. Only a fresh consent will.
The fix takes a minute per account:
npm run setupThe CLI tells you which accounts are affected before you do anything:
! 3 accounts need re-authorising:
· you@company.com
· billing@company.com
· you@gmail.com
They were connected before Drive and Calendar support existed, so they can only
do Gmail right now. A refresh cannot add the missing scopes: Google binds a
refresh token to the scopes it was issued with. Choose 1 and add each account
again — it re-consents and overwrites in place, so nothing has to be removed.Choose 1 and add each account again. You do not have to remove them first: the flow always asks for fresh consent and overwrites the stored token in place, keeping the alias.
Your assistant can check this for itself at any time — list_accounts reports
missingScopes and needsReauthorization per account, and a warning summarising them.
And do not forget step 2 of the Google Cloud setup: an account with the right scopes still fails if the Drive or Calendar API is not enabled on the project.
Usage
Ask in natural language. A few that exercise different corners:
What you say | What happens |
"Any unread invoices this week?" |
|
"Bin that newsletter" |
|
"Find the Q3 budget spreadsheet in any of my Drives" |
|
"Read me that spreadsheet" |
|
"Draft a reply to Ana with the numbers from it" |
|
"Write those notes up as a Google Doc in my work account" |
|
"Share it with marta@client.com as a commenter" |
|
"What's on my calendars tomorrow?" |
|
"Find me two free hours next week across all my accounts" |
|
"Book it as 'Budget review' and invite Marta" |
|
"Move Thursday's stand-up an hour later" |
|
"What's Marta's phone number?" |
|
"Add her to my work contacts with that email" |
|
Habits worth having:
Ask for a draft unless you really want the mail sent.
replyandsend_messagego out immediately and there is no undo.Name the account ("in my work account") when you mean one; leave it out when you want all of them.
calendar_delete_eventanddrive_shareare the two that change things you cannot take back from here. Be explicit with them.
The tools
account accepts an email address or an alias. Where it is optional, omitting it
means "every connected account, merged".
Accounts
Tool | Description | Parameters |
| Configured accounts, their scopes, and which need re-authorising |
|
Gmail
Tool | Description | Parameters |
| Gmail query syntax. Without |
|
| A full conversation, every message in order |
|
| One message: decoded body plus attachment metadata |
|
| Saves a draft without sending |
|
| Replies in-thread — handles |
|
| Composes and sends a new email. No undo. |
|
| Labels of an account, with their ids |
|
| Adds and/or removes labels; accepts names or ids |
|
| Moves a message to the bin. Reversible, not a permanent delete |
|
| Takes a message back out of the bin, restoring its labels |
|
Useful system labels for label_message: UNREAD (remove it to mark as read), STARRED,
IMPORTANT, SPAM.
On the bin. trash_email is the tool for it; adding the TRASH label by hand does the
same thing less clearly. Neither is a permanent delete — but "not permanent" deserves its
footnote: Gmail empties the bin by itself after about 30 days, so a trashed message is
recoverable for a month and then it is gone. search_emails with in:trash shows what is
still in there. There is no permanent-delete tool at all: users.messages.delete needs
the full-mailbox scope this server deliberately does not request.
Attachment metadata is returned (filename, MIME type, size); attachment contents are not downloaded.
Drive
Tool | Description | Parameters |
| Without |
|
| Reads a file as text |
|
| Lists a folder, sub-folders first |
|
| Lists shared drives with their ids. A drive is not a file, so it never appears in |
|
| Uploads a file. Exactly one of |
|
| Creates a real Google Doc from plain text |
|
| Grants access. Gives away real data. |
|
drive_read conversions: Google Docs → text/plain, Sheets → text/csv, Slides →
text/plain. Ordinary text files are downloaded as they are. Binary files (PDF, images,
archives) are refused with a link instead — this server does not download them. Content
over max_chars is cut and says so; files over 10 MB are refused.
drive_share types: user and group need email_address; domain needs domain;
anyone makes the file readable by anyone with the link. Roles are reader,
commenter or writer — ownership transfer is deliberately not offered. No notification
email is sent unless notify is true.
Contacts
Contacts are handled through Google’s People API.
Tool | Description | Parameters |
| Without |
|
| One contact in full: every email and phone with its label, organisation, addresses, links, notes |
|
| Pages through an address book, most recently changed first |
|
| Adds a person. At least one field required |
|
| Changes a person. Each named field is replaced, not merged |
|
resource_name looks like people/c1234567890 and comes from contacts_search or
contacts_list.
⚠️ contacts_update replaces whole fields. Sending emails: ["new@x.com"] leaves the
contact with exactly that one address and deletes the rest. Fields you do not name are left
alone — so read the contact first and send back the complete list of whatever you are
editing. The current version is re-read for its etag before writing, so an edit made on a
phone in the meantime cannot be silently overwritten.
The same person in two accounts is returned twice, on purpose. Each copy has its own
resource_name in its own address book; merging them would produce an id that edits the
wrong one.
Paging: contacts_list returns nextPageToken when there is more. Pass it back as
page_token.
Calendar
Tool | Description | Parameters |
| Without |
|
| One event in full, with attendees and their responses |
|
| Creates an event |
|
| Changes an event. |
|
| Deletes an event. Irreversible. |
|
| Gaps free on every account asked about |
|
Times must carry an explicit offset — 2026-09-10T09:00:00+02:00 or a Z. A bare
2026-09-10T09:00:00 is refused, on purpose: it would be read in whatever zone the
server happens to run in, and the resulting answer looks completely normal while being
hours wrong.
All-day events use plain dates (YYYY-MM-DD) with all_day: true. Google treats the
end date as exclusive, so a one-day event ends on the following day.
send_updates defaults to none — attendees are added to the event but not emailed
unless you ask. Values: all, externalOnly, none.
The calendar_find_free_time contract
This is the one tool where a wrong answer looks exactly like a right one, so its rules are worth stating:
time_min/time_maxmust carry an explicit offset. Maximum range: 62 days.Working hours (
workday_start,workday_end, default09:00–18:00) are interpreted intimezone, an IANA name likeEurope/Madrid. It defaults to the zone the server runs in, and is always echoed back in the result so you can see what was assumed. Daylight saving is handled; the offset is resolved per day, not fixed once.weekdays_onlydefaults to true.A busy block that overlaps a candidate window only partially still removes the overlapping part. Half a busy hour is not free.
Returned slots are the free segments of at least
minimum_minutes(default 30), in UTC, not a grid of fixed start times.If any calendar cannot be read, no slots are returned at all. The result carries
incomplete: true, thefailures, and an explanation. A gap computed without one person's calendar is a double-booking, not an answer.
Troubleshooting
MISSING_SCOPE — <account> is authorised, but not for Drive
The account was connected before Drive and Calendar support existed. See Re-authorise above. A refresh cannot fix it; only a fresh consent can.
Access denied by Drive: Request had insufficient authentication scopes
The scope gate let this through, so the stored token claims the scope but Google disagrees — usually a half-finished re-authorisation. Add the account again.
Access denied by Drive / by Calendar / by Contacts (HTTP 403)
The API is not enabled on the Google Cloud project. Enabling one does not enable the others — see step 2. For contacts the one you need is the People API.
contacts_search finds nothing, but the person is definitely there
Google’s contact search reads a server-side index that has to be warmed with an empty query before it answers anything. This server does that on the first search per account and retries once after a short pause, so you should not meet it — but if a brand new search comes back empty, ask again. A second attempt against a warm cache is the difference between "no such person" and "not ready yet".
A contact lost its other email addresses after an update
contacts_update replaces each field it is given. Passing one address removes the
others. Read the contact with contacts_get first and send the full list back.
HTTP ERROR 431 (Request Header Fields Too Large) on the callback
Fixed in current versions — if you see it, you are on an old build. The cause is worth
knowing: the redirect used to go to localhost, which is a shared cookie origin. Every
dev server you have ever run on any localhost port can leave cookies there, and the
browser sends all of them to the OAuth callback. Past roughly 16 KB that exceeds Node's
default header limit — and it fails after Google has already granted consent, so the
authorisation code is lost.
The redirect now targets 127.0.0.1, which does not receive localhost cookies, and the
callback server accepts 64 KB of headers. If it somehow still happens, clear cookies for
127.0.0.1 and retry.
Accounts stop working after 7 days
Your OAuth consent screen is still in Testing, and Google expires refresh tokens for apps in that state after seven days. Two options, neither free:
Stay in Testing and re-run
npm run setuponce a week.Publish the app (Google Auth Platform → Audience → Publish app). Refresh tokens stop expiring. In exchange, an app that has not been through Google's review shows an "unverified app" warning screen on every consent and is capped at a limited number of users.
Two of the four scopes here are restricted (gmail.modify, drive), the category
Google reviews most closely, and that review policy changes. Check the current requirements
on the consent screen itself before assuming publishing will be waved through.
Unknown account "..."
The reference did not match any configured email or alias. Run npm run setup and choose
2, or ask your assistant to call list_accounts.
Google did not return a refresh token
Google issues one only on first consent. Revoke the app at myaccount.google.com/permissions, then authorise again.
time_min must be ISO 8601 with an explicit UTC offset
Working as intended. Send 2026-09-10T09:00:00+02:00, not 2026-09-10T09:00:00.
You connected the same account twice
The browser reused an existing Google session. Remove the duplicate with option 3, then add the other account from a private window.
Your MCP client shows a JSON parse error and disconnects
Something wrote to stdout. On the stdio transport, stdout is the protocol. If you are
modifying this server, every diagnostic must go to stderr — never console.log.
A cross-account search returns fewer results than expected
Look at the failures field of the response. An account whose token was revoked is
reported there rather than silently skipped.
Security
Tokens never leave your machine. They live in
~/.gmail-multi-mcp/tokens/, one file per account,chmod 0600on Unix. Nothing is sent anywhere except Google.Nothing sensitive is in the repository.
.gitignorecovers.envand every.env.*except the example.CSRF-protected callback. A random
stateis generated per flow and compared in constant time; a callback that does not match is rejected.Loopback only. The callback server binds
127.0.0.1, on an ephemeral port, and shuts down as soon as the code arrives.Header injection blocked. CR/LF is stripped from every outgoing mail header, so a crafted subject or recipient cannot inject extra headers.
No permanent mail deletion. The Gmail scope cannot do it and no tool offers it.
About the client secret: for Desktop app clients Google does not treat it as confidential — an installed application cannot keep a secret, which is why this flow is designed not to depend on one. Keep it out of your repository anyway.
Removing an account deletes the local token but does not revoke Google's grant. Revoke it fully at myaccount.google.com/permissions.
⚠️ What you are handing an assistant, and the risk that comes with it
This server can read untrusted text (any email anyone sends you, any document shared with you) and, in the same session, read local files, upload them to Drive, and share them publicly. Those two capabilities together are an exfiltration path if the model acts on instructions it finds inside the content it is reading.
Nothing in this server can prevent that, because from the API's point of view a malicious instruction in an email body and a genuine request from you look identical. What it does instead is make the dangerous steps loud and deliberate:
drive_shareis annotated as destructive, defaults to sending no notification, refuses ownership transfer, and requirestype: "anyone"spelled out to create a public link.drive_uploadwill not read a local file unless given an explicitlocal_path.calendar_delete_eventis annotated as destructive and reads the event first so the answer says what it removed.
Use an MCP client that asks before running write tools, and read what it is about to do. Treat "an email told me to share this file" as the red flag it is.
How it works
MCP client ──stdio──▶ src/index.ts ──▶ src/server.ts 29 tools, zod-validated
│
┌──────────────┬────────────────────┼────────────────────┬──────────────┐
▼ ▼ ▼ ▼
src/gmail/ src/drive/ src/calendar/ src/contacts/
search, MIME queries, exports, events, free/busy People API,
parsing and uploads and and time-zone search warm-up,
building sharing arithmetic etag-guarded writes
└──────────────┴────────────────────┼────────────────────┴──────────────┘
▼
src/core/
AccountClientCache · the SCOPE GATE · error taxonomy
cross-account runner · bounded concurrency
│
▼
src/auth/oauth.ts
loopback consent, automatic refresh
│
▼
src/auth/token-store.ts
~/.gmail-multi-mcp/Path | Contents |
| OAuth client (if not in the environment) and the account list |
| One refresh token per account, with the scopes it was granted |
The scope gate lives in one place — AccountClientCache in
src/core/google-client.ts. Every Drive and Calendar call
has to go through a client, and every client comes from that class, so an account missing a
scope is refused by construction rather than by each of twenty-eight handlers remembering to
ask. A check spread across handlers has a blind spot the moment someone adds one more.
Auto-refresh is handled by Google's auth client: an expired access token is renewed
transparently, and a tokens listener writes the new one back to disk — preserving the
refresh token, which Google omits from refresh responses.
Development
npm install
npm run build # tsc → dist/
npm run typecheck # no emit
npm run dev # run the server from source with tsx
npm run setup # run the setup CLI from sourceSmoke-test the server without an MCP client:
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"t","version":"1"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}' \
| node dist/index.jsLayout
src/
index.ts entry point: server by default, `setup` for the CLI
server.ts MCP server and all twenty-eight tool definitions
core/
errors.ts the error taxonomy and Google's HTTP failures mapped onto it
google-client.ts per-account client cache AND the scope gate
accounts.ts run one operation across every account, failures as values
concurrency.ts bounded parallelism
auth/
oauth.ts scopes, consent flow, loopback callback, automatic refresh
token-store.ts config.json and per-account token files
gmail/
client.ts authenticated clients, MIME parsing and building
search.ts single- and multi-account search
types.ts Gmail domain types
drive/
client.ts queries, exports, uploads, sharing
types.ts Drive domain types
calendar/
client.ts events and free/busy arithmetic
timezone.ts wall-clock time in a named zone, and interval maths
types.ts Calendar domain types
contacts/
client.ts People API: search warm-up, etag-guarded writes
types.ts Contacts domain types
cli/
setup.ts interactive setup (chalk)
assets/ the project icon: PNG and SVG, with and without wordmarkThe SVG in assets/ is the master. Its geometry is also inlined in
src/auth/icon.ts as the favicon of the local OAuth callback page — the only web
page this server ever serves.
TypeScript runs with strict plus noUncheckedIndexedAccess, noImplicitReturns,
noUnusedLocals and verbatimModuleSyntax. There is no any in the source.
A note on the name
The project is still called gmail-multi-mcp, and that name is now too small for it.
Gmail is one of three services. google-workspace-mcp or google-multi-mcp would describe
it better.
The rename is deliberately not done yet, because it is wider than it looks. Its surface includes:
the package name, the binary name and the repository;
~/.gmail-multi-mcp/— the config and token directory, which means existing users would have to move their credentials or the server would look freshly installed;GmailMcpError, the error class every layer throws, including Drive and Calendar;SERVER_NAME, and the literal string"gmail-multi-mcp setup"in a dozen user-facing error messages.
Worth doing in one deliberate pass, with a migration for the token directory — not piecemeal.
Contributing
Issues and pull requests are welcome.
Fork the repository and create a branch:
git checkout -b feature/what-it-does.Make your change.
npm run typecheckandnpm run buildmust both pass.Smoke-test the server as shown above.
Open a pull request describing what changed, why, and how you verified it.
Please keep the five rules that hold the design together:
Nothing writes to stdout outside
src/cli/and the--help/--versionpaths. Stdout is the MCP transport.No
any. Google's types are noisy; convert them at the boundary in each service'sclient.tsand keep the rest of the codebase on the domain types.New service? Go through
AccountClientCache. That is where the scope gate lives. A client built any other way skips it.Cross-account operations report their failures. Never return a shorter list and stay quiet about the account that did not answer.
Nothing unrecoverable.
trash_emailis reversible; a permanent-delete counterpart is not, and does not belong behind an agent. The Gmail scope this server requests cannot do it, and that is the point — do not widen the scope to add one.
Good first contributions: attachment download, Google Docs formatting via the Docs API,
recurring-event rules, contact groups and labels, Google’s "other contacts", mark_read /
archive convenience tools, or shared-drive support.
Reporting a security issue
Please do not open a public issue. Contact the maintainers privately first.
License
MIT — do what you like, no warranty.
Who maintains this
Built and maintained by Black Wings — software, AI and managed IT for companies, from Valencia, Spain. We wrote this because we needed it: we run several Google Workspace mailboxes and no connector could see more than one at a time.
If it saves you the same afternoon it saved us, that is the whole point. Other things we have released are listed at blackwings.dev/codigo-abierto (in English: blackwings.dev/en/open-source).
Available Tools
29 toolscalendar_create_eventCreate a calendar eventA
Create an event. For a timed event, "start" and "end" are ISO 8601 WITH an offset. For an all-day event set "all_day" and give plain dates (YYYY-MM-DD) — note that Google treats the end date as EXCLUSIVE, so a one-day event ends on the following day. Listing attendees puts the event on their calendars; they are only emailed when "send_updates" says so.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | End. ISO 8601 with offset, or YYYY-MM-DD if all_day. | |
| start | Yes | Start. ISO 8601 with offset, or YYYY-MM-DD if all_day. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| all_day | No | Treat start and end as plain dates. | |
| summary | Yes | Event title. | |
| location | No | Where it happens. | |
| attendees | No | Attendee email addresses. | |
| time_zone | No | IANA zone stored with the event, e.g. "Europe/Madrid". | |
| calendar_id | No | Calendar id. Defaults to "primary", the account own calendar. | |
| description | No | Body of the event. | |
| send_updates | No | Whether Google emails the attendees. Defaults to "none": nobody is mailed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate a state-changing operation (readOnlyHint false). The description goes beyond the annotations by disclosing that adding attendees places the event on their calendars, and that emails are only sent when send_updates is set. It also explains an important API quirk (exclusive end date for all-day events). This adds behavioral context that the annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with zero fluff. It front-loads the primary action ('Create an event') and then packs the critical nuances (date formats, exclusive end, attendee/email behavior) into compact, useful clauses. Each sentence earns its place; no redundant or vague phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description covers the key gotchas: timed vs. all-day handling, date-exclusivity, attendee calendar placement, and send_updates default. It does not mention the return value, but that is acceptable given no output schema. It omits no critical information that an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter already has a description. The description adds value by clarifying the difference between timed ISO 8601 offsets and plain dates for all-day events, and by explaining the exclusive end-date behavior. It also notes the side effect of attendee emailing tied to send_updates. This goes beyond the schema's literal param descriptions, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific action 'Create an event,' which is a clear verb + resource. It further distinguishes the tool's purpose by detailing the two valid input formats (timed vs. all-day), which are unique to creating an event. This makes it unambiguous against sibling tools like calendar_update_event or calendar_delete_event.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: when to use timed vs. all-day formats, how to set attendees, and the effect of send_updates. It does not explicitly name alternatives (e.g., use calendar_update_event to modify) but the create/update distinction is implied by the verb. It provides situational guidance without excluding alternatives, which earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_delete_eventDelete a calendar eventADestructiveIdempotent
Delete an event. This is IRREVERSIBLE from here — there is no undo and no bin. The event is read before deletion so the answer says what was removed. Only do this when the user has clearly asked for it.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| event_id | Yes | Event to delete. | |
| calendar_id | No | Calendar id. Defaults to "primary", the account own calendar. | |
| send_updates | No | Whether Google emails the attendees. Defaults to "none": nobody is mailed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing that deletion is IRREVERSIBLE with no undo or bin, and that the event is read before deletion so the answer indicates what was removed. The destructiveHint annotation is reinforced with concrete behavioral consequences, adding real value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. The critical irreversibility warning is front-loaded, followed by the pre-read behavior and the invocation condition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with 4 parameters and no output schema, the description covers the essential behavioral context: irreversibility, what the response will indicate, and the condition for calling. Combined with the annotations and fully documented schema, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all four parameters. The description does not add parameter-specific detail beyond the schema, which is acceptable and earns the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Delete an event.' The action is unambiguous and distinct from sibling tools like calendar_update_event or calendar_get_event. The first sentence directly matches the title, leaving no doubt about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: 'Only do this when the user has clearly asked for it.' This is an explicit condition for invocation, especially important for a destructive action. It does not name alternatives or when-not-to-use scenarios, but the guidance is sufficient and clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_find_free_timeFind free timeARead-onlyIdempotent
Find gaps that are free on EVERY account asked about, within working hours. "time_min" and "time_max" must carry an explicit UTC offset. Working hours are interpreted in "timezone" (IANA), which defaults to the zone this server runs in and is always echoed back. IMPORTANT: if any calendar cannot be read, NO slots are returned and "incomplete" is set — a gap computed without one person's calendar is a double-booking, not an answer.
| Name | Required | Description | Default |
|---|---|---|---|
| accounts | No | Restrict to these accounts. Omit to use every configured account. | |
| time_max | Yes | Latest instant to consider. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00". | |
| time_min | Yes | Earliest instant to consider. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00". | |
| timezone | No | IANA zone for the working-hours bounds, e.g. "Europe/Madrid". | |
| calendar_id | No | Calendar id. Defaults to "primary", the account own calendar. | |
| workday_end | No | End of the working day, "HH:MM". Default 18:00. | |
| weekdays_only | No | Skip Saturdays and Sundays. Defaults to true. | |
| workday_start | No | Start of the working day, "HH:MM". Default 09:00. | |
| minimum_minutes | No | Ignore gaps shorter than this. Default 30. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, but the description adds crucial behavior beyond them: if any calendar is unreadable, NO slots are returned and 'incomplete' is set, preventing double-bookings. It also discloses the timezone default-and-echo behavior, which is valuable operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler: the first states the core function, the second gives hard input constraints, and the third highlights a critical safety behavior. The IMPORTANT warning is front-loaded enough to prevent misuse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers core behavior, defaults, working-hours constraints, and failure semantics, while parameter details are fully present in the schema. The only meaningful omission is a precise description of the successful output structure beyond 'gaps' and 'incomplete,' but the annotation suite and schema make this a fairly complete definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by emphasizing that time_min/time_max must carry an explicit UTC offset and by revealing the timezone default (server zone) plus the echo behavior, which the schema does not specify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource combination: find gaps that are free on EVERY account within working hours. This clearly differentiates the tool from event-listing siblings like calendar_list_events or calendar_get_event by emphasizing cross-account availability search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: when you need common free time across multiple accounts within working hours. It does not explicitly name an alternative or state a when-not-to-use condition, but the context is strong enough for an agent to select it over the event-centric calendar siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_get_eventRead a calendar eventARead-onlyIdempotent
Read one event in full: description, location, attendees and their responses, and the link to open it.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| event_id | Yes | Event id, as returned by calendar_list_events. | |
| calendar_id | No | Calendar id. Defaults to "primary", the account own calendar. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, and non-destructive behavior. The description adds useful return-content context by specifying exactly what is included in a full event read, which is valuable because no output schema is provided. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One crisp sentence that front-loads the action and resource, then enumerates the relevant return contents without any filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation with annotations covering safety and idempotency, the description and schema together are sufficient. The only minor gap is that the exact response shape is not fully specified, but the description's enumeration of returned fields largely compensates for the absent output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with sensible descriptions for account, event_id, and calendar_id including the default for calendar_id. The tool description adds no parameter-specific details, but the baseline of 3 is appropriate since the schema already carries that burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and resource ('one event') and clearly defines scope: description, location, attendees and their responses, and a link. This distinguishes it from calendar_list_events and other calendar operations without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the correct use case: retrieving the full content of a single event rather than a summary list or a mutation. It does not explicitly name alternatives like calendar_list_events, but the 'one event in full' framing gives clear contextual guidance with no exclusions needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_eventsList calendar eventsARead-onlyIdempotent
List events in a time range. Omit "account" to look at EVERY connected calendar at once, merged and sorted by start time. Repeating events are expanded into their real occurrences. Times must carry an explicit UTC offset.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Free text over summary, description, location and attendees. | |
| account | No | Restrict to one account. Omit to use every configured account at once. | |
| time_max | Yes | End of the range. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00". | |
| time_min | Yes | Start of the range. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00". | |
| calendar_id | No | Calendar id. Defaults to "primary", the account own calendar. | |
| max_results | No | Maximum events to return per account. Default 25. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this read-only, idempotent, and non-destructive. The description adds valuable behavioral context beyond those hints: recurrence expansion, cross-account merging and sorting, and the explicit UTC-offset requirement for timestamps. It does not contradict any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core action, the cross-account behavior, and the two critical gotchas (recurrence expansion and UTC offsets). The most important behavioral details are front-loaded and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with 100% schema coverage and helpful annotations, the description covers the essential operational nuances an agent needs: time-range usage, account omission behavior, recurrence handling, and timestamp format requirements. The absence of an output schema is mitigated by the clear 'list events' framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description mostly reinforces what the schema already says about 'account' and time parameters, though it adds the merged-and-sorted behavior and the recurrence-expansion detail. This is useful but does not substantially enhance parameter-level meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List events in a time range,' with concrete behavioral details like merging all calendars, sorting by start time, and expanding repeating events. This clearly distinguishes it from sibling tools such as calendar_get_event, which targets a single event, and calendar_find_free_time, which searches availability.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear contextual usage guidance: it explains how to query across every connected calendar by omitting 'account' and emphasizes the required time-range format with UTC offsets. It does not explicitly name alternatives or exclusion conditions, so it stops short of a full when-to-use versus when-not-to-use guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_update_eventUpdate a calendar eventAIdempotent
Change an existing event. Only the fields you give are touched — except "start" and "end", which must be given together or not at all, because moving one end alone reshapes the meeting. Passing "attendees" REPLACES the guest list rather than adding to it.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | New end. Must be given with "start". | |
| start | No | New start. Must be given with "end". | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| all_day | No | Treat the new start and end as plain dates. | |
| summary | No | New title. | |
| event_id | Yes | Event to change. | |
| location | No | New location. | |
| attendees | No | Replacement guest list, in full. | |
| time_zone | No | IANA zone stored with the event. | |
| calendar_id | No | Calendar id. Defaults to "primary", the account own calendar. | |
| description | No | New body. | |
| send_updates | No | Whether Google emails the attendees. Defaults to "none": nobody is mailed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation set says the tool is not read-only and is idempotent, but the description goes beyond annotations by disclosing a crucial behavioral nuance: the update is partial, and passing attendees replaces the full guest list. It also explains why start and end must be paired. This adds real value beyond the structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is three sentences with zero filler. The most important fact ('only the fields you give are touched') is front-loaded, and the start/end pairing and attendee replacement exceptions are stated precisely without redundant elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 12 parameters, the description covers the critical behaviors that could cause incorrect parameter usage: partial updates, paired temporal fields, and attendee list replacement. The remaining parameters are adequately documented by the complete schema, and no output schema is expected for an update operation. It is comprehensive enough without being verbose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 12 parameters with 100% coverage, so the baseline is 3. The description adds cross-parameter meaning: the partial-update behavior for all fields and the attendee replacement semantics. While some of this is echoed in the schema, the description frames it as tool-level behavior, which is useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Begins with a specific verb-plus-resource statement ('Change an existing event'), clearly distinguishing this from calendar_create_event and calendar_delete_event. It also adds partial-update semantics ('Only the fields you give are touched'), which sharpens what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes the context for use — modifying an existing event — and adds important constraints such as start/end pairing and attendee replacement. It does not explicitly name sibling alternatives or state when not to use it, but the usage context is concrete enough for an agent to select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_createCreate a contactA
Add a person to the account address book. At least one field is required. Google does not check for duplicates — search first if the person might already be there.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Free-text note stored on the contact. | |
| emails | No | Email addresses, in full. On update this REPLACES the existing list. | |
| phones | No | Phone numbers, in full. On update this REPLACES the existing list. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| job_title | No | Role within that organisation. | |
| given_name | No | First name. | |
| family_name | No | Surname. | |
| organization | No | Company or organisation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal a write operation (readOnlyHint=false) and non-idempotency (idempotentHint=false). The description adds concrete behavioral context beyond that: Google performs no duplicate check, so repeated calls create duplicate contacts, and 'at least one field is required' states an input constraint. This enriches what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse sentences, each earning its place: the core action, the required-input constraint, and the duplicate warning. The most important information is front-loaded and there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter write tool with no output schema, the description covers the key operational risks (duplicates, minimum fields) and the schema documents every parameter with 100% coverage. The only minor gap is that it does not address account context or the return value, but for a create action with no output schema this is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds one genuinely new semantic rule not derivable from the schema: despite only 'account' being formally required, at least one additional contact field must be supplied. This cross-parameter constraint adds value beyond the JSON Schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Add a person to the account address book.' It also scopes the tool to the 'account' address book, which combined with the sibling names (contacts_search, contacts_get, contacts_list, contacts_update) clearly distinguishes creation from the other contact operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance: 'Google does not check for duplicates — search first if the person might already be there.' This tells the agent to consult search before creating, which is clear when-to-use context. It does not name the sibling tool (contacts_search) explicitly, but the directive is actionable and sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_getRead a contactARead-onlyIdempotent
Read one contact in full: every email and phone with its label, organisation, job title, postal addresses, links and notes.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| resource_name | Yes | Contact id as returned by contacts_search or contacts_list, e.g. "people/c123456". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds value by specifying the return contents (every email, phone, label, organisation, job title, postal addresses, links, notes), which is behaviorally relevant and not repeated in annotations. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly written sentence. It is front-loaded with the core purpose ('Read one contact in full') and then efficiently lists the return fields. There is no filler, and every phrase contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no output schema), the description covers what is needed: it explains the scope ('in full') and the specific fields returned. It does not list error conditions or explicitly state the absence of side effects, but annotations already cover safety, and the field list is adequate for an agent to understand the returned data. Overall, it is sufficiently complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides full descriptions for both parameters (account and resource_name), and the schema coverage is 100%. The description does not add additional parameter-level meaning beyond what the schema already states (e.g., it does not clarify the 'full' return structure or the role of each parameter). Per the baseline rule for high schema coverage, a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action explicitly ('Read one contact in full') and enumerates the specific data fields returned (emails, phones, labels, organisation, job title, addresses, links, notes). This clearly distinguishes it from sibling tools like contacts_search (which searches) and contacts_list (which lists), making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: for a single known contact, use this tool. However, it does not explicitly state when to prefer it over contacts_search or contacts_list, nor does it mention any exclusions or prerequisites (e.g., that a resource_name is required). The guidance is implicit rather than explicit, so it falls short of a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_listList contactsARead-onlyIdempotent
Page through an account address book, most recently changed first. Returns "nextPageToken" when there is more; pass it back as "page_token" for the next page. Use contacts_search when you are looking for someone in particular.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| page_size | No | Contacts per page. Default 50, maximum 200. | |
| page_token | No | The "nextPageToken" from the previous call. Omit for the first page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false; the description adds behavioral detail beyond that: ordering by most recently changed and the token-based pagination contract. It does not cover rate limits or auth, but those are not essential given the annotations and the account parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core behavior, then token mechanics, then the sibling alternative. No filler or redundant clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only paged list tool, the description plus high-coverage schema and annotations is enough to invoke correctly. It covers ordering, pagination, and the search alternative; the absent output schema is not a blocker because the return shape is not needed to make the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so account, page_size, and page_token are already documented. The description mainly reinforces the page_token round-trip, which is already in the schema, so it adds little parameter meaning beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific action (paging through the address book) and resource (account address book), and adds the sort order. It also names contacts_search as the sibling for targeted lookup, so an agent can differentiate the tools without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to prefer contacts_search ('looking for someone in particular'), implying contacts_list is for browsing or enumerating contacts. It also frames the pagination workflow, making clear this is the paging tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_searchSearch contactsARead-onlyIdempotent
Find people in the account address book. The query matches names, nicknames, email addresses, phone numbers and organisations. Omit "account" to search EVERY connected address book at once. The same person in two accounts is returned TWICE on purpose — each copy has its own id and belongs to a different address book, so merging them would make contacts_update edit the wrong one.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Name, email, phone number or company. | |
| account | No | Restrict to one account. Omit to use every configured account at once. | |
| max_results | No | Maximum contacts to return (per account when searching all). Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the read-only/idempotent annotations, the description discloses that the same person can be returned twice across accounts, that each copy has its own id and address book, and that merging would break a later contacts_update. This is exactly the kind of non-obvious behavior an agent needs to safely use the results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the first states the purpose and matching behavior, the second explains the account-scoping option, and the third delivers the critical duplicate-result caveat. Each sentence earns its place and the most important usage detail is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read-only search with rich annotations and fully documented parameters, the description covers the core scenarios, the all-accounts exception, and the dangerous edge case around duplicates. No output schema is present, but the description provides enough return-related context (duplicate copies with distinct ids) to avoid misuse.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all three parameters with descriptions, so the baseline is 3. The description adds meaningful semantic detail: query covers nicknames and organisations, omitting account triggers a global search, and max results are per account when searching all. This exceeds the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Find people in the account address book') and enumerates the fields the query matches (names, nicknames, email addresses, phone numbers and organisations), so an agent can tell this is a query-based search rather than a plain list or fetch. It doesn't explicitly compare against sibling tools such as contacts_list or contacts_get, so it stops short of a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete scoping guidance: omit 'account' to search every connected address book at once, and it warns that duplicates across accounts are intentional and that merging them would make contacts_update edit the wrong record. It does not state explicit when-not-to-use rules or name alternatives, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_updateUpdate a contactAIdempotent
Change an existing contact. ⚠️ Each field you name is REPLACED, not merged: passing "emails" with one address removes every other address that contact had. Fields you do not name are left alone. Read the contact first with contacts_get and send back the full list of whichever field you are editing. The current version is re-read before writing, so a change made elsewhere in the meantime cannot be silently overwritten.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Free-text note stored on the contact. | |
| emails | No | Email addresses, in full. On update this REPLACES the existing list. | |
| phones | No | Phone numbers, in full. On update this REPLACES the existing list. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| job_title | No | Role within that organisation. | |
| given_name | No | First name. | |
| family_name | No | Surname. | |
| organization | No | Company or organisation. | |
| resource_name | Yes | Contact id as returned by contacts_search or contacts_list, e.g. "people/c123456". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=true, destructiveHint=false. The description goes far beyond these: it warns that named fields are REPLACED not merged, advises reading the contact first, and explains the optimistic concurrency mechanism. No contradiction with annotations; it adds critical behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-organized paragraph, front-loading the critical warning about replacement semantics. Each sentence adds essential value, and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an update tool with replace semantics, the description covers all key aspects: the replacement warning, the need to read first, the concurrent write protection, and the handling of omitted fields. It does not need to explain the return format (no output schema) and the sibling tools are clear enough. The description is complete for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for all parameters, but the description adds crucial semantics beyond the schema, particularly the replacement behavior for emails and phones and the 'full list' requirement. This clarifies usage of parameters in a way the schema alone does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Change an existing contact' with a specific verb and resource, clearly distinguishing it from siblings like contacts_create and contacts_get. It also mentions the field-level replacement semantics, which differentiates it from a generic update.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs the agent to read the contact first via contacts_get and send back the full list of the field being edited, and warns against the merge pitfall. It also notes the re-read before writing, which guides safe concurrent usage. This is explicit when/how guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_draftCreate a draftA
Save a draft without sending it. Use this whenever the user has not explicitly asked for the message to go out.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | Carbon-copy recipients. | |
| to | Yes | Recipients. Plain addresses or "Name <a@b.c>". | |
| bcc | No | Blind carbon-copy recipients. | |
| body | Yes | Message body. Plain text unless is_html is true. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| is_html | No | Send the body as text/html instead of text/plain. | |
| subject | Yes | Subject line. Non-ASCII is encoded automatically. | |
| thread_id | No | Attach to an existing thread. Prefer the "reply" tool for answering a message. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description clarifies that the draft is not sent, which is a key behavioral trait, and it complements the openWorldHint by implying side effects (draft creation). However, it does not elaborate on other potential side effects (e.g., threading behavior via thread_id) or what the draft persistence entails, leaving some uncertainty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exactly two sentences, front-loading the core purpose and usage condition. There is no redundant or filler content; every word earns its place. It is ideal for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters (4 required) and no output schema, the description is minimal but adequate for the core action. It does not mention return behavior, error handling, or prerequisites beyond what the schema states. The openWorldHint suggests potential side effects that are not disclosed, so the description could be more thorough, but it covers the essential intent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with each parameter described, so the baseline is met. The description adds no parameter-specific guidance beyond what the schema already provides; it simply states the overall action. It does not explain relationships between parameters (e.g., how thread_id interacts with 'to' or 'subject'), but the schema descriptions are sufficient for basic use.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair: 'Save a draft' and clearly states the non-sending behavior. It distinguishes itself from send_message by explicitly noting when to use it ('whenever the user has not explicitly asked for the message to go out'), which differentiates it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit when-to-use condition ('Use this whenever the user has not explicitly asked for the message to go out'), which is clear and actionable. It does not name alternative tools like send_message or reply, but the implication is strong enough for an agent to infer the correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_create_docCreate a Google DocA
Create a real Google Doc — not a text file — from plain text. Line breaks are kept; formatting is not, because the content is converted from plain text on the way in. Returns the document id and a link to open it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Document title. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| content | No | Initial body text. May be empty. | |
| parent_folder_id | No | Destination folder. Defaults to the root of My Drive. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It adds genuine behavioral detail beyond annotations: line breaks are preserved, formatting is lost, content is converted from plain text, and the return value is a document id plus open link. This is exactly the nuance needed for a creation tool and goes well beyond what readOnlyHint/destructiveHint provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and the key distinction ('not a text file') is front-loaded. Every sentence earns its place, including the return-value note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple creation tool with no output schema, the description covers behavior, formatting constraints, and return value. Combined with a fully documented schema, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description only loosely maps 'from plain text' to the content parameter and does not add new per-parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a real Google Doc') and explicitly contrasts it with a text file, which distinguishes it from sibling tools like drive_upload. It also tells the agent what the tool returns, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage signal: use this when you need a real Google Doc from plain text, not when you want a plain text file. It does not name an alternative sibling tool explicitly, but the 'not a text file' exclusion plus the sibling list is enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_listList a Drive folderARead-onlyIdempotent
List what is inside a folder, sub-folders first. Omit "folder_id" for the root of My Drive. Use drive_search when you know what you are looking for but not where it lives.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| folder_id | No | Folder id. Defaults to "root", the top of My Drive. | |
| max_results | No | Maximum entries to return. Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful behavioral context beyond annotations by specifying the listing order ('sub-folders first'), which an agent would not otherwise know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences carry the core behavior, the root handling, and the main sibling alternative. Every sentence earns its place, and the primary action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple non-destructive listing tool with fully documented parameterschers, the description is mostly complete. It could be slightly stronger by mentioning what the returned entries look like, but since no output schema is present and the tool name/title make the listing nature clear, the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents account, folder_id, and max_results. The description's note about omitting folder_id for the root essentially restates the schema's default-to-root behavior, adding no new parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('List') and resource ('what is inside a folder'), and even states the traversal order ('sub-folders first'). It clearly distinguishes itself from drive_search by contrasting the use case, so an agent can tell this is the enumeration tool for a known folder.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance: omit folder_id for root, and use drive_search when you know what you are looking for but not where it lives. This directly tells the agent when to use this tool and when to prefer a sibling tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_readRead a Drive fileARead-onlyIdempotent
Read a file as text. Google Docs are exported to plain text, Sheets to CSV and Slides to plain text; ordinary text files are downloaded as they are. Binary files (PDF, images, archives) are refused with a link instead — this tool does not download them. Long content is truncated and says so.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| file_id | Yes | File id, as returned by drive_search. | |
| max_chars | No | Cut the content at this many characters. Default 60000. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as read-only, idempotent, open-world, and non-destructive, but the description adds valuable behavioral context beyond those: Google format exports, binary-file refusal with a link, and truncation behavior. It explicitly says truncation is indicated in the output, which an agent needs to interpret the result correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded. The opening sentence names the operation immediately, and each subsequent sentence adds a distinct, necessary behavioral fact without repetition or padding. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description covers what the tool returns: text content, exported formats, refusal with a link for binaries, and truncation notice. Combined with annotations that cover safety and the schema that covers parameters, this is sufficiently complete for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage and already documents each parameter clearly, including file_id provenance, account aliasing, and max_chars bounds/default. The description adds no additional parameter-level meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read a file as text.' It further describes the exact transformation behavior for different file types (Google Docs to plain text, Sheets to CSV, Slides to plain text), which clearly distinguishes it from sibling tools like drive_search, drive_list, drive_upload, and drive_create_doc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys when it is appropriate to use this tool: when the agent needs the textual content of a file. It also gives a when-not signal by saying binary files are refused and not downloaded. However, it does not explicitly name alternative sibling tools or provide a direct comparison, so it stops short of full exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_searchSearch DriveARead-onlyIdempotent
Find files in Google Drive. Omit "account" to search EVERY connected Drive in parallel and get one merged list, most recently modified first; if one account fails the others still return and the failure is reported in "failures". Use "query" for plain words (matched against filename and file contents) and "drive_query" only when you need raw Drive query syntax, e.g. "mimeType = 'application/pdf' and modifiedTime > '2026-01-01T00:00:00'".
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Free text. Matched against both the file name and the file contents. | |
| account | No | Restrict to one account. Omit to use every configured account at once. | |
| drive_query | No | Raw Drive query syntax, ANDed with "query" when both are given. | |
| max_results | No | Maximum files to return (per account when searching all). Default 20. | |
| include_trashed | No | Include files in the bin. Defaults to false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses parallel account search, merged result ordering, and partial failure semantics — behavior not captured in the annotations. This goes well beyond the readOnlyHint and idempotentHint annotations by explaining what happens when one account fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, default behavior, failure handling, and parameter selection guidance are packed into a compact, front-loaded description. No fluff or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers the key return behavior (merged list, ordering, failure reporting) and explains all important parameter interactions. An agent has enough information to call the tool correctly without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds crucial meaning: 'query' matches filename and contents, 'drive_query' uses raw syntax and is ANDed with 'query', and 'max_results' is per account when searching all. These details are not inferable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Find files in Google Drive', a specific verb and resource, then immediately explains the multi-account search behavior that sets it apart from sibling tools like drive_list or drive_read. The distinction between plain query and raw drive_query further clarifies the tool's scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear conditional guidance: omit 'account' to search all connected drives, use 'query' for plain words, and use 'drive_query' only when raw Drive syntax is required. It does not explicitly compare against sibling search tools like search_emails, but the usage context within this tool is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drive_uploadUpload a file to DriveA
Put a file in Drive. Give either "content" for text you composed, or "local_path" for a file that already exists on the machine running this server — exactly one of the two. The file is private to the account until something shares it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | File name in Drive, including its extension. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| content | No | Inline text content. | |
| mime_type | No | MIME type. Guessed from the file extension when omitted. | |
| local_path | No | Absolute path to a file on the machine running this server. | |
| parent_folder_id | No | Destination folder. Defaults to the root of My Drive. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, so the write nature is known. The description adds a useful privacy note ('The file is private to the account until something shares it'), which goes beyond annotations. It does not describe overwrite behavior, return format, or side effects like idempotency, but the annotations cover the basic safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The action is front-loaded, and the exclusivity constraint and privacy note are delivered efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema, the description covers the main decision (content vs. path) and the default folder behavior is in the schema. However, it does not mention what the tool returns (e.g., file ID or URL), which is important for chaining operations. It also omits any mention of authentication requirements or failure modes, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter has a description. The tool description adds critical meaning by specifying that content and local_path are mutually exclusive ('exactly one of the two'), a constraint not present in the schema. This clarifies a non-obvious relationship and aids correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource ('Put a file in Drive') and immediately distinguishes two input modes (inline content vs. local file). It is specific enough to separate this from sibling tools like drive_create_doc or drive_read, and the title reinforces the action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use content versus local_path ('exactly one of the two'), which is valuable. However, it does not explicitly contrast this tool with alternatives (e.g., drive_create_doc for Google Docs) or state when not to use it. The usage context is implied but not explicit about exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
label_messageAdd or remove labelsAIdempotent
Apply and/or remove labels on a message. Accepts label names or ids — names are resolved case-insensitively. Useful system labels: UNREAD, STARRED, IMPORTANT, SPAM, TRASH. Removing UNREAD marks a message as read.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| add_labels | No | Labels to apply. | |
| message_id | Yes | Message to modify. | |
| remove_labels | No | Labels to remove. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a non-read-only, idempotent, non-destructive mutation. The description adds real behavioral detail beyond the annotations: label names are matched case-insensitively, either names or IDs are accepted, and removing UNREAD has the side effect of marking the message read. This meaningfully improves the agent's understanding of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each carrying distinct information: the action, label identifier flexibility, useful system labels, and the UNREAD-to-read side effect. There is no redundancy or filler, and the key action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation with four parameterscraper, the description plus schema covers the required fields, label value formats, and important side-effect semantics. It does not explicitly warn that calling it with neither add_labels nor remove_labels is a no-op, nor describe the response shape, but these are minor gaps given the idempotentHint and the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining that add_labels and remove_labels accept both label names and IDs, that names resolve case-insensitively, and that the UNREAD label has special read-marking semantics. This helps the agent construct correct label values beyond just reading the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Apply and/or remove labels on a message') and clearly distinguishes this from sibling tools like list_labels and trash_email by focusing on label mutation. It also clarifies the label namespace and case-insensitive name resolution, leaving no doubt what operation is performed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it—when a message's labels need changing—and lists useful system labels, which gives practical context. However, it never explicitly contrasts this with list_labels, trash_email, or untrash_email, nor does it state conditions when an alternative should be preferred. The guidance is adequate but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_accountsList Gmail accountsARead-onlyIdempotent
List every configured Gmail account and whether it is usable. Call this first when you do not know which accounts exist, or when another tool reports an unknown account.
| Name | Required | Description | Default |
|---|---|---|---|
| probe | No | Contact Gmail to confirm each account really works. Slower, but detects revoked tokens that look fine on disk. Defaults to false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false), so the bar is lower. The description adds genuine behavioral context beyond that: accounts may exist yet be unusable, and this tool reports that state. It is consistent with annotations — listing is read-only and idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Exactly two sentences: the first delivers the core purpose, the second front-loads the usage trigger. No filler, no repetition of schema or annotation content, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool (one optional param, no required params, no output schema) with annotations covering safety, the description covers purpose and usage triggers well. The only minor gap is that the return format is not described and there is no output schema to fill it, but the shape of a discovery list is largely predictable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the single "probe" parameter is fully documented in the schema (boolean, contacts Gmail, slower, detects revoked tokens, defaults false). The description only hints at this via "whether it is usable," so the schema does the heavy lifting — baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"List every configured Gmail account and whether it is usable" names a specific verb (list), resource (configured Gmail accounts), and added scope (usability status). It is unambiguous against the sibling list_labels — that tool lists labels, not accounts — so an agent can tell them apart without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Call this first when you do not know which accounts exist, or when another tool reports an unknown account" gives explicit, actionable trigger conditions. For a discovery tool with no genuine sibling alternative, the absence of a when-not clause is not a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_labelsList labelsARead-onlyIdempotent
List the labels of an account, system and user-created, with their ids. Call this before label_message if you are unsure a label exists.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, and non-destructive behavior, so the description does not need to repeat them. It adds context about what categories of labels are returned and how results relate to label_message. It does not describe pagination or exact response shape, but those are minor for a simple label listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, with the primary action and scope in the first and usage guidance in the second. No filler, no repetition of schema content, and the information is effectively front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only listing with rich annotations, the description gives the required scope, the output contents (labels and ids), and when to call it relative to label_message. The absence of an output schema is compensated by explicitly stating that ids are included.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with the single account parameter already described as an email address or alias. The description only restates the account scoping ('of an account') and adds no new syntactic or format guidance. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List'), the resource ('labels of an account'), and the useful detail that both system and user-created labels are included with ids. The mention of 'Call this before label_message' also distinguishes it from the only sibling that operates on labels.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to call: before label_message when unsure a label exists. It points at the alternative label_message by name, giving clear routing guidance. It doesn't enumerate other alternatives, but none are needed for a pure list operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_messageRead a messageARead-onlyIdempotent
Read one message in full, including its decoded body and attachment metadata. Attachment contents are not downloaded.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| message_id | Yes | Message id, as returned by search_emails. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and idempotent behavior, and the description adds the valuable caveat that attachment contents are not downloaded. This goes beyond the structured metadata by clarifying the scope of the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary purpose is front-loaded and the crucial limitation about attachments is stated immediately after the core behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only tool with fully described parameters and safety annotations, the description covers the important behavior and limits. It could mention what is returned in more detail, but the absence of an output schema is mitigated by describing the body and attachment metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides full descriptions for both parameters (100% coverage). The tool description adds no parameter-specific detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Read one message in full') and explicitly states the scope of the result (decoded body and attachment metadata). This makes the tool easy to distinguish from siblings like search_emails and read_thread.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool versus alternatives like read_thread or how it fits into a workflow after search_emails. The description implies its use but does not state when not to use it or name any companion tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_threadRead a threadARead-onlyIdempotent
Read a full conversation, every message in order. Prefer this over read_message when you need the context of an exchange rather than a single email.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| thread_id | Yes | Thread id, as returned by search_emails. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish readOnly, openWorld, idempotent, and non-destructive behavior, so the description's added value is in stating that the tool returns the complete thread ordered as a full conversation. It also communicates that the output provides exchange-level context rather than an isolated message, which goes beyond the annotations. It does not describe the response fields or pagination, but no annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tightly written sentences with no filler. The first sentence states the core behavior and scope, and the second sentence adds the comparative usage guidance. Every word contributes to tool selection or invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only two required parameters and a full schema, the description covers the main behavior and expected output order. The only gap is the lack of detail about the exact structure of returned messages, such as headers or bodies, but the absence of an output schema and the simple read-only nature keep this close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, documenting both required parameters: thread_id is described as returned by search_emailseding an email address or alias for account. The description adds no additional parameter-level details, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Read', and a clear resource, 'a full conversation,' and states that it returns 'every message in order.' It explicitly contrasts with read_message by targeting the context of an exchange rather than a single email, making the tool's purpose unambiguous and distinguishing it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives direct selection guidance: 'Prefer this over read_message when you need the context of an exchange rather than a single email.' This names the alternative tool and the condition that should trigger the choice, so an agent knows when to use this tool versus read_message.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replyReply to a messageA
Send a reply that stays in the original thread. Handles the subject prefix and the In-Reply-To/References headers, which is what makes it appear as a reply rather than a new conversation. This SENDS immediately.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Message body. Plain text unless is_html is true. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| is_html | No | Send the body as text/html. | |
| reply_all | No | Also copy everyone in the original Cc. Defaults to false. | |
| message_id | Yes | The message being answered. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate non-readonly, non-idempotent, and non-destructive behavior. The description adds valuable context beyond that: it handles subject prefixes and In-Reply-To/References headersWorld, and it emphasizes that the action sends immediately. This is critical behavioral information for an agent deciding whether to invoke a side-effecting operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences deliver the essential message efficiently: what the tool does, why it behaves as a reply, and that it sends immediately. Every sentence earns its place, and the immediate-send warning is front-loaded at the end for emphasis without being wordy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter send action with no output schema, the description covers the key behavioral risks and thread semantics. It does not mention return values, errors, or delivery guarantees, but the schema fully documents parameters and the tool is simple enough that those omissions are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter clearly. The description adds no parameter-level meaning beyond the schema, such as how body, is_html, or reply_all interact. A baseline score of 3 is appropriate because the schema carries the full semantic weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Send a reply that stays in the original thread.' It clearly distinguishes the tool from sending a new message by emphasizing thread continuity and header handling. It also differentiates from draft creation with the explicit warning 'This SENDS immediately.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when the intent is to reply within an existing thread rather than start a new conversation. It also signals that the tool sends immediately, contrasting with a draft workflow. However, it never explicitly names alternatives like send_message or create_draft, nor states when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_emailsSearch emailsARead-onlyIdempotent
Search with Gmail query syntax (e.g. "from:ana@x.com is:unread newer_than:7d"). Omit "account" to search EVERY configured mailbox in parallel and get one merged, date-sorted list. If one account fails the others still return, and the failure is reported in "failures".
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Gmail search query, exactly as typed in the Gmail search box. | |
| account | No | Restrict to one account. Omit to search all of them. | |
| max_results | No | Maximum messages to return (per account when searching all). Default 20. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive behavior. The description adds meaningful behavioral context beyond that: parallel search across accounts, a merged and date-sorted result, partial failure tolerance, and an explicit 'failures' field. This substantially helps an agent anticipate what happens at runtime.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler. The query syntax example is front-loaded, followed by the key scoping behavior and then failure semantics. Every sentence contributes essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter search tool, the description plus annotation set covers defining query syntax, account scoping, result ordering, and partial failure behavior. It does not detail the shape of the returned message objects, but the absence of an output schema is partially mitigated by the clear description of the merged date-sorted list and failures field.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value by showing a concrete Gmail query example and emphasizing that omitting 'account' triggers a parallel search across all mailboxes. The max_results parameter is left to the schema, which documents it clearly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('search') and resource ('emails') and adds the essential Gmail query syntax. It clearly distinguishes this from sibling retrieval tools like read_message by framing it as query-based search, and the cross-mailbox behavior further pins down what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational guidance: omit 'account' to search all mailboxes, or include it to restrict the search. It describes a concrete invocation strategy and failure behavior. It does not explicitly name alternatives or exclusion criteria, but none of the sibling tools perform the same email search role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageSend an emailA
Compose and SEND a new email immediately. There is no undo. If the user has not clearly asked for it to be sent, use create_draft instead.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | Carbon-copy recipients. | |
| to | Yes | Recipients. Plain addresses or "Name <a@b.c>". | |
| bcc | No | Blind carbon-copy recipients. | |
| body | Yes | Message body. Plain text unless is_html is true. | |
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| is_html | No | Send the body as text/html instead of text/plain. | |
| subject | Yes | Subject line. Non-ASCII is encoded automatically. | |
| thread_id | No | Attach to an existing thread. Prefer the "reply" tool for answering a message. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds important behavioral context beyond the annotations: the send is immediate and irreversible ('There is no undo'). This is a critical warning for a mutating side-effectful action. While it doesn't discuss failures or delivery semantics, the core risk is clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the action and immediacy are front-loaded, the no-undo warning is prominent, and the alternative tool is named concisely. No filler or redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutating tool with no output schema, the description provides the necessary side-effect warning, user-intent safeguards, and routing guidance. Combined with the fully documented input schema, an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all 8 properties with meaningful descriptions, so the tool description adds little parameter-specific detail. Baseline 3 is appropriate because the schema carries the explanatory burden fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Uses a specific verb and resource ('Compose and SEND a new email immediately') and clearly differentiates from create_draft by stating that sending happens now. The title and description align, and a sibling tool shares the email domain without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs the agent to use create_draft when the user has not clearly asked for the message to be sent. This is a concrete conditional routing rule, which is ideal guidance for choosing between two otherwise similar sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trash_emailMove an email to the binADestructiveIdempotent
Move a message to the bin. This is NOT a permanent delete: the message keeps existing and untrash_email brings it back. Gmail does empty the bin by itself after 30 days, so it becomes permanent eventually — treat it as reversible for a month, not for ever. This server has no permanent-delete tool at all. The answer names the message that moved, so a wrong id is visible immediately.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| message_id | Yes | Message to move to the bin. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=false and destructiveHint=true, but the description adds significant behavior beyond that: it is not permanent, the message stays existing, untrash_email can restore it, Gmail auto-empties after 30 days, and a wrong ID is visible in the response. This gives the agent a much richer model of the tool's side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and then adds only high-value caveats: reversibility, the 30-day limit, absence of permanent delete, and response feedback. No sentence is wasted, and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers everything needed: what the tool does, its side effects, how to undo it, and what to look for in the response. It also addresses the relevant sibling untrash_email and the lack of a permanent-delete alternative.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters ('Configured account' and 'Message to move to the bin'). The description adds little beyond naming the bin, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Move a message to the bin.' It distinguishes this from a permanent delete and from recovery via untrash_email, so an agent can immediately tell what the tool is for and how it differs from related actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear by explaining that this is reversible for 30 days and that no permanent-delete tool exists. It doesn't literally say 'use this when you want to soft-delete and use untrash_email to undo,' but that is strongly implied by the contrast with untrash_email and the warning about Gmail's 30-day emptying.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
untrash_emailRecover an email from the binAIdempotent
Take a message back out of the bin and restore the labels it had. Only works while the message is still there — once Gmail has emptied the bin, after about 30 days, there is nothing left to recover. Use search_emails with "in:trash" to find what is in there.
| Name | Required | Description | Default |
|---|---|---|---|
| account | Yes | Configured account: the email address, or the alias given at setup. | |
| message_id | Yes | Message to take out of the bin. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond the annotations: it restores prior labels and has a 30-day recovery window. It does not describe failure modes or authorization requirements, but the annotations already signal mutation and idempotency, and nothing contradicts the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences deliver purpose, limitation, and the discovery alternative in order. Every sentence earns its place and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter recovery tool, this description is nearly complete: what it does, when it works, and how to find eligible messages are all present. The only missing piece is the response format, but with no output schema and a side-effect-driven operation, this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and both account and message_id have meaningful schema descriptions. The tool description does not add extra format or syntax guidance for the parameters, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a concrete action, 'Take a message back out of the bin and restore the labels it had,' specifying both the resource (trashed email) and the expected effect. This clearly distinguishes untrash_email from trash_email and search_emails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States a hard precondition: only works while the message is still in the bin, with a 30-day Gmail expiration limit. It also explicitly directs the agent to search_emails with 'in:trash' to find recoverable messages, giving a clear when-to-use and when-to-look-elsewhere guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
29 tool updates
v0.3.0- First observed
calendar_create_event - First observed
calendar_delete_event - First observed
calendar_find_free_time - First observed
calendar_get_event - First observed
calendar_list_events - First observed
calendar_update_event - First observed
contacts_create - First observed
contacts_get - First observed
contacts_list - First observed
contacts_search - First observed
contacts_update - First observed
create_draft - First observed
drive_create_doc - First observed
drive_list - First observed
drive_read - First observed
drive_search - First observed
drive_share - First observed
drive_shared_drives - First observed
drive_upload - First observed
label_message - First observed
list_accounts - First observed
list_labels - First observed
read_message - First observed
read_thread - First observed
reply - First observed
search_emails - First observed
send_message - First observed
trash_email - First observed
untrash_email
TDQS
Scored across 29 tools
Every tool maps to a distinct resource and action: search/read/draft/send/reply/label/trash for email, search/read/list/upload/share for Drive, CRUD plus free-time for Calendar, and search/get/list/create/update for Contacts. read_thread vs read_message and create_draft vs send_message are explicitly scoped to avoid confusion.
Most tools follow a clear verb_noun pattern, and drive_*, calendar_*, and contacts_* prefixes make domains predictable. Minor deviations like the bare verb "reply", the unprefixed list_labels/label_message, and the noun-phrase drive_shared_drives keep it from being perfectly consistent.
29 tools is high in absolute terms, but the server covers four distinct Google services plus multi-account handling, giving each domain a reasonable 5-8 tool surface. It is slightly heavy but every tool serves a distinct purpose and none feel redundant.
Core workflows are well covered: Gmail has search/read/draft/send/reply/label/trash, Calendar has full CRUD and free-time search, Contacts has create/read/update/search, and Drive has search/read/upload/share/create. Minor gaps exist such as no contacts_delete, no Drive delete or folder creation, and no attachment downloading, but agents can complete most real tasks.
Maintenance
Related MCP Connectors
Multiple Google accounts (Gmail, Calendar, Drive, Contacts, Tasks) in one Claude connector.
Multiple Google accounts (Gmail, Calendar, Drive, Contacts, Tasks) in one Claude connector.
Multiple Gmail accounts, editable Google Sheets & Docs for AI agents. Deny-by-default access rules.
Gmail, Outlook, Drive, OneDrive and calendars for AI agents. Many accounts, one endpoint, audit log.
Related MCP Servers
- FlicenseAqualityBmaintenanceConnects AI assistants to multiple Gmail accounts simultaneously, enabling search, read, draft, send, and reply operations with per-account permission controls.54-
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to access multiple Gmail/Google Workspace accounts for searching and reading mail, calendars, and attachments.-
- AlicenseNot gradedqualityCmaintenanceConnects Gmail to AI assistants via the MCP protocol, enabling search, read, send, and manage emails across multiple Google accounts simultaneously.MIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude to access and manage multiple Google accounts at once across Gmail, Calendar, Drive, Contacts, and Tasks, with unified search and document text extraction.MIT