Skip to main content
Glama

gmail-multi-mcp

An MCP server that connects your AI assistant to several Google accounts at once — Gmail, Drive, Calendar and Contacts.

License: MIT Node TypeScript MCP Tools


Why this exists

The official connectors handle one account. If your work lives in you@company.com, your invoices arrive at billing@company.com and your life happens in you@gmail.com, you end up switching connectors — or simply cannot ask a question that spans two of them.

This server makes the account a first-class parameter, and then goes one step further:

"Any unread invoices this week, and do I have a free hour to deal with them?"

  → search_emails          searches work@, billing@ and personal@ in parallel
  → calendar_find_free_time gaps that are free on ALL THREE calendars
  → one answer

Omit the account and it searches every mailbox, every Drive, every calendar and every address book you have connected. One Google Cloud project, as many accounts as you want.


Related MCP server: multi-mail-mcp

Features

  • Multi-account by design. Add as many Google accounts as you like; each keeps its own credentials.

  • Cross-account by default. search_emails, drive_search, calendar_list_events, calendar_find_free_time and contacts_search all work across every connected account when you omit account.

  • Shared drives included. Drive returns only My Drive unless a request opts in, so every call here does. drive_search looks across My Drive and every shared drive the account can reach; drive_list, drive_read, drive_upload and drive_share all work on shared-drive items.

  • A shared drive is not a file, so it has its own tool. A drive never shows up in drive_search — not even searched by its exact name, with full access. drive_shared_drives is the only way to turn a drive's name into the id every other tool wants.

  • A partial file list says so. Searching every drive lets Drive quietly omit documents. When it admits that, the affected accounts come back in incompleteAccounts instead of a count that looks complete and is not.

  • Degrades instead of failing. If one account's token is revoked, the others still return — and the broken one is reported in a failures field rather than silently dropped.

  • Free time across accounts. calendar_find_free_time intersects the busy blocks of several calendars. If it cannot read one of them it returns no slots at all rather than a confident answer built on half the picture.

  • Aliases. Call an account work or personal instead of typing the full address.

  • Standard OAuth, no password ever. Consent happens in your browser, on Google's own screen, through a CSRF-protected loopback callback.

  • Tokens stay on your machine, in your home directory, 0600 on Unix — never in the repository, never in the MCP client configuration.

  • Automatic token refresh, written back to disk so a restart does not re-refresh.

  • Per-service scope gate. An account authorised before Drive existed keeps doing Gmail and gets a clear MISSING_SCOPE for Drive — not a baffling Google 403.

  • Twenty-nine tools across mail, files, calendar and contacts.

  • Contacts that cannot be clobbered. contacts_update re-reads the contact for its etag before writing, so a change made on a phone thirty seconds earlier is not silently overwritten.

  • Cannot permanently delete email. trash_email moves a message to the bin and untrash_email brings it back. There is no permanent-delete tool: it would need a scope this server does not request, and an unrecoverable action does not belong behind an agent.

  • Strict TypeScript, no any, and a dependency list you can read in one glance.


Prerequisites

Node.js

≥ 20.12 — see the note below

Google Cloud

A free account. No billing needed for personal use.

An MCP client

Claude Code, Claude Desktop, or anything that speaks MCP over stdio.

Why 20.12 and not 18. The server loads a local .env with Node's built-in process.loadEnvFile(), which landed in 20.12 (and 21.7). On Node 18 that file would be ignored silently, and you would get a confusing "no credentials" error with a perfectly good .env sitting right there. The floor is real, not cautious.


Installation

git clone https://github.com/blackwings-dev/gmail-multi-mcp.git
cd gmail-multi-mcp
npm install
npm run build

That produces dist/index.js, which is what your MCP client will run. Check it:

node dist/index.js --version

Google Cloud setup

You do this once, no matter how many accounts you connect. About five minutes.

1. Create a project

Open the Google Cloud Console and create a project — call it whatever you like, e.g. google-mcp.

2. Enable four APIs

With your new project selected, enable each of these. Missing one produces a 403 that says nothing useful, and only for the tools that need it:

Contacts live behind the People API, not a "Contacts API". Looking for the latter in the library and finding the long-dead Contacts API v3 is the usual wrong turn.

Press Enable on each. Nothing else on those pages matters.

In the sidebar this now lives under Google Auth Platform; older projects show it as APIs & Services → OAuth consent screen.

Field

What to put

App name

Anything, e.g. Google MCP. Only you will see it.

User support email

Your own address.

Audience / User type

External

Developer contact

Your own address again.

Then, under Audience → Test users, press Add users and add every Google address you intend to connect.

⚠️ This step is not optional. While the app is in Testing, an account that is not listed as a test user cannot grant consent — Google shows an "access blocked" screen that does not explain why.

4. Create the OAuth client

APIs & Services → Credentials → Create Credentials → OAuth client ID

Field

Value

Application type

Desktop app ← this exact type

Name

Anything, e.g. gmail-multi-mcp

Copy the Client ID and Client secret from the dialog that appears.

Why "Desktop app" specifically. That client type accepts any loopback port as a redirect URI. The setup flow opens a throwaway server on a random free port and hands that URL to Google, so you never register redirect URIs by hand — which is exactly what lets a single project serve any number of accounts.


Scopes, and what they cost you

This server asks for four scopes:

Scope

What it buys

Google's tier

.../auth/gmail.modify

Read, draft, send, label. Not permanent deletion.

Restricted

.../auth/drive

Full access to the user's files.

Restricted

.../auth/calendar

Read and write events, and free/busy.

Sensitive

.../auth/contacts

Read and write the account’s own contacts.

Sensitive

Be aware of what full drive costs, because it changes your options later. Google classes it as a restricted scope, the same tier as gmail.modify. The narrower drive.file — access limited to files this app created or that the user explicitly picked in a Google file picker — is not restricted and would keep the verification story simpler. But drive.file cannot see files you already have, which means drive_search over your existing Drive would return nothing. That is the whole trade-off: searchable Drive, or an easier path through Google's review.

Full drive is the deliberate choice here. Switching is a one-line change in src/auth/oauth.ts (DRIVE_SCOPES) followed by re-authorising your accounts.

contacts is read and write because contacts_create and contacts_update exist; contacts.readonly would do if you only ever look. It does not cover Google’s "other contacts" — the addresses Gmail collects by itself — which sit behind a separate scope this server does not ask for.


Configuring the server

1. Credentials

cp .env.example .env
GOOGLE_CLIENT_ID=1234567890-abcdefg.apps.googleusercontent.com
GOOGLE_CLIENT_SECRET=GOCSPX-your-secret-here

.env is git-ignored. You can also skip this file entirely — npm run setup will ask for the credentials and store them in ~/.gmail-multi-mcp/config.json. If both exist, the environment wins.

2. Connect your accounts

npm run setup
  gmail-multi-mcp
  Gmail, Drive and Calendar. One Google Cloud project, as many accounts as you need.

  config  C:\Users\you\.gmail-multi-mcp\config.json
  tokens  C:\Users\you\.gmail-multi-mcp\tokens

✔ OAuth credentials found in the environment.

What now?
─────────
  1  Add a Google account
  2  List accounts
  3  Remove an account
  4  Show MCP client configuration
  q  Quit

Choose 1. A browser opens and you sign in. Google’s consent screen lists Gmail, Drive, Calendar and Contacts; grant them, and the tab says "Account connected". The CLI then asks for an optional alias — work, personal, billing — which you can use instead of the full address from then on.

⚠️ Adding a second account

Sign out of Google first, or use a private/incognito window.

Otherwise the browser reuses the session you already have, Google skips the account chooser, and you connect the same account twice without noticing.

Repeat for each account, then check the result with option 2.

3. Register the server with Claude Code

claude mcp add gmail-multi-mcp \
  --scope user \
  -e GOOGLE_CLIENT_ID=your-client-id.apps.googleusercontent.com \
  -e GOOGLE_CLIENT_SECRET=GOCSPX-your-secret \
  -- node /absolute/path/to/gmail-multi-mcp/dist/index.js

On Windows the path looks like D:\Dev\gmail-multi-mcp\dist\index.js. Verify with claude mcp list, then restart Claude Code so it picks the server up.

{
  "mcpServers": {
    "gmail-multi-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/gmail-multi-mcp/dist/index.js"],
      "env": {
        "GOOGLE_CLIENT_ID": "your-client-id.apps.googleusercontent.com",
        "GOOGLE_CLIENT_SECRET": "GOCSPX-your-secret"
      }
    }
  }
}

The env block is optional if the credentials are already in ~/.gmail-multi-mcp/config.json. Account tokens are always read from there — they never go in the client configuration.

Option 4 in the setup CLI prints this snippet with the correct absolute path already filled in.


⚠️ Upgrading from a Gmail-only version? Re-authorise

If you connected your accounts when this server only did Gmail, they will keep working for Gmail and fail for everything else. Drive and Calendar tools will answer:

MISSING_SCOPE — <account> is authorised, but not for Drive.

A refresh will not fix this. Google binds a refresh token to the exact set of scopes granted when it was issued, so no amount of restarting, re-refreshing or waiting will add Drive and Calendar. Only a fresh consent will.

The fix takes a minute per account:

npm run setup

The CLI tells you which accounts are affected before you do anything:

! 3 accounts need re-authorising:
    · you@company.com
    · billing@company.com
    · you@gmail.com
  They were connected before Drive and Calendar support existed, so they can only
  do Gmail right now. A refresh cannot add the missing scopes: Google binds a
  refresh token to the scopes it was issued with. Choose 1 and add each account
  again — it re-consents and overwrites in place, so nothing has to be removed.

Choose 1 and add each account again. You do not have to remove them first: the flow always asks for fresh consent and overwrites the stored token in place, keeping the alias.

Your assistant can check this for itself at any time — list_accounts reports missingScopes and needsReauthorization per account, and a warning summarising them.

And do not forget step 2 of the Google Cloud setup: an account with the right scopes still fails if the Drive or Calendar API is not enabled on the project.


Usage

Ask in natural language. A few that exercise different corners:

What you say

What happens

"Any unread invoices this week?"

search_emails across all mailboxes, merged

"Bin that newsletter"

trash_email — and untrash_email if it was the wrong one

"Find the Q3 budget spreadsheet in any of my Drives"

drive_search across all accounts

"Read me that spreadsheet"

drive_read — exported to CSV

"Draft a reply to Ana with the numbers from it"

drive_read + create_draft

"Write those notes up as a Google Doc in my work account"

drive_create_doc

"Share it with marta@client.com as a commenter"

drive_share

"What's on my calendars tomorrow?"

calendar_list_events across all accounts

"Find me two free hours next week across all my accounts"

calendar_find_free_time

"Book it as 'Budget review' and invite Marta"

calendar_create_event

"Move Thursday's stand-up an hour later"

calendar_update_event

"What's Marta's phone number?"

contacts_search across all address books

"Add her to my work contacts with that email"

contacts_create

Habits worth having:

  • Ask for a draft unless you really want the mail sent. reply and send_message go out immediately and there is no undo.

  • Name the account ("in my work account") when you mean one; leave it out when you want all of them.

  • calendar_delete_event and drive_share are the two that change things you cannot take back from here. Be explicit with them.


The tools

account accepts an email address or an alias. Where it is optional, omitting it means "every connected account, merged".

Accounts

Tool

Description

Parameters

list_accounts

Configured accounts, their scopes, and which need re-authorising

probe? — contact Google to confirm (slower, catches revoked tokens)

Gmail

Tool

Description

Parameters

search_emails

Gmail query syntax. Without account, searches every mailbox

query, account?, max_results? (1–100, default 20)

read_thread

A full conversation, every message in order

thread_id, account

read_message

One message: decoded body plus attachment metadata

message_id, account

create_draft

Saves a draft without sending

account, to[], subject, body, cc?, bcc?, is_html?, thread_id?

reply

Replies in-thread — handles Re:, In-Reply-To, References. Sends.

account, message_id, body, reply_all?, is_html?

send_message

Composes and sends a new email. No undo.

account, to[], subject, body, cc?, bcc?, is_html?, thread_id?

list_labels

Labels of an account, with their ids

account

label_message

Adds and/or removes labels; accepts names or ids

account, message_id, add_labels?, remove_labels?

trash_email

Moves a message to the bin. Reversible, not a permanent delete

account, message_id

untrash_email

Takes a message back out of the bin, restoring its labels

account, message_id

Useful system labels for label_message: UNREAD (remove it to mark as read), STARRED, IMPORTANT, SPAM.

On the bin. trash_email is the tool for it; adding the TRASH label by hand does the same thing less clearly. Neither is a permanent delete — but "not permanent" deserves its footnote: Gmail empties the bin by itself after about 30 days, so a trashed message is recoverable for a month and then it is gone. search_emails with in:trash shows what is still in there. There is no permanent-delete tool at all: users.messages.delete needs the full-mailbox scope this server deliberately does not request.

Attachment metadata is returned (filename, MIME type, size); attachment contents are not downloaded.

Drive

Tool

Description

Parameters

drive_search

Without account, searches every Drive. Covers My Drive and shared drives. query is free text over name and contents; drive_query takes raw Drive query syntax. Returns incompleteAccounts when Drive admits it dropped documents — narrow the query if you see it

query?, drive_query?, account?, max_results? (1–100, default 20), include_trashed?

drive_read

Reads a file as text

file_id, account, max_chars? (default 60000)

drive_list

Lists a folder, sub-folders first

account, folder_id? (default root), max_results?

drive_shared_drives

Lists shared drives with their ids. A drive is not a file, so it never appears in drive_search; this is the only way to get its id. canAddChildren says whether that account can write to it

account?, max_results?

drive_upload

Uploads a file. Exactly one of content or local_path

account, name, content?, local_path?, mime_type?, parent_folder_id?

drive_create_doc

Creates a real Google Doc from plain text

account, name, content?, parent_folder_id?

drive_share

Grants access. Gives away real data.

account, file_id, type, role, email_address?, domain?, notify?, message?

drive_read conversions: Google Docs → text/plain, Sheets → text/csv, Slides → text/plain. Ordinary text files are downloaded as they are. Binary files (PDF, images, archives) are refused with a link instead — this server does not download them. Content over max_chars is cut and says so; files over 10 MB are refused.

drive_share types: user and group need email_address; domain needs domain; anyone makes the file readable by anyone with the link. Roles are reader, commenter or writer — ownership transfer is deliberately not offered. No notification email is sent unless notify is true.

Contacts

Contacts are handled through Google’s People API.

Tool

Description

Parameters

contacts_search

Without account, searches every address book. Matches names, nicknames, emails, phone numbers and organisations

query, account?, max_results? (1–100, default 20)

contacts_get

One contact in full: every email and phone with its label, organisation, addresses, links, notes

account, resource_name

contacts_list

Pages through an address book, most recently changed first

account, page_size? (1–200, default 50), page_token?

contacts_create

Adds a person. At least one field required

account, given_name?, family_name?, emails?, phones?, organization?, job_title?, notes?

contacts_update

Changes a person. Each named field is replaced, not merged

account, resource_name, + any of the create fields

resource_name looks like people/c1234567890 and comes from contacts_search or contacts_list.

⚠️ contacts_update replaces whole fields. Sending emails: ["new@x.com"] leaves the contact with exactly that one address and deletes the rest. Fields you do not name are left alone — so read the contact first and send back the complete list of whatever you are editing. The current version is re-read for its etag before writing, so an edit made on a phone in the meantime cannot be silently overwritten.

The same person in two accounts is returned twice, on purpose. Each copy has its own resource_name in its own address book; merging them would produce an id that edits the wrong one.

Paging: contacts_list returns nextPageToken when there is more. Pass it back as page_token.

Calendar

Tool

Description

Parameters

calendar_list_events

Without account, reads every calendar. Repeating events are expanded into real occurrences

time_min, time_max, account?, calendar_id?, max_results? (1–250, default 25), query?

calendar_get_event

One event in full, with attendees and their responses

account, event_id, calendar_id?

calendar_create_event

Creates an event

account, summary, start, end, all_day?, time_zone?, description?, location?, attendees?, calendar_id?, send_updates?

calendar_update_event

Changes an event. start and end go together or not at all

account, event_id, + any of the create fields

calendar_delete_event

Deletes an event. Irreversible.

account, event_id, calendar_id?, send_updates?

calendar_find_free_time

Gaps free on every account asked about

time_min, time_max, accounts?, timezone?, workday_start?, workday_end?, weekdays_only?, minimum_minutes?, calendar_id?

Times must carry an explicit offset — 2026-09-10T09:00:00+02:00 or a Z. A bare 2026-09-10T09:00:00 is refused, on purpose: it would be read in whatever zone the server happens to run in, and the resulting answer looks completely normal while being hours wrong.

All-day events use plain dates (YYYY-MM-DD) with all_day: true. Google treats the end date as exclusive, so a one-day event ends on the following day.

send_updates defaults to none — attendees are added to the event but not emailed unless you ask. Values: all, externalOnly, none.

The calendar_find_free_time contract

This is the one tool where a wrong answer looks exactly like a right one, so its rules are worth stating:

  • time_min / time_max must carry an explicit offset. Maximum range: 62 days.

  • Working hours (workday_start, workday_end, default 09:00–18:00) are interpreted in timezone, an IANA name like Europe/Madrid. It defaults to the zone the server runs in, and is always echoed back in the result so you can see what was assumed. Daylight saving is handled; the offset is resolved per day, not fixed once.

  • weekdays_only defaults to true.

  • A busy block that overlaps a candidate window only partially still removes the overlapping part. Half a busy hour is not free.

  • Returned slots are the free segments of at least minimum_minutes (default 30), in UTC, not a grid of fixed start times.

  • If any calendar cannot be read, no slots are returned at all. The result carries incomplete: true, the failures, and an explanation. A gap computed without one person's calendar is a double-booking, not an answer.


Troubleshooting

MISSING_SCOPE — <account> is authorised, but not for Drive

The account was connected before Drive and Calendar support existed. See Re-authorise above. A refresh cannot fix it; only a fresh consent can.

Access denied by Drive: Request had insufficient authentication scopes

The scope gate let this through, so the stored token claims the scope but Google disagrees — usually a half-finished re-authorisation. Add the account again.

Access denied by Drive / by Calendar / by Contacts (HTTP 403)

The API is not enabled on the Google Cloud project. Enabling one does not enable the others — see step 2. For contacts the one you need is the People API.

Google’s contact search reads a server-side index that has to be warmed with an empty query before it answers anything. This server does that on the first search per account and retries once after a short pause, so you should not meet it — but if a brand new search comes back empty, ask again. A second attempt against a warm cache is the difference between "no such person" and "not ready yet".

A contact lost its other email addresses after an update

contacts_update replaces each field it is given. Passing one address removes the others. Read the contact with contacts_get first and send the full list back.

HTTP ERROR 431 (Request Header Fields Too Large) on the callback

Fixed in current versions — if you see it, you are on an old build. The cause is worth knowing: the redirect used to go to localhost, which is a shared cookie origin. Every dev server you have ever run on any localhost port can leave cookies there, and the browser sends all of them to the OAuth callback. Past roughly 16 KB that exceeds Node's default header limit — and it fails after Google has already granted consent, so the authorisation code is lost.

The redirect now targets 127.0.0.1, which does not receive localhost cookies, and the callback server accepts 64 KB of headers. If it somehow still happens, clear cookies for 127.0.0.1 and retry.

Accounts stop working after 7 days

Your OAuth consent screen is still in Testing, and Google expires refresh tokens for apps in that state after seven days. Two options, neither free:

  • Stay in Testing and re-run npm run setup once a week.

  • Publish the app (Google Auth Platform → Audience → Publish app). Refresh tokens stop expiring. In exchange, an app that has not been through Google's review shows an "unverified app" warning screen on every consent and is capped at a limited number of users.

Two of the four scopes here are restricted (gmail.modify, drive), the category Google reviews most closely, and that review policy changes. Check the current requirements on the consent screen itself before assuming publishing will be waved through.

Unknown account "..."

The reference did not match any configured email or alias. Run npm run setup and choose 2, or ask your assistant to call list_accounts.

Google did not return a refresh token

Google issues one only on first consent. Revoke the app at myaccount.google.com/permissions, then authorise again.

time_min must be ISO 8601 with an explicit UTC offset

Working as intended. Send 2026-09-10T09:00:00+02:00, not 2026-09-10T09:00:00.

You connected the same account twice

The browser reused an existing Google session. Remove the duplicate with option 3, then add the other account from a private window.

Your MCP client shows a JSON parse error and disconnects

Something wrote to stdout. On the stdio transport, stdout is the protocol. If you are modifying this server, every diagnostic must go to stderr — never console.log.

A cross-account search returns fewer results than expected

Look at the failures field of the response. An account whose token was revoked is reported there rather than silently skipped.


Security

  • Tokens never leave your machine. They live in ~/.gmail-multi-mcp/tokens/, one file per account, chmod 0600 on Unix. Nothing is sent anywhere except Google.

  • Nothing sensitive is in the repository. .gitignore covers .env and every .env.* except the example.

  • CSRF-protected callback. A random state is generated per flow and compared in constant time; a callback that does not match is rejected.

  • Loopback only. The callback server binds 127.0.0.1, on an ephemeral port, and shuts down as soon as the code arrives.

  • Header injection blocked. CR/LF is stripped from every outgoing mail header, so a crafted subject or recipient cannot inject extra headers.

  • No permanent mail deletion. The Gmail scope cannot do it and no tool offers it.

  • About the client secret: for Desktop app clients Google does not treat it as confidential — an installed application cannot keep a secret, which is why this flow is designed not to depend on one. Keep it out of your repository anyway.

  • Removing an account deletes the local token but does not revoke Google's grant. Revoke it fully at myaccount.google.com/permissions.

⚠️ What you are handing an assistant, and the risk that comes with it

This server can read untrusted text (any email anyone sends you, any document shared with you) and, in the same session, read local files, upload them to Drive, and share them publicly. Those two capabilities together are an exfiltration path if the model acts on instructions it finds inside the content it is reading.

Nothing in this server can prevent that, because from the API's point of view a malicious instruction in an email body and a genuine request from you look identical. What it does instead is make the dangerous steps loud and deliberate:

  • drive_share is annotated as destructive, defaults to sending no notification, refuses ownership transfer, and requires type: "anyone" spelled out to create a public link.

  • drive_upload will not read a local file unless given an explicit local_path.

  • calendar_delete_event is annotated as destructive and reads the event first so the answer says what it removed.

Use an MCP client that asks before running write tools, and read what it is about to do. Treat "an email told me to share this file" as the red flag it is.


How it works

MCP client ──stdio──▶ src/index.ts ──▶ src/server.ts        29 tools, zod-validated
                                          │
     ┌──────────────┬────────────────────┼────────────────────┬──────────────┐
     ▼               ▼                    ▼                    ▼
src/gmail/      src/drive/          src/calendar/        src/contacts/
search, MIME    queries, exports,   events, free/busy    People API,
parsing and     uploads and         and time-zone        search warm-up,
building        sharing             arithmetic           etag-guarded writes
     └──────────────┴────────────────────┼────────────────────┴──────────────┘
                                          ▼
                                      src/core/
                    AccountClientCache · the SCOPE GATE · error taxonomy
                    cross-account runner · bounded concurrency
                                          │
                                          ▼
                                  src/auth/oauth.ts
                          loopback consent, automatic refresh
                                          │
                                          ▼
                              src/auth/token-store.ts
                                ~/.gmail-multi-mcp/

Path

Contents

~/.gmail-multi-mcp/config.json

OAuth client (if not in the environment) and the account list

~/.gmail-multi-mcp/tokens/<email>.json

One refresh token per account, with the scopes it was granted

The scope gate lives in one place — AccountClientCache in src/core/google-client.ts. Every Drive and Calendar call has to go through a client, and every client comes from that class, so an account missing a scope is refused by construction rather than by each of twenty-eight handlers remembering to ask. A check spread across handlers has a blind spot the moment someone adds one more.

Auto-refresh is handled by Google's auth client: an expired access token is renewed transparently, and a tokens listener writes the new one back to disk — preserving the refresh token, which Google omits from refresh responses.


Development

npm install
npm run build        # tsc → dist/
npm run typecheck    # no emit
npm run dev          # run the server from source with tsx
npm run setup        # run the setup CLI from source

Smoke-test the server without an MCP client:

printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"t","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}' \
  | node dist/index.js

Layout

src/
  index.ts            entry point: server by default, `setup` for the CLI
  server.ts           MCP server and all twenty-eight tool definitions
  core/
    errors.ts         the error taxonomy and Google's HTTP failures mapped onto it
    google-client.ts  per-account client cache AND the scope gate
    accounts.ts       run one operation across every account, failures as values
    concurrency.ts    bounded parallelism
  auth/
    oauth.ts          scopes, consent flow, loopback callback, automatic refresh
    token-store.ts    config.json and per-account token files
  gmail/
    client.ts         authenticated clients, MIME parsing and building
    search.ts         single- and multi-account search
    types.ts          Gmail domain types
  drive/
    client.ts         queries, exports, uploads, sharing
    types.ts          Drive domain types
  calendar/
    client.ts         events and free/busy arithmetic
    timezone.ts       wall-clock time in a named zone, and interval maths
    types.ts          Calendar domain types
  contacts/
    client.ts         People API: search warm-up, etag-guarded writes
    types.ts          Contacts domain types
  cli/
    setup.ts          interactive setup (chalk)
assets/             the project icon: PNG and SVG, with and without wordmark

The SVG in assets/ is the master. Its geometry is also inlined in src/auth/icon.ts as the favicon of the local OAuth callback page — the only web page this server ever serves.

TypeScript runs with strict plus noUncheckedIndexedAccess, noImplicitReturns, noUnusedLocals and verbatimModuleSyntax. There is no any in the source.


A note on the name

The project is still called gmail-multi-mcp, and that name is now too small for it. Gmail is one of three services. google-workspace-mcp or google-multi-mcp would describe it better.

The rename is deliberately not done yet, because it is wider than it looks. Its surface includes:

  • the package name, the binary name and the repository;

  • ~/.gmail-multi-mcp/ — the config and token directory, which means existing users would have to move their credentials or the server would look freshly installed;

  • GmailMcpError, the error class every layer throws, including Drive and Calendar;

  • SERVER_NAME, and the literal string "gmail-multi-mcp setup" in a dozen user-facing error messages.

Worth doing in one deliberate pass, with a migration for the token directory — not piecemeal.


Contributing

Issues and pull requests are welcome.

  1. Fork the repository and create a branch: git checkout -b feature/what-it-does.

  2. Make your change. npm run typecheck and npm run build must both pass.

  3. Smoke-test the server as shown above.

  4. Open a pull request describing what changed, why, and how you verified it.

Please keep the five rules that hold the design together:

  • Nothing writes to stdout outside src/cli/ and the --help / --version paths. Stdout is the MCP transport.

  • No any. Google's types are noisy; convert them at the boundary in each service's client.ts and keep the rest of the codebase on the domain types.

  • New service? Go through AccountClientCache. That is where the scope gate lives. A client built any other way skips it.

  • Cross-account operations report their failures. Never return a shorter list and stay quiet about the account that did not answer.

  • Nothing unrecoverable. trash_email is reversible; a permanent-delete counterpart is not, and does not belong behind an agent. The Gmail scope this server requests cannot do it, and that is the point — do not widen the scope to add one.

Good first contributions: attachment download, Google Docs formatting via the Docs API, recurring-event rules, contact groups and labels, Google’s "other contacts", mark_read / archive convenience tools, or shared-drive support.

Reporting a security issue

Please do not open a public issue. Contact the maintainers privately first.


License

MIT — do what you like, no warranty.


Who maintains this

Built and maintained by Black Wings — software, AI and managed IT for companies, from Valencia, Spain. We wrote this because we needed it: we run several Google Workspace mailboxes and no connector could see more than one at a time.

If it saves you the same afternoon it saved us, that is the whole point. Other things we have released are listed at blackwings.dev/codigo-abierto (in English: blackwings.dev/en/open-source).

Available Tools

29 tools
calendar_create_eventCreate a calendar eventA

Create an event. For a timed event, "start" and "end" are ISO 8601 WITH an offset. For an all-day event set "all_day" and give plain dates (YYYY-MM-DD) — note that Google treats the end date as EXCLUSIVE, so a one-day event ends on the following day. Listing attendees puts the event on their calendars; they are only emailed when "send_updates" says so.

ParametersJSON Schema
NameRequiredDescriptionDefault
endYesEnd. ISO 8601 with offset, or YYYY-MM-DD if all_day.
startYesStart. ISO 8601 with offset, or YYYY-MM-DD if all_day.
accountYesConfigured account: the email address, or the alias given at setup.
all_dayNoTreat start and end as plain dates.
summaryYesEvent title.
locationNoWhere it happens.
attendeesNoAttendee email addresses.
time_zoneNoIANA zone stored with the event, e.g. "Europe/Madrid".
calendar_idNoCalendar id. Defaults to "primary", the account own calendar.
descriptionNoBody of the event.
send_updatesNoWhether Google emails the attendees. Defaults to "none": nobody is mailed.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate a state-changing operation (readOnlyHint false). The description goes beyond the annotations by disclosing that adding attendees places the event on their calendars, and that emails are only sent when send_updates is set. It also explains an important API quirk (exclusive end date for all-day events). This adds behavioral context that the annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with zero fluff. It front-loads the primary action ('Create an event') and then packs the critical nuances (date formats, exclusive end, attendee/email behavior) into compact, useful clauses. Each sentence earns its place; no redundant or vague phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema, the description covers the key gotchas: timed vs. all-day handling, date-exclusivity, attendee calendar placement, and send_updates default. It does not mention the return value, but that is acceptable given no output schema. It omits no critical information that an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter already has a description. The description adds value by clarifying the difference between timed ISO 8601 offsets and plain dates for all-day events, and by explaining the exclusive end-date behavior. It also notes the side effect of attendee emailing tied to send_updates. This goes beyond the schema's literal param descriptions, so a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the specific action 'Create an event,' which is a clear verb + resource. It further distinguishes the tool's purpose by detailing the two valid input formats (timed vs. all-day), which are unique to creating an event. This makes it unambiguous against sibling tools like calendar_update_event or calendar_delete_event.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: when to use timed vs. all-day formats, how to set attendees, and the effect of send_updates. It does not explicitly name alternatives (e.g., use calendar_update_event to modify) but the create/update distinction is implied by the verb. It provides situational guidance without excluding alternatives, which earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calendar_delete_eventDelete a calendar eventA
DestructiveIdempotent

Delete an event. This is IRREVERSIBLE from here — there is no undo and no bin. The event is read before deletion so the answer says what was removed. Only do this when the user has clearly asked for it.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
event_idYesEvent to delete.
calendar_idNoCalendar id. Defaults to "primary", the account own calendar.
send_updatesNoWhether Google emails the attendees. Defaults to "none": nobody is mailed.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing that deletion is IRREVERSIBLE with no undo or bin, and that the event is read before deletion so the answer indicates what was removed. The destructiveHint annotation is reinforced with concrete behavioral consequences, adding real value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted words. The critical irreversibility warning is front-loaded, followed by the pre-read behavior and the invocation condition. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with 4 parameters and no output schema, the description covers the essential behavioral context: irreversibility, what the response will indicate, and the condition for calling. Combined with the annotations and fully documented schema, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all four parameters. The description does not add parameter-specific detail beyond the schema, which is acceptable and earns the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Delete an event.' The action is unambiguous and distinct from sibling tools like calendar_update_event or calendar_get_event. The first sentence directly matches the title, leaving no doubt about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context: 'Only do this when the user has clearly asked for it.' This is an explicit condition for invocation, especially important for a destructive action. It does not name alternatives or when-not-to-use scenarios, but the guidance is sufficient and clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calendar_find_free_timeFind free timeA
Read-onlyIdempotent

Find gaps that are free on EVERY account asked about, within working hours. "time_min" and "time_max" must carry an explicit UTC offset. Working hours are interpreted in "timezone" (IANA), which defaults to the zone this server runs in and is always echoed back. IMPORTANT: if any calendar cannot be read, NO slots are returned and "incomplete" is set — a gap computed without one person's calendar is a double-booking, not an answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountsNoRestrict to these accounts. Omit to use every configured account.
time_maxYesLatest instant to consider. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00".
time_minYesEarliest instant to consider. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00".
timezoneNoIANA zone for the working-hours bounds, e.g. "Europe/Madrid".
calendar_idNoCalendar id. Defaults to "primary", the account own calendar.
workday_endNoEnd of the working day, "HH:MM". Default 18:00.
weekdays_onlyNoSkip Saturdays and Sundays. Defaults to true.
workday_startNoStart of the working day, "HH:MM". Default 09:00.
minimum_minutesNoIgnore gaps shorter than this. Default 30.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, but the description adds crucial behavior beyond them: if any calendar is unreadable, NO slots are returned and 'incomplete' is set, preventing double-bookings. It also discloses the timezone default-and-echo behavior, which is valuable operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler: the first states the core function, the second gives hard input constraints, and the third highlights a critical safety behavior. The IMPORTANT warning is front-loaded enough to prevent misuse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers core behavior, defaults, working-hours constraints, and failure semantics, while parameter details are fully present in the schema. The only meaningful omission is a precise description of the successful output structure beyond 'gaps' and 'incomplete,' but the annotation suite and schema make this a fairly complete definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by emphasizing that time_min/time_max must carry an explicit UTC offset and by revealing the timezone default (server zone) plus the echo behavior, which the schema does not specify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource combination: find gaps that are free on EVERY account within working hours. This clearly differentiates the tool from event-listing siblings like calendar_list_events or calendar_get_event by emphasizing cross-account availability search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use the tool: when you need common free time across multiple accounts within working hours. It does not explicitly name an alternative or state a when-not-to-use condition, but the context is strong enough for an agent to select it over the event-centric calendar siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calendar_get_eventRead a calendar eventA
Read-onlyIdempotent

Read one event in full: description, location, attendees and their responses, and the link to open it.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
event_idYesEvent id, as returned by calendar_list_events.
calendar_idNoCalendar id. Defaults to "primary", the account own calendar.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld, and non-destructive behavior. The description adds useful return-content context by specifying exactly what is included in a full event read, which is valuable because no output schema is provided. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One crisp sentence that front-loads the action and resource, then enumerates the relevant return contents without any filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read operation with annotations covering safety and idempotency, the description and schema together are sufficient. The only minor gap is that the exact response shape is not fully specified, but the description's enumeration of returned fields largely compensates for the absent output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with sensible descriptions for account, event_id, and calendar_id including the default for calendar_id. The tool description adds no parameter-specific details, but the baseline of 3 is appropriate since the schema already carries that burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and resource ('one event') and clearly defines scope: description, location, attendees and their responses, and a link. This distinguishes it from calendar_list_events and other calendar operations without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the correct use case: retrieving the full content of a single event rather than a summary list or a mutation. It does not explicitly name alternatives like calendar_list_events, but the 'one event in full' framing gives clear contextual guidance with no exclusions needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calendar_list_eventsList calendar eventsA
Read-onlyIdempotent

List events in a time range. Omit "account" to look at EVERY connected calendar at once, merged and sorted by start time. Repeating events are expanded into their real occurrences. Times must carry an explicit UTC offset.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoFree text over summary, description, location and attendees.
accountNoRestrict to one account. Omit to use every configured account at once.
time_maxYesEnd of the range. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00".
time_minYesStart of the range. ISO 8601 WITH an explicit UTC offset or "Z" — for example "2026-09-10T09:00:00+02:00".
calendar_idNoCalendar id. Defaults to "primary", the account own calendar.
max_resultsNoMaximum events to return per account. Default 25.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this read-only, idempotent, and non-destructive. The description adds valuable behavioral context beyond those hints: recurrence expansion, cross-account merging and sorting, and the explicit UTC-offset requirement for timestamps. It does not contradict any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the core action, the cross-account behavior, and the two critical gotchas (recurrence expansion and UTC offsets). The most important behavioral details are front-loaded and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only listing tool with 100% schema coverage and helpful annotations, the description covers the essential operational nuances an agent needs: time-range usage, account omission behavior, recurrence handling, and timestamp format requirements. The absence of an output schema is mitigated by the clear 'list events' framing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mostly reinforces what the schema already says about 'account' and time parameters, though it adds the merged-and-sorted behavior and the recurrence-expansion detail. This is useful but does not substantially enhance parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List events in a time range,' with concrete behavioral details like merging all calendars, sorting by start time, and expanding repeating events. This clearly distinguishes it from sibling tools such as calendar_get_event, which targets a single event, and calendar_find_free_time, which searches availability.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear contextual usage guidance: it explains how to query across every connected calendar by omitting 'account' and emphasizes the required time-range format with UTC offsets. It does not explicitly name alternatives or exclusion conditions, so it stops short of a full when-to-use versus when-not-to-use guide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calendar_update_eventUpdate a calendar eventA
Idempotent

Change an existing event. Only the fields you give are touched — except "start" and "end", which must be given together or not at all, because moving one end alone reshapes the meeting. Passing "attendees" REPLACES the guest list rather than adding to it.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNoNew end. Must be given with "start".
startNoNew start. Must be given with "end".
accountYesConfigured account: the email address, or the alias given at setup.
all_dayNoTreat the new start and end as plain dates.
summaryNoNew title.
event_idYesEvent to change.
locationNoNew location.
attendeesNoReplacement guest list, in full.
time_zoneNoIANA zone stored with the event.
calendar_idNoCalendar id. Defaults to "primary", the account own calendar.
descriptionNoNew body.
send_updatesNoWhether Google emails the attendees. Defaults to "none": nobody is mailed.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation set says the tool is not read-only and is idempotent, but the description goes beyond annotations by disclosing a crucial behavioral nuance: the update is partial, and passing attendees replaces the full guest list. It also explains why start and end must be paired. This adds real value beyond the structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is three sentences with zero filler. The most important fact ('only the fields you give are touched') is front-loaded, and the start/end pairing and attendee replacement exceptions are stated precisely without redundant elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with 12 parameters, the description covers the critical behaviors that could cause incorrect parameter usage: partial updates, paired temporal fields, and attendee list replacement. The remaining parameters are adequately documented by the complete schema, and no output schema is expected for an update operation. It is comprehensive enough without being verbose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 12 parameters with 100% coverage, so the baseline is 3. The description adds cross-parameter meaning: the partial-update behavior for all fields and the attendee replacement semantics. While some of this is echoed in the schema, the description frames it as tool-level behavior, which is useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Begins with a specific verb-plus-resource statement ('Change an existing event'), clearly distinguishing this from calendar_create_event and calendar_delete_event. It also adds partial-update semantics ('Only the fields you give are touched'), which sharpens what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly establishes the context for use — modifying an existing event — and adds important constraints such as start/end pairing and attendee replacement. It does not explicitly name sibling alternatives or state when not to use it, but the usage context is concrete enough for an agent to select it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contacts_createCreate a contactA

Add a person to the account address book. At least one field is required. Google does not check for duplicates — search first if the person might already be there.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoFree-text note stored on the contact.
emailsNoEmail addresses, in full. On update this REPLACES the existing list.
phonesNoPhone numbers, in full. On update this REPLACES the existing list.
accountYesConfigured account: the email address, or the alias given at setup.
job_titleNoRole within that organisation.
given_nameNoFirst name.
family_nameNoSurname.
organizationNoCompany or organisation.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal a write operation (readOnlyHint=false) and non-idempotency (idempotentHint=false). The description adds concrete behavioral context beyond that: Google performs no duplicate check, so repeated calls create duplicate contacts, and 'at least one field is required' states an input constraint. This enriches what annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three terse sentences, each earning its place: the core action, the required-input constraint, and the duplicate warning. The most important information is front-loaded and there is no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter write tool with no output schema, the description covers the key operational risks (duplicates, minimum fields) and the schema documents every parameter with 100% coverage. The only minor gap is that it does not address account context or the return value, but for a create action with no output schema this is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds one genuinely new semantic rule not derivable from the schema: despite only 'account' being formally required, at least one additional contact field must be supplied. This cross-parameter constraint adds value beyond the JSON Schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Add a person to the account address book.' It also scopes the tool to the 'account' address book, which combined with the sibling names (contacts_search, contacts_get, contacts_list, contacts_update) clearly distinguishes creation from the other contact operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance: 'Google does not check for duplicates — search first if the person might already be there.' This tells the agent to consult search before creating, which is clear when-to-use context. It does not name the sibling tool (contacts_search) explicitly, but the directive is actionable and sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contacts_getRead a contactA
Read-onlyIdempotent

Read one contact in full: every email and phone with its label, organisation, job title, postal addresses, links and notes.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
resource_nameYesContact id as returned by contacts_search or contacts_list, e.g. "people/c123456".

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds value by specifying the return contents (every email, phone, label, organisation, job title, postal addresses, links, notes), which is behaviorally relevant and not repeated in annotations. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly written sentence. It is front-loaded with the core purpose ('Read one contact in full') and then efficiently lists the return fields. There is no filler, and every phrase contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 parameters, no output schema), the description covers what is needed: it explains the scope ('in full') and the specific fields returned. It does not list error conditions or explicitly state the absence of side effects, but annotations already cover safety, and the field list is adequate for an agent to understand the returned data. Overall, it is sufficiently complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides full descriptions for both parameters (account and resource_name), and the schema coverage is 100%. The description does not add additional parameter-level meaning beyond what the schema already states (e.g., it does not clarify the 'full' return structure or the role of each parameter). Per the baseline rule for high schema coverage, a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action explicitly ('Read one contact in full') and enumerates the specific data fields returned (emails, phones, labels, organisation, job title, addresses, links, notes). This clearly distinguishes it from sibling tools like contacts_search (which searches) and contacts_list (which lists), making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: for a single known contact, use this tool. However, it does not explicitly state when to prefer it over contacts_search or contacts_list, nor does it mention any exclusions or prerequisites (e.g., that a resource_name is required). The guidance is implicit rather than explicit, so it falls short of a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contacts_listList contactsA
Read-onlyIdempotent

Page through an account address book, most recently changed first. Returns "nextPageToken" when there is more; pass it back as "page_token" for the next page. Use contacts_search when you are looking for someone in particular.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
page_sizeNoContacts per page. Default 50, maximum 200.
page_tokenNoThe "nextPageToken" from the previous call. Omit for the first page.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false; the description adds behavioral detail beyond that: ordering by most recently changed and the token-based pagination contract. It does not cover rate limits or auth, but those are not essential given the annotations and the account parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core behavior, then token mechanics, then the sibling alternative. No filler or redundant clauses.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only paged list tool, the description plus high-coverage schema and annotations is enough to invoke correctly. It covers ordering, pagination, and the search alternative; the absent output schema is not a blocker because the return shape is not needed to make the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so account, page_size, and page_token are already documented. The description mainly reinforces the page_token round-trip, which is already in the schema, so it adds little parameter meaning beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific action (paging through the address book) and resource (account address book), and adds the sort order. It also names contacts_search as the sibling for targeted lookup, so an agent can differentiate the tools without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to prefer contacts_search ('looking for someone in particular'), implying contacts_list is for browsing or enumerating contacts. It also frames the pagination workflow, making clear this is the paging tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contacts_updateUpdate a contactA
Idempotent

Change an existing contact. ⚠️ Each field you name is REPLACED, not merged: passing "emails" with one address removes every other address that contact had. Fields you do not name are left alone. Read the contact first with contacts_get and send back the full list of whichever field you are editing. The current version is re-read before writing, so a change made elsewhere in the meantime cannot be silently overwritten.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoFree-text note stored on the contact.
emailsNoEmail addresses, in full. On update this REPLACES the existing list.
phonesNoPhone numbers, in full. On update this REPLACES the existing list.
accountYesConfigured account: the email address, or the alias given at setup.
job_titleNoRole within that organisation.
given_nameNoFirst name.
family_nameNoSurname.
organizationNoCompany or organisation.
resource_nameYesContact id as returned by contacts_search or contacts_list, e.g. "people/c123456".

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=true, destructiveHint=false. The description goes far beyond these: it warns that named fields are REPLACED not merged, advises reading the contact first, and explains the optimistic concurrency mechanism. No contradiction with annotations; it adds critical behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single well-organized paragraph, front-loading the critical warning about replacement semantics. Each sentence adds essential value, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an update tool with replace semantics, the description covers all key aspects: the replacement warning, the need to read first, the concurrent write protection, and the handling of omitted fields. It does not need to explain the return format (no output schema) and the sibling tools are clear enough. The description is complete for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all parameters, but the description adds crucial semantics beyond the schema, particularly the replacement behavior for emails and phones and the 'full list' requirement. This clarifies usage of parameters in a way the schema alone does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Change an existing contact' with a specific verb and resource, clearly distinguishing it from siblings like contacts_create and contacts_get. It also mentions the field-level replacement semantics, which differentiates it from a generic update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs the agent to read the contact first via contacts_get and send back the full list of the field being edited, and warns against the merge pitfall. It also notes the re-read before writing, which guides safe concurrent usage. This is explicit when/how guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_draftCreate a draftA

Save a draft without sending it. Use this whenever the user has not explicitly asked for the message to go out.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoCarbon-copy recipients.
toYesRecipients. Plain addresses or "Name <a@b.c>".
bccNoBlind carbon-copy recipients.
bodyYesMessage body. Plain text unless is_html is true.
accountYesConfigured account: the email address, or the alias given at setup.
is_htmlNoSend the body as text/html instead of text/plain.
subjectYesSubject line. Non-ASCII is encoded automatically.
thread_idNoAttach to an existing thread. Prefer the "reply" tool for answering a message.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description clarifies that the draft is not sent, which is a key behavioral trait, and it complements the openWorldHint by implying side effects (draft creation). However, it does not elaborate on other potential side effects (e.g., threading behavior via thread_id) or what the draft persistence entails, leaving some uncertainty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loading the core purpose and usage condition. There is no redundant or filler content; every word earns its place. It is ideal for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters (4 required) and no output schema, the description is minimal but adequate for the core action. It does not mention return behavior, error handling, or prerequisites beyond what the schema states. The openWorldHint suggests potential side effects that are not disclosed, so the description could be more thorough, but it covers the essential intent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with each parameter described, so the baseline is met. The description adds no parameter-specific guidance beyond what the schema already provides; it simply states the overall action. It does not explain relationships between parameters (e.g., how thread_id interacts with 'to' or 'subject'), but the schema descriptions are sufficient for basic use.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair: 'Save a draft' and clearly states the non-sending behavior. It distinguishes itself from send_message by explicitly noting when to use it ('whenever the user has not explicitly asked for the message to go out'), which differentiates it from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit when-to-use condition ('Use this whenever the user has not explicitly asked for the message to go out'), which is clear and actionable. It does not name alternative tools like send_message or reply, but the implication is strong enough for an agent to infer the correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

drive_create_docCreate a Google DocA

Create a real Google Doc — not a text file — from plain text. Line breaks are kept; formatting is not, because the content is converted from plain text on the way in. Returns the document id and a link to open it.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesDocument title.
accountYesConfigured account: the email address, or the alias given at setup.
contentNoInitial body text. May be empty.
parent_folder_idNoDestination folder. Defaults to the root of My Drive.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It adds genuine behavioral detail beyond annotations: line breaks are preserved, formatting is lost, content is converted from plain text, and the return value is a document id plus open link. This is exactly the nuance needed for a creation tool and goes well beyond what readOnlyHint/destructiveHint provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler, and the key distinction ('not a text file') is front-loaded. Every sentence earns its place, including the return-value note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple creation tool with no output schema, the description covers behavior, formatting constraints, and return value. Combined with a fully documented schema, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description only loosely maps 'from plain text' to the content parameter and does not add new per-parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a real Google Doc') and explicitly contrasts it with a text file, which distinguishes it from sibling tools like drive_upload. It also tells the agent what the tool returns, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage signal: use this when you need a real Google Doc from plain text, not when you want a plain text file. It does not name an alternative sibling tool explicitly, but the 'not a text file' exclusion plus the sibling list is enough for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

drive_listList a Drive folderA
Read-onlyIdempotent

List what is inside a folder, sub-folders first. Omit "folder_id" for the root of My Drive. Use drive_search when you know what you are looking for but not where it lives.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
folder_idNoFolder id. Defaults to "root", the top of My Drive.
max_resultsNoMaximum entries to return. Default 20.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful behavioral context beyond annotations by specifying the listing order ('sub-folders first'), which an agent would not otherwise know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences carry the core behavior, the root handling, and the main sibling alternative. Every sentence earns its place, and the primary action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple non-destructive listing tool with fully documented parameterschers, the description is mostly complete. It could be slightly stronger by mentioning what the returned entries look like, but since no output schema is present and the tool name/title make the listing nature clear, the gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents account, folder_id, and max_results. The description's note about omitting folder_id for the root essentially restates the schema's default-to-root behavior, adding no new parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('List') and resource ('what is inside a folder'), and even states the traversal order ('sub-folders first'). It clearly distinguishes itself from drive_search by contrasting the use case, so an agent can tell this is the enumeration tool for a known folder.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance: omit folder_id for root, and use drive_search when you know what you are looking for but not where it lives. This directly tells the agent when to use this tool and when to prefer a sibling tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

drive_readRead a Drive fileA
Read-onlyIdempotent

Read a file as text. Google Docs are exported to plain text, Sheets to CSV and Slides to plain text; ordinary text files are downloaded as they are. Binary files (PDF, images, archives) are refused with a link instead — this tool does not download them. Long content is truncated and says so.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
file_idYesFile id, as returned by drive_search.
max_charsNoCut the content at this many characters. Default 60000.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation as read-only, idempotent, open-world, and non-destructive, but the description adds valuable behavioral context beyond those: Google format exports, binary-file refusal with a link, and truncation behavior. It explicitly says truncation is indicated in the output, which an agent needs to interpret the result correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded. The opening sentence names the operation immediately, and each subsequent sentence adds a distinct, necessary behavioral fact without repetition or padding. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description covers what the tool returns: text content, exported formats, refusal with a link for binaries, and truncation notice. Combined with annotations that cover safety and the schema that covers parameters, this is sufficiently complete for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage and already documents each parameter clearly, including file_id provenance, account aliasing, and max_chars bounds/default. The description adds no additional parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Read a file as text.' It further describes the exact transformation behavior for different file types (Google Docs to plain text, Sheets to CSV, Slides to plain text), which clearly distinguishes it from sibling tools like drive_search, drive_list, drive_upload, and drive_create_doc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly conveys when it is appropriate to use this tool: when the agent needs the textual content of a file. It also gives a when-not signal by saying binary files are refused and not downloaded. However, it does not explicitly name alternative sibling tools or provide a direct comparison, so it stops short of full exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

drive_shareShare a Drive fileA
DestructiveIdempotent

Grant someone access to a file. This GIVES AWAY ACCESS to real data and cannot be undone from here. type "user" or "group" needs "email_address"; "domain" needs "domain"; type "anyone" makes the file readable by ANYONE WITH THE LINK, so only use it when the user asked for a public link in so many words. No notification email is sent unless "notify" is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYesWhat they may do. Ownership transfer is deliberately not offered.
typeYesWho the permission is for. "anyone" means a public link.
domainNoRequired for type "domain".
notifyNoSend Google’s notification email. Defaults to false.
accountYesConfigured account: the email address, or the alias given at setup.
file_idYesFile to share.
messageNoNote included in the notification email.
email_addressNoRequired for type "user" or "group".

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint, readOnlyHint false), the description adds essential behavioral warnings: 'GIVES AWAY ACCESS to real data and cannot be undone from here', 'anyone ... readable by ANYONE WITH THE LINK', and the notification email default. This significantly exceeds what the schema and annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence contributes meaningful guidance, with the most critical warning front-loaded ('cannot be undone from here'). Despite covering several parameter nuances, the description remains tight and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given eight parameters and no output schema, the description covers the highest-risk decision points (permission type, irrevocability, notify behavior). Remaining parameters like account, file_id, role, and message are adequately documented in the schema, so little is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds valuable parameter-level context: it explains the 'type' conditional requirements, the public-link implication of 'anyone', and the notify default. This goes beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Grant someone access to a file') and is distinct from sibling tools like drive_search, drive_read, or drive_upload. The resource and verb are specific, and the title reinforces the action without being a mere tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage guidance for permission types, including when to use 'anyone' ('only use it when the user asked for a public link in so many words') and conditional requirements for user/group versus domain. It doesn't explicitly name alternative tools, but the sharing action is clear enough that no alternative appears likely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

drive_shared_drivesList shared drivesA
Read-onlyIdempotent

List the shared drives an account can reach, with their ids. A shared drive is NOT a file: it never appears in drive_search, not even searched by its exact name, so this is the only way to turn a drive name into the id that drive_list, drive_upload and drive_share need. Without "account", asks every configured account. "canAddChildren" tells you whether that account may write to it.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
max_resultsNoMaximum drives to return (per account when asking all). Default 20.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds useful behavior beyond the readOnly/idempotent annotations: omitting account queries every configured account, and the returned canAddChildren field indicates write permission. It also clarifies that shared drives cannot be discovered through drive_search.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, and every sentence earns its place. The NOT-a-file clarification, id-resolution role, and account-scoping/output-field note are all necessary and free of repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, the description still explains the returned ids and canAddChildren field, account selection behavior, and relationship to sibling tools. The main completeness gap is the unresolved conflict between the description and the required account parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description's claim that omitting account asks every configured account directly contradicts the schema's required: ['account']. This is misleading rather than purely additive. The max_results parameter is left to the schema, which is acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and object: list shared drives and return their ids. It also distinguishes itself from drive_search by clarifying shared drives are not files and never appear in search results, making its role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly explains when to use it: when you need to convert a shared drive name into an id for drive_list, drive_upload, or drive_share. It gives a concrete decision rule by noting that drive_search cannot find shared drives, and it explains account-scoping behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

drive_uploadUpload a file to DriveA

Put a file in Drive. Give either "content" for text you composed, or "local_path" for a file that already exists on the machine running this server — exactly one of the two. The file is private to the account until something shares it.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesFile name in Drive, including its extension.
accountYesConfigured account: the email address, or the alias given at setup.
contentNoInline text content.
mime_typeNoMIME type. Guessed from the file extension when omitted.
local_pathNoAbsolute path to a file on the machine running this server.
parent_folder_idNoDestination folder. Defaults to the root of My Drive.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the write nature is known. The description adds a useful privacy note ('The file is private to the account until something shares it'), which goes beyond annotations. It does not describe overwrite behavior, return format, or side effects like idempotency, but the annotations cover the basic safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The action is front-loaded, and the exclusivity constraint and privacy note are delivered efficiently. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no output schema, the description covers the main decision (content vs. path) and the default folder behavior is in the schema. However, it does not mention what the tool returns (e.g., file ID or URL), which is important for chaining operations. It also omits any mention of authentication requirements or failure modes, leaving some gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter has a description. The tool description adds critical meaning by specifying that content and local_path are mutually exclusive ('exactly one of the two'), a constraint not present in the schema. This clarifies a non-obvious relationship and aids correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb+resource ('Put a file in Drive') and immediately distinguishes two input modes (inline content vs. local file). It is specific enough to separate this from sibling tools like drive_create_doc or drive_read, and the title reinforces the action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use content versus local_path ('exactly one of the two'), which is valuable. However, it does not explicitly contrast this tool with alternatives (e.g., drive_create_doc for Google Docs) or state when not to use it. The usage context is implied but not explicit about exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

label_messageAdd or remove labelsA
Idempotent

Apply and/or remove labels on a message. Accepts label names or ids — names are resolved case-insensitively. Useful system labels: UNREAD, STARRED, IMPORTANT, SPAM, TRASH. Removing UNREAD marks a message as read.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
add_labelsNoLabels to apply.
message_idYesMessage to modify.
remove_labelsNoLabels to remove.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a non-read-only, idempotent, non-destructive mutation. The description adds real behavioral detail beyond the annotations: label names are matched case-insensitively, either names or IDs are accepted, and removing UNREAD has the side effect of marking the message read. This meaningfully improves the agent's understanding of side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each carrying distinct information: the action, label identifier flexibility, useful system labels, and the UNREAD-to-read side effect. There is no redundancy or filler, and the key action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation with four parameterscraper, the description plus schema covers the required fields, label value formats, and important side-effect semantics. It does not explicitly warn that calling it with neither add_labels nor remove_labels is a no-op, nor describe the response shape, but these are minor gaps given the idempotentHint and the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining that add_labels and remove_labels accept both label names and IDs, that names resolve case-insensitively, and that the UNREAD label has special read-marking semantics. This helps the agent construct correct label values beyond just reading the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Apply and/or remove labels on a message') and clearly distinguishes this from sibling tools like list_labels and trash_email by focusing on label mutation. It also clarifies the label namespace and case-insensitive name resolution, leaving no doubt what operation is performed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it—when a message's labels need changing—and lists useful system labels, which gives practical context. However, it never explicitly contrasts this with list_labels, trash_email, or untrash_email, nor does it state conditions when an alternative should be preferred. The guidance is adequate but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_accountsList Gmail accountsA
Read-onlyIdempotent

List every configured Gmail account and whether it is usable. Call this first when you do not know which accounts exist, or when another tool reports an unknown account.

ParametersJSON Schema
NameRequiredDescriptionDefault
probeNoContact Gmail to confirm each account really works. Slower, but detects revoked tokens that look fine on disk. Defaults to false.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false), so the bar is lower. The description adds genuine behavioral context beyond that: accounts may exist yet be unusable, and this tool reports that state. It is consistent with annotations — listing is read-only and idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Exactly two sentences: the first delivers the core purpose, the second front-loads the usage trigger. No filler, no repetition of schema or annotation content, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool (one optional param, no required params, no output schema) with annotations covering safety, the description covers purpose and usage triggers well. The only minor gap is that the return format is not described and there is no output schema to fill it, but the shape of a discovery list is largely predictable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the single "probe" parameter is fully documented in the schema (boolean, contacts Gmail, slower, detects revoked tokens, defaults false). The description only hints at this via "whether it is usable," so the schema does the heavy lifting — baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

"List every configured Gmail account and whether it is usable" names a specific verb (list), resource (configured Gmail accounts), and added scope (usability status). It is unambiguous against the sibling list_labels — that tool lists labels, not accounts — so an agent can tell them apart without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Call this first when you do not know which accounts exist, or when another tool reports an unknown account" gives explicit, actionable trigger conditions. For a discovery tool with no genuine sibling alternative, the absence of a when-not clause is not a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_labelsList labelsA
Read-onlyIdempotent

List the labels of an account, system and user-created, with their ids. Call this before label_message if you are unsure a label exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, and non-destructive behavior, so the description does not need to repeat them. It adds context about what categories of labels are returned and how results relate to label_message. It does not describe pagination or exact response shape, but those are minor for a simple label listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, with the primary action and scope in the first and usage guidance in the second. No filler, no repetition of schema content, and the information is effectively front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only listing with rich annotations, the description gives the required scope, the output contents (labels and ids), and when to call it relative to label_message. The absence of an output schema is compensated by explicitly stating that ids are included.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with the single account parameter already described as an email address or alias. The description only restates the account scoping ('of an account') and adds no new syntactic or format guidance. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List'), the resource ('labels of an account'), and the useful detail that both system and user-created labels are included with ids. The mention of 'Call this before label_message' also distinguishes it from the only sibling that operates on labels.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to call: before label_message when unsure a label exists. It points at the alternative label_message by name, giving clear routing guidance. It doesn't enumerate other alternatives, but none are needed for a pure list operation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_messageRead a messageA
Read-onlyIdempotent

Read one message in full, including its decoded body and attachment metadata. Attachment contents are not downloaded.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
message_idYesMessage id, as returned by search_emails.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and idempotent behavior, and the description adds the valuable caveat that attachment contents are not downloaded. This goes beyond the structured metadata by clarifying the scope of the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The primary purpose is front-loaded and the crucial limitation about attachments is stated immediately after the core behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, read-only tool with fully described parameters and safety annotations, the description covers the important behavior and limits. It could mention what is returned in more detail, but the absence of an output schema is mitigated by describing the body and attachment metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides full descriptions for both parameters (100% coverage). The tool description adds no parameter-specific detail, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Read one message in full') and explicitly states the scope of the result (decoded body and attachment metadata). This makes the tool easy to distinguish from siblings like search_emails and read_thread.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given about when to use this tool versus alternatives like read_thread or how it fits into a workflow after search_emails. The description implies its use but does not state when not to use it or name any companion tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_threadRead a threadA
Read-onlyIdempotent

Read a full conversation, every message in order. Prefer this over read_message when you need the context of an exchange rather than a single email.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
thread_idYesThread id, as returned by search_emails.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already establish readOnly, openWorld, idempotent, and non-destructive behavior, so the description's added value is in stating that the tool returns the complete thread ordered as a full conversation. It also communicates that the output provides exchange-level context rather than an isolated message, which goes beyond the annotations. It does not describe the response fields or pagination, but no annotation contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tightly written sentences with no filler. The first sentence states the core behavior and scope, and the second sentence adds the comparative usage guidance. Every word contributes to tool selection or invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has only two required parameters and a full schema, the description covers the main behavior and expected output order. The only gap is the lack of detail about the exact structure of returned messages, such as headers or bodies, but the absence of an output schema and the simple read-only nature keep this close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, documenting both required parameters: thread_id is described as returned by search_emailseding an email address or alias for account. The description adds no additional parameter-level details, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, 'Read', and a clear resource, 'a full conversation,' and states that it returns 'every message in order.' It explicitly contrasts with read_message by targeting the context of an exchange rather than a single email, making the tool's purpose unambiguous and distinguishing it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives direct selection guidance: 'Prefer this over read_message when you need the context of an exchange rather than a single email.' This names the alternative tool and the condition that should trigger the choice, so an agent knows when to use this tool versus read_message.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replyReply to a messageA

Send a reply that stays in the original thread. Handles the subject prefix and the In-Reply-To/References headers, which is what makes it appear as a reply rather than a new conversation. This SENDS immediately.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesMessage body. Plain text unless is_html is true.
accountYesConfigured account: the email address, or the alias given at setup.
is_htmlNoSend the body as text/html.
reply_allNoAlso copy everyone in the original Cc. Defaults to false.
message_idYesThe message being answered.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only indicate non-readonly, non-idempotent, and non-destructive behavior. The description adds valuable context beyond that: it handles subject prefixes and In-Reply-To/References headersWorld, and it emphasizes that the action sends immediately. This is critical behavioral information for an agent deciding whether to invoke a side-effecting operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences deliver the essential message efficiently: what the tool does, why it behaves as a reply, and that it sends immediately. Every sentence earns its place, and the immediate-send warning is front-loaded at the end for emphasis without being wordy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter send action with no output schema, the description covers the key behavioral risks and thread semantics. It does not mention return values, errors, or delivery guarantees, but the schema fully documents parameters and the tool is simple enough that those omissions are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter clearly. The description adds no parameter-level meaning beyond the schema, such as how body, is_html, or reply_all interact. A baseline score of 3 is appropriate because the schema carries the full semantic weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Send a reply that stays in the original thread.' It clearly distinguishes the tool from sending a new message by emphasizing thread continuity and header handling. It also differentiates from draft creation with the explicit warning 'This SENDS immediately.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: when the intent is to reply within an existing thread rather than start a new conversation. It also signals that the tool sends immediately, contrasting with a draft workflow. However, it never explicitly names alternatives like send_message or create_draft, nor states when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_emailsSearch emailsA
Read-onlyIdempotent

Search with Gmail query syntax (e.g. "from:ana@x.com is:unread newer_than:7d"). Omit "account" to search EVERY configured mailbox in parallel and get one merged, date-sorted list. If one account fails the others still return, and the failure is reported in "failures".

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesGmail search query, exactly as typed in the Gmail search box.
accountNoRestrict to one account. Omit to search all of them.
max_resultsNoMaximum messages to return (per account when searching all). Default 20.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive behavior. The description adds meaningful behavioral context beyond that: parallel search across accounts, a merged and date-sorted result, partial failure tolerance, and an explicit 'failures' field. This substantially helps an agent anticipate what happens at runtime.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The query syntax example is front-loaded, followed by the key scoping behavior and then failure semantics. Every sentence contributes essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter search tool, the description plus annotation set covers defining query syntax, account scoping, result ordering, and partial failure behavior. It does not detail the shape of the returned message objects, but the absence of an output schema is partially mitigated by the clear description of the merged date-sorted list and failures field.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds value by showing a concrete Gmail query example and emphasizing that omitting 'account' triggers a parallel search across all mailboxes. The max_results parameter is left to the schema, which documents it clearly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('search') and resource ('emails') and adds the essential Gmail query syntax. It clearly distinguishes this from sibling retrieval tools like read_message by framing it as query-based search, and the cross-mailbox behavior further pins down what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational guidance: omit 'account' to search all mailboxes, or include it to restrict the search. It describes a concrete invocation strategy and failure behavior. It does not explicitly name alternatives or exclusion criteria, but none of the sibling tools perform the same email search role.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_messageSend an emailA

Compose and SEND a new email immediately. There is no undo. If the user has not clearly asked for it to be sent, use create_draft instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoCarbon-copy recipients.
toYesRecipients. Plain addresses or "Name <a@b.c>".
bccNoBlind carbon-copy recipients.
bodyYesMessage body. Plain text unless is_html is true.
accountYesConfigured account: the email address, or the alias given at setup.
is_htmlNoSend the body as text/html instead of text/plain.
subjectYesSubject line. Non-ASCII is encoded automatically.
thread_idNoAttach to an existing thread. Prefer the "reply" tool for answering a message.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds important behavioral context beyond the annotations: the send is immediate and irreversible ('There is no undo'). This is a critical warning for a mutating side-effectful action. While it doesn't discuss failures or delivery semantics, the core risk is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the action and immediacy are front-loaded, the no-undo warning is prominent, and the alternative tool is named concisely. No filler or redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutating tool with no output schema, the description provides the necessary side-effect warning, user-intent safeguards, and routing guidance. Combined with the fully documented input schema, an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all 8 properties with meaningful descriptions, so the tool description adds little parameter-specific detail. Baseline 3 is appropriate because the schema carries the explanatory burden fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Uses a specific verb and resource ('Compose and SEND a new email immediately') and clearly differentiates from create_draft by stating that sending happens now. The title and description align, and a sibling tool shares the email domain without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly directs the agent to use create_draft when the user has not clearly asked for the message to be sent. This is a concrete conditional routing rule, which is ideal guidance for choosing between two otherwise similar sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trash_emailMove an email to the binA
DestructiveIdempotent

Move a message to the bin. This is NOT a permanent delete: the message keeps existing and untrash_email brings it back. Gmail does empty the bin by itself after 30 days, so it becomes permanent eventually — treat it as reversible for a month, not for ever. This server has no permanent-delete tool at all. The answer names the message that moved, so a wrong id is visible immediately.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
message_idYesMessage to move to the bin.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=false and destructiveHint=true, but the description adds significant behavior beyond that: it is not permanent, the message stays existing, untrash_email can restore it, Gmail auto-empties after 30 days, and a wrong ID is visible in the response. This gives the agent a much richer model of the tool's side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then adds only high-value caveats: reversibility, the 30-day limit, absence of permanent delete, and response feedback. No sentence is wasted, and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the description covers everything needed: what the tool does, its side effects, how to undo it, and what to look for in the response. It also addresses the relevant sibling untrash_email and the lack of a permanent-delete alternative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters ('Configured account' and 'Message to move to the bin'). The description adds little beyond naming the bin, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and resource: 'Move a message to the bin.' It distinguishes this from a permanent delete and from recovery via untrash_email, so an agent can immediately tell what the tool is for and how it differs from related actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the usage context clear by explaining that this is reversible for 30 days and that no permanent-delete tool exists. It doesn't literally say 'use this when you want to soft-delete and use untrash_email to undo,' but that is strongly implied by the contrast with untrash_email and the warning about Gmail's 30-day emptying.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

untrash_emailRecover an email from the binA
Idempotent

Take a message back out of the bin and restore the labels it had. Only works while the message is still there — once Gmail has emptied the bin, after about 30 days, there is nothing left to recover. Use search_emails with "in:trash" to find what is in there.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountYesConfigured account: the email address, or the alias given at setup.
message_idYesMessage to take out of the bin.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond the annotations: it restores prior labels and has a 30-day recovery window. It does not describe failure modes or authorization requirements, but the annotations already signal mutation and idempotency, and nothing contradicts the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences deliver purpose, limitation, and the discovery alternative in order. Every sentence earns its place and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter recovery tool, this description is nearly complete: what it does, when it works, and how to find eligible messages are all present. The only missing piece is the response format, but with no output schema and a side-effect-driven operation, this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and both account and message_id have meaningful schema descriptions. The tool description does not add extra format or syntax guidance for the parameters, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a concrete action, 'Take a message back out of the bin and restore the labels it had,' specifying both the resource (trashed email) and the expected effect. This clearly distinguishes untrash_email from trash_email and search_emails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States a hard precondition: only works while the message is still in the bin, with a 30-day Gmail expiration limit. It also explicitly directs the agent to search_emails with 'in:trash' to find recoverable messages, giving a clear when-to-use and when-to-look-elsewhere guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 29 tool updatesv0.3.0
    • First observedcalendar_create_event
    • First observedcalendar_delete_event
    • First observedcalendar_find_free_time
    • First observedcalendar_get_event
    • First observedcalendar_list_events
    • First observedcalendar_update_event
    • First observedcontacts_create
    • First observedcontacts_get
    • First observedcontacts_list
    • First observedcontacts_search
    • First observedcontacts_update
    • First observedcreate_draft
    • First observeddrive_create_doc
    • First observeddrive_list
    • First observeddrive_read
    • First observeddrive_search
    • First observeddrive_share
    • First observeddrive_shared_drives
    • First observeddrive_upload
    • First observedlabel_message
    • First observedlist_accounts
    • First observedlist_labels
    • First observedread_message
    • First observedread_thread
    • First observedreply
    • First observedsearch_emails
    • First observedsend_message
    • First observedtrash_email
    • First observeduntrash_email

TDQS

A4.2/5.0

Scored across 29 tools

Disambiguation5/5

Every tool maps to a distinct resource and action: search/read/draft/send/reply/label/trash for email, search/read/list/upload/share for Drive, CRUD plus free-time for Calendar, and search/get/list/create/update for Contacts. read_thread vs read_message and create_draft vs send_message are explicitly scoped to avoid confusion.

Naming Consistency4/5

Most tools follow a clear verb_noun pattern, and drive_*, calendar_*, and contacts_* prefixes make domains predictable. Minor deviations like the bare verb "reply", the unprefixed list_labels/label_message, and the noun-phrase drive_shared_drives keep it from being perfectly consistent.

Tool Count4/5

29 tools is high in absolute terms, but the server covers four distinct Google services plus multi-account handling, giving each domain a reasonable 5-8 tool surface. It is slightly heavy but every tool serves a distinct purpose and none feel redundant.

Completeness4/5

Core workflows are well covered: Gmail has search/read/draft/send/reply/label/trash, Calendar has full CRUD and free-time search, Contacts has create/read/update/search, and Drive has search/read/upload/share/create. Minor gaps exist such as no contacts_delete, no Drive delete or folder creation, and no attachment downloading, but agents can complete most real tasks.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    B
    maintenance
    Connects AI assistants to multiple Gmail accounts simultaneously, enabling search, read, draft, send, and reply operations with per-account permission controls.
    54
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to access multiple Gmail/Google Workspace accounts for searching and reading mail, calendars, and attachments.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Connects Gmail to AI assistants via the MCP protocol, enabling search, read, send, and manage emails across multiple Google accounts simultaneously.
    MIT