Sales Prep MCP
Allows reading email threads, drafting and sending follow-up emails, and performing mail merges from a spreadsheet, enabling pre-call briefs, call recaps, and outbound communication.
Provides OAuth authentication to Google services, enabling the server to access Google Calendar, Gmail, Google Sheets, Google Forms, Google Docs, and Google Drive for reading and writing data.
Enables reading calendar events to prepare pre-call briefs, schedule follow-up meetings, book demos, flag deals without meetings, and handle no-show recoveries.
Allows reading documents to extract line items from signed quotes for invoice creation.
Provides access to files for retrieving policies to prefill security questionnaires.
Enables reading form responses to qualify inbound demo requests and book meetings.
Allows reading and writing spreadsheet data for lead enrichment, prefilling security questionnaires, and mail merges with status updates.
Enables creating issues or projects for closed-won handoffs, such as setting up internal tasks.
Provides tools for creating and updating pages and databases, used for CRM updates, account research briefs, pipeline digests, and security questionnaire prefill.
Allows sending direct messages and posting to channels for pre-call briefs, demo booking alerts, closed-won notifications, pipeline digests, and trial alerts.
Enables creating payment links and invoices, and reading subscription/trial data for deal closure, invoicing, and trial management.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Sales Prep MCPprepare a pre-call brief for tomorrow's external meetings"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sales Prep and Follow-up
Pre-call briefs, follow-up emails, enriched lead sheets and payment links, from your calendar.
An MCP server with 14 workflows across Google Calendar, Tavily, Firecrawl, Gmail, Slack, Granola, Notion, Google Sheets, Google Forms, Stripe, Google Docs, Linear and Google Drive. Each workflow is a prompt your agent runs as a slash command, over the 33 tools it needs and no others.
uv tool install https://github.com/r28ai/sales-prep-mcp/releases/download/v0.1.0/sales_prep_mcp-0.1.0-py3-none-any.whl
claude mcp add sales -- sales-prep-mcpIt installs with uv from this repository's release, with no git and nothing to build; nothing but Charter and the libraries it uses comes from PyPI. To update, run the install line from the latest release. If a desktop app cannot find sales-prep-mcp, give it the full path from which sales-prep-mcp (where sales-prep-mcp on Windows).
Then ask your agent to connect your apps, or run /mcp__sales__setup.
Connect your apps
Ask the agent to connect one ("connect Linear"). It tells you where to get that app's key and the command that stores it, and the next call works, with no restart. The agent never asks for a key in the chat.
Or connect everything this server uses from a terminal:
sales-prep-mcp login # each app in turn
sales-prep-mcp login tavily # just one
sales-prep-mcp status # what is connectedTokens and keys go to your operating system's keychain (macOS Keychain, Windows Credential Manager, the Secret Service on Linux), and are checked with one read-only call to the app's own API before they are kept. Every key, token and OAuth client is yours: we register no app with any of these services, and nothing passes through a server of ours, because there isn't one.
App | How it connects | Or set |
Browser sign-in, over your own OAuth client (make one). |
| |
Tavily | Your own key (get one), entered once. |
|
Firecrawl | Your own key (get one), entered once. |
|
Slack | Your own key (get one), entered once. A bot token from your own Slack app, which the guide sets up in about three minutes. |
|
Granola | Your own key (get one), entered once. In the Granola app: Settings → Connectors → API keys. Business plan or above. |
|
Notion | Your own key (get one), entered once. Then share the pages it should see with the integration. |
|
Stripe | Your own key (get one), entered once. |
|
Linear | Your own key (get one), entered once. |
|
A variable set in your client's config always wins over the keychain.
Related MCP server: Chief of Staff
Workflows
Workflow | What you get | Apps |
Pre-call brief for tomorrow's external meetings | Company news, what they sell and the last email thread, in a DM the night before. | Google Calendar, Tavily, Firecrawl, Gmail, Slack |
Call → follow-up email and next step | The recap email is drafted, the next meeting held and the CRM updated before the rep leaves the room. | Granola, Gmail, Google Calendar, Notion |
Lead list enrichment | A column of domains becomes size, product, pricing model and a recent signal per row. | Google Sheets, Firecrawl, Tavily |
Inbound demo request → booked | A qualified request gets a meeting and a researched AE within the hour. | Google Forms, Firecrawl, Google Calendar, Gmail, Slack |
Scope agreed on a call → payment link | The deal closes while the buyer still remembers saying yes. | Granola, Stripe, Gmail |
Signed quote → Stripe invoice | The quote's line items become the invoice's, with no retyping. | Google Docs, Stripe |
Closed-won handoff | Account page, customer record and internal channel exist the day the deal closes. | Stripe, Notion, Linear, Slack |
Target account research | Every account in the target list gets a cited brief on its page. | Notion, Tavily |
Pipeline digest | Deals with no meeting booked in two weeks get flagged to their owner. | Notion, Google Calendar, Slack |
Trial ending → rep alert and nudge | Trials that end in three days get a human touch instead of a silent expiry. | Stripe, Gmail, Slack |
No-show recovery | A meeting with no notes was a no-show, and it gets a reschedule email. | Google Calendar, Granola, Gmail |
Security questionnaire, prefilled | Two hundred rows answered from your own policies, for a human to review. | Google Sheets, Notion, Google Drive |
New thread from an unknown domain → CRM | Inbound that didn't come through a form still lands in the CRM with context. | Gmail, Notion, Firecrawl |
Mail merge from a sheet | Personalised drafts per row, status written back, nothing sent without review. | Google Sheets, Gmail |
Every prompt takes one optional argument, details: the repo, team, channel, customer or date range you mean, so the agent does not have to ask. In Claude Code, put it in quotes, or only its first word arrives:
/mcp__sales__pre_call_brief_for_tomorrow_s_external_meetings "tomorrow's calls only, skip internal ones"Reads run without asking. Before anything that creates, sends, changes or deletes, the prompt tells the agent to show you the call and wait.
2 of the 14 workflows need no Google or Granola credential.
Other clients
Claude Desktop: install uv if you have not, since Claude Desktop starts the server with it, then open the .mcpb from the latest release. Claude asks for any keys in its own settings and keeps them in your keychain. The first start takes a few seconds longer, while uv installs it.
VS Code (.vscode/mcp.json): VS Code asks for each key the first time the server starts and stores it securely. Leave out any you stored with login.
{
"inputs": [
{
"type": "promptString",
"id": "google-client-secret",
"description": "Google: OAuth client secret",
"password": true
},
{
"type": "promptString",
"id": "tavily-api-key",
"description": "Tavily: API key",
"password": true
},
{
"type": "promptString",
"id": "firecrawl-api-key",
"description": "Firecrawl: API key",
"password": true
},
{
"type": "promptString",
"id": "slack-bot-token",
"description": "Slack: Bot token (xoxb-\u2026)",
"password": true
},
{
"type": "promptString",
"id": "granola-api-key",
"description": "Granola: API key",
"password": true
},
{
"type": "promptString",
"id": "notion-api-key",
"description": "Notion: Integration secret (ntn_\u2026)",
"password": true
},
{
"type": "promptString",
"id": "stripe-api-key",
"description": "Stripe: Secret or restricted key",
"password": true
},
{
"type": "promptString",
"id": "linear-api-key",
"description": "Linear: Personal API key",
"password": true
}
],
"servers": {
"sales": {
"type": "stdio",
"command": "sales-prep-mcp",
"env": {
"GOOGLE_CLIENT_SECRET": "${input:google-client-secret}",
"TAVILY_API_KEY": "${input:tavily-api-key}",
"FIRECRAWL_API_KEY": "${input:firecrawl-api-key}",
"SLACK_BOT_TOKEN": "${input:slack-bot-token}",
"GRANOLA_API_KEY": "${input:granola-api-key}",
"NOTION_API_KEY": "${input:notion-api-key}",
"STRIPE_API_KEY": "${input:stripe-api-key}",
"LINEAR_API_KEY": "${input:linear-api-key}",
"GOOGLE_CLIENT_ID": ""
}
}
}
}Cursor (.cursor/mcp.json) starts it the same way:
{
"mcpServers": {
"sales": {
"command": "sales-prep-mcp"
}
}
}Codex (~/.codex/config.toml) starts a turn without waiting for a server unless it is required, and then the agent has none of its tools. required = true makes the session wait for it, and startup_readiness = "catalog" waits for its tool list rather than just its connection:
[mcp_servers.sales]
command = "sales-prep-mcp"
required = true
startup_readiness = "catalog"
startup_timeout_sec = 30Name the server sales. A host builds each tool's name from that key, and a longer one can push a tool past the 64 characters a function name allows.
Built with Charter
Every tool here is a Charter declaration: a Pydantic schema saying where each field goes on the wire. Charter's runtime builds the request, attaches and refreshes the credential, and trims the response before the model reads it. It runs in your process, with no proxy and no telemetry.
The 33 tool schemas come to 47,240 tokens.
The same tools work in your own agent, without MCP:
from charter.adapters.openai import to_openai_tools
from charter_packs_mcp import FAMILIES
tools = FAMILIES["sales"].tools()
definitions = to_openai_tools(tools) # or charter.adapters.langchainNeed an API that isn't here? Write a pack: your coding agent writes the declarations, and Charter's conformance suite checks them.
Google Calendar:
gcalendar_events_list,gcalendar_events_insertTavily:
tavily_search,tavily_research_create,tavily_research_getFirecrawl:
firecrawl_scrape,firecrawl_extractGmail:
gmail_threads_list,gmail_drafts_create,gmail_messages_sendSlack:
slack_chat_post_message,slack_conversations_createGranola:
granola_notes_get,granola_notes_listNotion:
notion_pages_update,notion_pages_create,notion_data_sources_query,notion_searchGoogle Sheets:
gsheets_spreadsheets_values_get,gsheets_spreadsheets_values_updateGoogle Forms:
gforms_forms_responses_listStripe:
stripe_prices_list,stripe_checkout_sessions_create,stripe_customers_create,stripe_invoice_items_create,stripe_invoices_create,stripe_invoices_finalize,stripe_subscriptions_list,stripe_customers_retrieveGoogle Docs:
gdocs_documents_getLinear:
linear_customer_createGoogle Drive:
gdrive_files_list,gdrive_files_export
License
Apache 2.0.
Available Tools
35 toolsconnectA
Connect one app this server uses. For an app that issues keys, says where to get one and the terminal command that stores it. For Google, once the user's own OAuth client is set, starts the browser sign-in and returns at once: the user approves in the browser and the next call works. To see which apps are connected, call connection_status. Never ask the user for a key in the chat.
| Name | Required | Description | Default |
|---|---|---|---|
| app | Yes | The app to connect. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false and openWorldHint=false; the description adds the important behavior: for key-based apps it surfaces where to obtain the key plus the terminal command to store it, and for Google it kicks off browser sign-in and returns immediately (user approves out-of-band, next call works). This is exactly the async/side-effect context the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences, front-loaded with the core action, then mode-specific behavior, then the sibling pointer and the safety rule. Every sentence carries distinct information; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of return expectations and does so for the Google path ('returns at once... the next call works'), though what connect returns for key-issuing apps is only implied. Failure modes (invalid app, already-connected app) and the OAuth-client prerequisite for Google are mentioned but not fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum already lists the eight apps, so the baseline would be 3. The description adds value beyond the schema by tying semantics to specific enum values: key-issuing apps produce key instructions, and Google requires the user's own OAuth client to be set before the browser flow. It does not elaborate on the remaining enum members.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Connect one app this server uses'), names the sibling to use for the read-only counterpart, and differentiates the two behavioral modes (key-issuing apps vs. Google OAuth). An agent can identify this tool's job without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use routing is given: call `connection_status` to see which apps are connected, and a hard negative rule ('Never ask the user for a key in the chat'). The Google flow's async condition is also stated so the agent knows a follow-up call is expected rather than an immediate result.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connection_statusARead-only
See which apps this server is connected to, and how to connect each one that is not. Changes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered by structured data. 'Changes nothing' restates the readOnly hint rather than adding new behavior; the only incremental value is noting that connect instructions are returned for unconnected apps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the primary purpose and appends the secondary benefit with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-param, read-only status tool with no output schema, the description covers both what is inspected and the shape of the useful payload (connect guidance). Return format details are absent but minimal given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly implies no input is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('See') and resource ('which apps this server is connected to'), and adds the secondary payload of connect instructions for missing apps. This distinguishes it from the sibling 'connect' tool, which performs the connection rather than reporting status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'how to connect each one that is not' implies this tool is the discovery step before using 'connect', but the sibling is never named and there is no explicit when-to-use/when-not statement. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_extractA
Extract structured data from one or more URLs using an LLM. Poll results with extract_status.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | The URLs to extract data from. URLs should be in glob format. | |
| prompt | No | Prompt to guide the extraction process. | |
| schema | No | Schema to define the structure of the extracted data. Must conform to JSON Schema. | |
| showSources | No | When true, the sources used to extract the data will be included in the response as `sources`. | |
| ignoreSitemap | No | When true, sitemap.xml files will be ignored during website scanning. | |
| scrapeOptions | No | Options applied when scraping pages for extraction. | |
| enableWebSearch | No | When true, the extraction will use web search to find additional data. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| ignoreInvalidURLs | No | If invalid URLs are specified, they are ignored and returned in invalidURLs instead of failing the request. The server applies true when this is absent. | |
| includeSubdomains | No | When true, subdomains of the provided URLs will also be scanned. The server applies true when this is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, openWorldHint=true) are minimal and do not reveal the async job lifecycle, so the description correctly surfaces the most important behavioral trait: results must be polled via extract_status. It stops there, omitting what the initial call returns (a job id?), failure behavior, and whether web-search/credit costing applies. Decent added context over annotations, but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The core purpose is front-loaded and the polling instruction follows immediately. Nothing to trim and nothing misplaced.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, no output schema, and an async job model, the description covers the poll step but not the full lifecycle: it does not explain what the initial invocation returns or how to correlate extract_status results back to it. For a complex async tool with no output schema to fall back on, this is a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 10 parameters (urls, prompt, schema, scrapeOptions, enableWebSearch, etc.) are already documented in the schema. The description adds no parameter-level detail (e.g., the glob URL format or JSON Schema requirement) beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Extract structured data from one or more URLs') and adds the key mechanism ('using an LLM'), so the agent understands this is LLM-driven structured extraction rather than raw scraping. It does not, however, distinguish itself from the very close sibling firecrawl_scrape, leaving the boundary between the two to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence ('Poll results with extract_status') is genuine usage guidance: it tells the agent this is an async operation that must be polled. But it gives no guidance on when to choose this tool over firecrawl_scrape or other extraction options, and does not name any exclusion criteria. Usage is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_scrapeB
Scrape a single URL and optionally extract information. Use when the user wants to read or summarize a specific webpage. Supports markdown, HTML, screenshots, and structured JSON extraction.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| proxy | No | Specifies the type of proxy to use. | |
| maxAge | No | Returns a cached version of the page if it is younger than this age in milliseconds. The server applies 172800000 (2 days) when this is absent. | |
| minAge | No | When set, the request only checks the cache and never triggers a fresh scrape. | |
| mobile | No | Emulate scraping from a mobile device. | |
| actions | No | Actions to perform on the page before grabbing the content. | |
| formats | No | Output formats to include in the response. Strings or objects. The server applies markdown when this is absent. | |
| headers | No | Headers to send with the request. | |
| parsers | No | Controls how files are processed during scraping. | |
| profile | No | Persistent browser storage across scrape and interact sessions. | |
| timeout | No | Timeout in milliseconds. The server applies 60000 when this is absent. | |
| waitFor | No | Specify a delay in milliseconds before fetching the content. The server applies 0 when this is absent. | |
| blockAds | No | Enables ad-blocking and cookie popup blocking. | |
| location | No | Location settings for the request. | |
| lockdown | No | Serve from cache only and never make an outbound request. On miss, returns 404 SCRAPE_LOCKDOWN_CACHE_MISS. | |
| redactPII | No | Redact personally identifiable information from returned markdown. Pass true for defaults, or an object to tune it. | |
| excludeTags | No | Tags to exclude from the output. | |
| includeTags | No | Tags to include in the output. | |
| storeInCache | No | If true, the page will be stored in the Firecrawl index and cache. | |
| auditMetadata | No | User attribution included with SIEM logging events when SIEM is enabled. | |
| onlyMainContent | No | Only return the main content of the page excluding headers, navs, footers, etc. The server applies true when this is absent. | |
| onlyCleanContent | No | Beta. LLM pass over markdown to remove residual boilerplate that onlyMainContent can miss. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| zeroDataRetention | No | If true, this will enable zero data retention for this scrape. To enable this feature, please contact help@firecrawl.dev | |
| removeBase64Images | No | Removes all base64 images from the markdown output. | |
| skipTlsVerification | No | Skip TLS certificate verification when making requests. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the agent knows this touches external state. The description adds the supported output formats, which is useful, but it omits behavior implied by the schema — that actions (click/write/executeJavascript) mutate the page, that storeInCache writes to an external index, and that some features cost credits. Nothing contradicts the annotations, but the added behavioral detail is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core purpose front-loaded and no filler. The trailing format list is somewhat redundant with the schema's formats enum, which keeps it short of ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a 26-parameter tool with no output schema and only minimal annotations, yet the description is three sentences long. It omits cost/credit implications, caching/lockdown semantics, the relationship to firecrawl_extract, and any hint of what the response looks like, so an agent invoking it correctly still depends almost entirely on reading the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 26 parameters are already documented in-schema and the baseline is 3. The description echoes the formats dimension ('markdown, HTML, screenshots, structured JSON extraction') but adds no format syntax, precedence, or interaction detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Scrape a single URL') plus an optional outgrowth ('optionally extract information'), so an agent can tell it is a per-URL content fetcher. The phrase 'a single URL' gestures at the multi-URL alternative but never names firecrawl_extract, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one clear trigger ('Use when the user wants to read or summarize a specific webpage'), which is real usage guidance. But it never states when NOT to use it, nor does it point to firecrawl_extract for bulk/structured extraction, so the routing decision against the closest sibling is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gcalendar_events_insertC
Create a calendar event; returns details of the event.
| Name | Required | Description | Default |
|---|---|---|---|
| event | Yes | The event to insert. | |
| calendarId | Yes | Calendar identifier. To retrieve calendar IDs call the calendarList.list method. If you want to access the primary calendar of the currently logged in user, use the "primary" keyword. | |
| sendUpdates | No | Guests who should receive notifications about the change. Acceptable values are: "all" (notifications are sent to all guests), "externalOnly" (notifications are sent to non-Google Calendar guests only), "none" (no notifications are sent; for calendar migration tasks, consider using the Events.import method instead). | |
| maxAttendees | No | The maximum number of attendees to include in the response. If there are more than the specified number of attendees, only the participant is returned. Optional. | |
| eventLabelVersion | No | Version number of the event label feature supported by the API client. Version 0 assumes no event label support and processes the colorId field for color management. Version 1 enables support for event labels, and processes the eventLabelId in the event's body. In this case, the colorId field is ignored. The default is 0. Acceptable values are 0 to 1, inclusive. | |
| supportsAttachments | No | Whether API client performing operation supports event attachments. Optional. The default is False. | |
| conferenceDataVersion | No | Version number of conference data supported by the API client. Version 0 assumes no conference data support and ignores conference data in the event's body. Version 1 enables support for copying of ConferenceData as well as for creating new conferences using the createRequest field of conferenceData. The default is 0. Acceptable values are 0 to 1, inclusive. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false (a mutation) and openWorldHint=true. The description adds only 'returns details of the event' and omits meaningful behavioral context for a mutation tool of this complexity: guest notifications, the sendUpdates/conferenceDataVersion/supportsAttachments side-effect flags, permission requirements, and attendee-invitation behavior are all undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. It is efficient, though the trailing return-value clause could arguably be dropped since it is the only content beyond the verb.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the brief note about returned details is helpful, but for a mutation tool with 7 parameters and rich side-effect potential (invitations, notifications, conference generation) the description is thin. It is minimally adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema fully documents all 7 parameters including sendUpdates, conferenceDataVersion, and maxAttendees. Per the rubric, that establishes a baseline of 3; the description adds nothing beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a calendar event'), which clearly distinguishes it from the sibling gcalendar_events_list. It does not, however, differentiate itself from similar write operations or mention scope (which calendar), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use context, prerequisites, or alternatives. An agent must infer that this is the creation counterpart to gcalendar_events_list without any explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gcalendar_events_listCRead-onlyIdempotent
List events matching a given search filter.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Free text search terms to find events that match these terms in the following fields: summary, description, location, attendee's displayName, attendee's email, organizer's displayName, organizer's email, workingLocationProperties.officeLocation.buildingId, workingLocationProperties.officeLocation.deskId, workingLocationProperties.officeLocation.label, workingLocationProperties.customLocation.label. These search terms also match predefined keywords against all display title translations of working location, out-of-office, and focus-time events. Optional. | |
| iCalUID | No | Specifies an event ID in the iCalendar format to be provided in the response. Optional. Use this if you want to search for an event by its iCalendar ID. | |
| orderBy | No | The order of the events returned in the result. Optional. The default is an unspecified, stable order. Acceptable values are: "startTime" (order by the start date/time, ascending; this is only available when querying single events, i.e. the parameter singleEvents is True), "updated" (order by last modification time, ascending). | |
| timeMax | No | Upper bound (exclusive) for an event's start time to filter by. Optional. The default is not to filter by start time. Must be an RFC3339 timestamp with mandatory time zone offset, for example, 2011-06-03T10:00:00-07:00, 2011-06-03T10:00:00Z. Milliseconds may be provided but are ignored. If timeMin is set, timeMax must be greater than timeMin. | |
| timeMin | No | Lower bound (exclusive) for an event's end time to filter by. Optional. The default is not to filter by end time. Must be an RFC3339 timestamp with mandatory time zone offset, for example, 2011-06-03T10:00:00-07:00, 2011-06-03T10:00:00Z. Milliseconds may be provided but are ignored. If timeMax is set, timeMin must be smaller than timeMax. | |
| timeZone | No | Time zone used in the response. Optional. The default is the time zone of the calendar. | |
| pageToken | No | Token specifying which result page to return. Optional. | |
| syncToken | No | Token obtained from the nextSyncToken field returned on the last page of results from the previous list request. It makes the result of this list request contain only entries that have changed since then. All events deleted since the previous list request will always be in the result set and it is not allowed to set showDeleted to False. There are several query parameters that cannot be specified together with nextSyncToken to ensure consistency of the client state. These are: iCalUID, orderBy, privateExtendedProperty, q, sharedExtendedProperty, timeMin, timeMax, updatedMin. All other query parameters should be the same as for the initial synchronization to avoid undefined behavior. If the syncToken expires, the server will respond with a 410 GONE response code and the client should clear its storage and perform a full synchronization without any syncToken. Optional. The default is to return all entries. | |
| calendarId | Yes | Calendar identifier. To retrieve calendar IDs call the calendarList.list method. If you want to access the primary calendar of the currently logged in user, use the "primary" keyword. | |
| eventTypes | No | Event types to return. Optional. This parameter can be repeated multiple times to return events of different types. If unset, returns all event types. Acceptable values are: "birthday" (special all-day events with an annual recurrence), "default" (regular events), "focusTime" (focus time events), "fromGmail" (events from Gmail), "outOfOffice" (out of office events), "workingLocation" (working location events). | |
| maxResults | No | Maximum number of events returned on one result page. The number of events in the resulting page may be less than this value, or none at all, even if there are more events matching the query. Incomplete pages can be detected by a non-empty nextPageToken field in the response. By default the value is 250 events. The page size can never be larger than 2500 events. Optional. | |
| updatedMin | No | Lower bound for an event's last modification time (as a RFC3339 timestamp) to filter by. When specified, entries deleted since this time will always be included regardless of showDeleted. Optional. The default is not to filter by last modification time. | |
| showDeleted | No | Whether to include deleted events (with status equals "cancelled") in the result. Cancelled instances of recurring events (but not the underlying recurring event) will still be included if showDeleted and singleEvents are both False. If showDeleted and singleEvents are both True, only single instances of deleted events (but not the underlying recurring events) are returned. Optional. The default is False. | |
| maxAttendees | No | The maximum number of attendees to include in the response. If there are more than the specified number of attendees, only the participant is returned. Optional. | |
| singleEvents | No | Whether to expand recurring events into instances and only return single one-off events and instances of recurring events, but not the underlying recurring events themselves. Optional. The default is False. | |
| showHiddenInvitations | No | Whether to include hidden invitations in the result. Optional. The default is False. | |
| sharedExtendedProperty | No | Extended properties constraint specified as propertyName=value. Matches only shared properties. This parameter might be repeated multiple times to return events that match all given constraints. | |
| privateExtendedProperty | No | Extended properties constraint specified as propertyName=value. Matches only private properties. This parameter might be repeated multiple times to return events that match all given constraints. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that — no pagination behavior, no sync-token semantics, no note that results are scoped to a single calendar. It contributes no behavioral context of its own.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with zero waste, and the action is front-loaded. But the brevity is under-specification rather than disciplined conciseness — it omits any usable context an agent would want for a 18-parameter list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with 18 parameters, no output schema, and paging/sync behavior, a single sentence is insufficient. The rich schema mitigates some of this, but the description supplies no operational framing (single-calendar scope, pagination, recurring-event expansion) that an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 18 parameters, so the schema carries full semantic burden (time bounds, orderBy constraints, syncToken exclusions, etc.). The description adds no parameter meaning whatsoever, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (events) plus a scope qualifier ('matching a given search filter'). However, 'search filter' is vague against an 18-parameter schema and the description offers no differentiation from the sibling gcalendar_events_insert or any hint of what the tool actually filters on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives, no mention of prerequisites (calendarId required, calendarList.list to discover IDs), and no exclusions. The single sentence gives only the bare action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gdocs_documents_getARead-onlyIdempotent
Read a document's full structural content. The document ID is the long string in its URL, between '/d/' and '/edit'. Pass include_tabs_content=true for a document with several tabs, since the default response covers only the first.
| Name | Required | Description | Default |
|---|---|---|---|
| documentId | Yes | The ID of the document to retrieve. This is the long string in the document's URL, between `/d/` and `/edit`. | |
| includeTabsContent | No | When true, content is returned in the `tabs` field, covering every tab in the document. When false or omitted, only the first tab's content is returned, in the legacy top-level `body` field. | |
| suggestionsViewMode | No | How to render suggested edits. Defaults to DEFAULT_FOR_CURRENT_ACCESS, which shows suggestions inline if the caller may see them. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered structurally. The description adds a genuinely non-obvious behavioral default: the default response covers only the first tab, which would silently truncate output for multi-tab documents. It does not mention auth scopes, payload size, or rate limits, so it is not fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose, then invocation details. It loses a point because the documentId URL-parsing sentence exactly duplicates the schema description and earns nothing new.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only get with no output schema, the description covers what the operation does and the one behavioral trap (single-tab default) that would cause an agent to return incomplete content. Nothing further is required to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters, including the tabs/body return-shape distinction for includeTabsContent. The description's documentId explanation is a verbatim restatement of the schema and its tab note largely duplicates the schema text, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('a document's full structural content'), which is enough to separate it from gdrive_files_export, gsheets_spreadsheets_values_get and granola_notes_get. No ambiguity about what is fetched.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The sentence about passing include_tabs_content=true for multi-tab documents is conditional guidance on invocation, which is useful. However, there is no guidance on when to choose this tool over sibling read tools (gdrive_files_export, gsheets values get), or any stated prerequisites beyond the ID format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gdrive_files_exportARead-onlyIdempotent
Export a Google Workspace document to the requested MIME type and return the exported content. Limited to 10 MB. Common MIME types: text/plain, text/html, text/csv, application/pdf.
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | The ID of the file. | |
| mimeType | Yes | Required. The MIME type of the format requested for this export. For a list of supported MIME types, see Export MIME types for Google Workspace documents. Common values: `text/plain`, `text/html`, `text/csv`, `application/pdf`, `application/vnd.openxmlformats-officedocument.wordprocessingml.document`, `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds a genuinely useful behavioral constraint — the 10 MB size limit — and notes that content is returned rather than a link, which is meaningful beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and outcome. The trailing MIME type list is somewhat redundant with the schema, but the whole thing is tight and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description states that the exported content is returned and flags the 10 MB ceiling, which is the key expectation-setting fact. Combined with 100% schema coverage and clear annotations, an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented in the schema, including a longer MIME type list identical to the one the description gives. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export) and resource (Google Workspace document) plus the outcome (returns exported content to a requested MIME type). It's clearly distinguishable in kind from list/read siblings like gdrive_files_list or gdocs_documents_get, though it never explicitly names an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what export does but gives no when-to-use guidance, prerequisites (e.g., this only works on native Google Workspace docs, not binary files), or conditions that would route an agent here versus gdocs_documents_get. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gdrive_files_listARead-onlyIdempotent
List the user's files. Returns all files by default, including trashed files; add trashed = false to q to hide them. q is Drive's search grammar — name contains 'Q3' and mimeType = 'application/vnd.google-apps.folder'.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | A query for filtering the file results. For supported syntax, see Search for files and folders. This method returns all files by default, including trashed files. If you don't want trashed files to appear in the list, use `trashed = false` in `q`. | |
| spaces | No | A comma-separated list of spaces to query within the corpora. Supported values are `drive` and `appDataFolder`. If omitted, the server queries the `drive` space. | |
| corpora | No | Specifies a collection of items (files or documents) to which the query applies. Supported items include: `user`, `domain`, `drive`, `allDrives`. Prefer `user` or `drive` to `allDrives` for efficiency. By default, corpora is set to `user`. However, this can change depending on the filter set through the `q` parameter. If `driveId` is specified, corpora must be `drive`. | |
| driveId | No | ID of the shared drive to search. | |
| orderBy | No | A comma-separated list of sort keys. Valid keys are: `createdTime` (when the file was created; avoid using this key for queries on large item collections as it might result in timeouts or other issues; for time-related sorting on large item collections, use `modifiedTime desc` instead); `folder` (the folder ID, sorted using alphabetical ordering); `modifiedByMeTime`; `modifiedTime`; `name` (alphabetical, so 1, 12, 2, 22); `name_natural` (natural sort, so 1, 2, 12, 22); `quotaBytesUsed`; `recency`; `sharedWithMeTime`; `starred`; `viewedByMeTime`. Each key sorts ascending by default, but can be reversed with the `desc` modifier. Example usage: `folder,modifiedTime desc,name`. | |
| pageSize | No | The maximum number of files to return. The service may return fewer than this value. If unspecified, at most 100 files will be returned for shared drives, and the entire list of files for non-shared drives. The maximum value is 1000; values above 1000 will be coerced to 1000. | |
| pageToken | No | The token for continuing a previous list request on the next page. This should be set to the value of `nextPageToken` from the previous response. | |
| includeLabels | No | A comma-separated list of IDs of labels to include in the `labelInfo` part of the response. | |
| supportsAllDrives | No | Whether the requesting application supports both My Drives and shared drives. | |
| includeItemsFromAllDrives | No | Whether both My Drive and shared drive items should be included in results. | |
| includePermissionsForView | No | Specifies which additional view's permissions to include in the response. Only `published` is supported. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld, so safety is covered. The description adds a non-obvious behavioral default (trashed files are returned unless filtered) plus a worked example of the query grammar, which is real value beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the operation, then the default-scope caveat, then the query syntax. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, zero-required list tool with no output schema, the description covers the default behavior and the query parameter that most affects results. Pagination, corpora, and driveId nuances are left entirely to the (rich) schema, which is acceptable but leaves the description slightly thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description goes further by supplying a concrete `q` example (`name contains 'Q3' and mimeType = 'application/vnd.google-apps.folder'`) that the schema defers to external docs for, making the query parameter usable without leaving the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the user's files') and immediately clarifies the default scope (all files, including trashed). It does not explicitly contrast with siblings like gdrive_files_export, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives useful in-scope guidance ('add `trashed = false` to `q` to hide them') but never states when to use this list tool versus alternatives such as gdrive_files_export or gdocs_documents_get. Usage is implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gforms_forms_responses_listARead-onlyIdempotent
List a form's submitted responses, newest page first, up to 5000 per page. The only supported filter is on submission time: pass filter='timestamp >= 2026-01-01T00:00:00Z' to read what has arrived since a point in time. Answers come back keyed by questionId.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Which form responses to return. Currently, the only supported filters are: `timestamp > N` which means to get all form responses submitted after (but not at) timestamp N, and `timestamp >= N` which means to get all form responses submitted at and after timestamp N. For both supported filters, timestamp must be formatted in RFC3339 UTC "Zulu" format. Examples: "2014-10-02T15:01:23Z" and "2014-10-02T15:01:23.045123456Z". The whole filter is one string, operator included: 'timestamp >= 2014-10-02T15:01:23Z'. There is no other filterable field — a question, an email or a score cannot be filtered here. | |
| formId | Yes | Required. ID of the Form whose responses to list. | |
| pageSize | No | The maximum number of responses to return. The service may return fewer than this value. If unspecified or zero, at most 5000 responses are returned. | |
| pageToken | No | A page token returned by a previous list response. If this field is set, the form and the values of the filter must be the same as for the original request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds behavior the annotations do not: newest-first ordering, the 5000-per-page ceiling, and that answers are keyed by questionId. It omits pagination semantics (that pageToken must repeat the same filter), which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the resource and ordering, then the filter rule, then the return shape. Every sentence carries information an agent needs and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description partially compensates by noting answers are keyed by questionId, plus it gives ordering and page limits. It stops short of describing the response envelope (e.g. nextPageToken, responseId) or how to continue paging, which is the remaining gap for a paginated list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including the full filter grammar, RFC3339 format and the pageToken consistency rule. The description's filter example and questionId note mostly restate that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List a form's submitted responses') plus ordering ('newest page first') and a hard cap ('up to 5000 per page'). No other Google Forms tool exists among the siblings, so there is nothing to disambiguate against, and the purpose is unambiguous on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operative constraint clearly: the only supported filter is on submission time, with a worked example that shows what the filter buys you ('read what has arrived since a point in time'). It does not state when *not* to use it or how pagination resumes, but for a single-purpose list tool the context given is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_drafts_createC
Save an email draft to Gmail.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The draft to create. | |
| userId | No | The user's email address. The special value 'me' can be used to indicate the authenticated user. | me |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/non-read nature is covered structurally. The description adds nothing beyond that: it does not say the draft is not sent, whether creation is idempotent, or what auth is required, so the behavioral burden is largely unmet.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is well structured. It is arguably under-specified rather than bloated, but as a size/structure judgment it is tight and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with only minimal annotations and no output schema, the description should at least clarify that it saves rather than sends and hint at the returned draft. The rich input schema compensates for parameters, but the core behavioral distinction is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the nested Message/Draft/EmailContent fields are richly documented (threadId rules, bodyHtml multipart behavior, in_reply_to threading). The description adds no parameter meaning at all, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Save an email draft') and destination (Gmail), which is enough to know it creates a draft rather than sending. However, it does not distinguish itself from the adjacent gmail_messages_send sibling, so an agent gets no explicit routing cue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the alternative tool (gmail_messages_send) for actually delivering mail. The agent must infer that this only persists a draft.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_messages_sendC
Send an email via the Gmail API.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The email message data. | |
| userId | No | The user's email address. The special value 'me' can be used to indicate the authenticated user. | me |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the safety profile is partially covered. The description adds no behavioral traits beyond that, such as immediate sending, authentication requirements, irreversibility, or threading behavior; it essentially restates the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded sentence with no wasted words. While it is sparse, the structure is efficient and the core action is stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with open-world annotations and no output schema, the description is incomplete. It does not clarify that the message is sent immediately rather than saved as a draft, nor does it describe return behavior or error handling, leaving important context gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the nested EmailContent fields are fully documented. The description adds no parameter-level meaning beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Send) and resource (email via Gmail API), which is clearer than a generic action. However, it does not differentiate from sibling tools like gmail_drafts_create or slack_chat_post_message, so it misses the top mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It simply states the action without any contextual routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_threads_listCRead-onlyIdempotent
List Gmail threads.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Only return threads matching this Gmail search query string. | |
| userId | No | The user's email address. The special value 'me' can be used to indicate the authenticated user. | me |
| labelIds | No | Return only threads with all of these label IDs. | |
| pageToken | No | Page token to retrieve a specific page of results in the list. | |
| maxResults | No | Maximum number of threads to return (default 100, max 500). | |
| includeSpamTrash | No | Include threads from SPAM and TRASH in the results. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered by structured data. The description adds nothing beyond that — no note on pagination behavior, result ordering, or the fact that a full mailbox scan may be needed without filters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler and the key information front-loaded. It is efficient, though the extreme brevity shades into under-specification rather than pure conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter read tool with no output schema, the description is minimally viable: the schema covers all inputs, but the description omits pagination semantics and what a thread result contains. Nothing is misleading, but an agent gets no help beyond the structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (q, userId, labelIds, pageToken, maxResults, includeSpamTrash) is already documented in the schema. The description contributes no additional meaning, which is the baseline 3 case when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('Gmail threads'), which is unambiguous and distinguishable from write-oriented siblings like gmail_messages_send and gmail_drafts_create. It stops short of scope details (e.g. mailbox-wide vs. label-filtered), so it is clear but not maximally informative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no mention of prerequisites or the conditions under which a caller should prefer it. The sibling set contains other Gmail operations but the description offers no routing signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
granola_notes_getARead-onlyIdempotent
Read one meeting note: its AI summary, the people who attended, the calendar event it was taken against, and the folders it belongs to. Pass include='transcript' to get the transcript inline as well. A transcript too large to inline answers 413 TRANSCRIPT_TOO_LARGE. Read it with notes_transcript_get instead.
| Name | Required | Description | Default |
|---|---|---|---|
| noteId | Yes | The ID of the note, as returned by the list endpoint — a `not_` prefix followed by fourteen alphanumeric characters. The UUID in a Granola web app URL is a different identifier and is not accepted here. | |
| include | No | Include the note transcript in the response. If it is too large to return inline, Get Note returns `TRANSCRIPT_TOO_LARGE`; retrieve it in pages from `/v1/notes/{note_id}/transcript`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds materially useful behavior beyond that: the 413 TRANSCRIPT_TOO_LARGE failure mode and the fallback path to notes_transcript_get, which an agent cannot infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose and followed by the opt-in parameter and its failure path. No filler and every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden, and it does so by enumerating the returned fields (summary, attendees, calendar event, folders). Combined with the error-handling note, nothing needed to call or interpret this read tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both noteId (format constraints) and include are already fully documented in the schema, making 3 the baseline. The description's 'Pass include=transcript' largely restates the schema description and its TRANSCRIPT_TOO_LARGE note, adding little syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Read one meeting note') and immediately distinguishes scope from the sibling granola_notes_list by emphasizing the singular note. It enumerates exactly what comes back (AI summary, attendees, calendar event, folders), so an agent knows the payload without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditional guidance for the transcript case ('Pass include=transcript') and routes to the alternative tool (notes_transcript_get) when the transcript is too large. It does not explicitly contrast with granola_notes_list, which is left implicit via 'one meeting note'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
granola_notes_listARead-onlyIdempotent
List meeting notes, filtered by when they were created or last updated and optionally narrowed to one folder and its subfolders. Returns each note's id, title, owner and timestamps, not its content. Fetch that with notes_get. Only notes that already have a generated AI summary appear here.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | The cursor to continue from | |
| folderId | No | Return notes in this folder and any of its child folders. Use the list folders endpoint to discover folder IDs. | |
| pageSize | No | Maximum number of notes to return per page. The server returns 10 when this is absent. | |
| createdAfter | No | Return notes created after this date. A date (`2026-01-27`) or a date-time (`2026-01-27T15:30:00Z`). | |
| updatedAfter | No | Return notes updated after this date. A date (`2026-01-27`) or a date-time (`2026-01-27T15:30:00Z`). | |
| createdBefore | No | Return notes created before this date. A date (`2026-01-27`) or a date-time (`2026-01-27T15:30:00Z`). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld), so the bar is lower, and the description adds real behavioral context: the response is metadata-only (id, title, owner, timestamps) and only AI-summarized notes are returned. It does not mention pagination or cursor behavior, which is a notable omission for a list tool returning partial pages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what is listed and filtered, then the return shape, then the sibling routing. Every sentence carries distinct information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields and the AI-summary precondition, which is the key thing an agent needs to know before calling. Pagination behavior is left entirely to the cursor parameter's schema description, a minor gap for a paged list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including folder recursion, pageSize default of 10, cursor, and date formats. The description only restates the existence of the time and folder filters, adding essentially nothing beyond the schema, which is the baseline-3 case for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list meeting notes) plus the two filtering axes (creation/update time, folder scope) and explicitly distinguishes the resource from its sibling by noting that content is fetched with notes_get. An agent can differentiate it from granola_notes_get without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to the right sibling for content ('Fetch that with notes_get') and discloses a decisive selection condition ('Only notes that already have a generated AI summary appear here'). It stops short of explicit when-not-to-use guidance, such as what to do if the summary requirement excludes a desired note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gsheets_spreadsheets_values_getCRead-onlyIdempotent
Returns a range of values from a spreadsheet.
| Name | Required | Description | Default |
|---|---|---|---|
| range | Yes | The A1 notation or R1C1 notation of the range to retrieve values from. | |
| spreadsheetId | Yes | The ID of the spreadsheet to retrieve data from. | |
| majorDimension | No | The major dimension that results should use. For example, if the spreadsheet data in Sheet1 is: A1=1,B1=2,A2=3,B2=4, then requesting range=Sheet1!A1:B2?majorDimension=ROWS returns [[1,2],[3,4]], whereas requesting range=Sheet1!A1:B2?majorDimension=COLUMNS returns [[1,3],[2,4]]. | |
| valueRenderOption | No | How values should be represented in the output. The default render option is FORMATTED_VALUE. | |
| dateTimeRenderOption | No | How dates, times, and durations should be represented in the output. This is ignored if valueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no mention of behavior for missing/empty ranges, error cases, or whether the range must already exist — so it contributes essentially no behavioral context of its own.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero padding. It is efficient, though arguably terse given the tool takes five parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations cover the safety profile and the schema fully documents inputs, so the description's burden is lighter. Still, with no output schema, it says only that 'values' are returned without hinting at the row/column shape that majorDimension controls, which is the main thing an agent must reason about.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the parameters (range, spreadsheetId, majorDimension, valueRenderOption, dateTimeRenderOption) are already richly documented in the schema, including an example for majorDimension. The description adds no parameter meaning at all, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns a range of values from a spreadsheet'), which is clearly readable. However, it does not differentiate this from siblings such as gsheets_spreadsheets_values_update or gdocs_documents_get, leaving the agent to infer the read-vs-write distinction from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when NOT to use it, and no reference to the sibling update tool. The agent gets an implied read-only purpose from the verb but nothing explicit about context or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gsheets_spreadsheets_values_updateC
Sets values in a range of a spreadsheet.
| Name | Required | Description | Default |
|---|---|---|---|
| range | Yes | The A1 notation of the values to update. | |
| valueRange | Yes | The request body contains an instance of ValueRange. | |
| spreadsheetId | Yes | The ID of the spreadsheet to update. | |
| valueInputOption | Yes | How the input data should be interpreted. | |
| includeValuesInResponse | No | Determines if the update response should include the values of the cells that were updated. By default, responses do not include the updated values. If the range to write was larger than the range actually written, the response includes all values in the requested range (excluding trailing empty rows and columns). | |
| responseValueRenderOption | No | Determines how values in the response should be rendered. The default render option is FORMATTED_VALUE. | |
| responseDateTimeRenderOption | No | Determines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so the write nature is known. But the description adds nothing beyond the name – it does not disclose that existing cell values are overwritten, how valueInputOption affects interpretation, or what the update returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is tight. But for a 7-parameter mutation tool it is under-specified rather than appropriately sized; conciseness here is closer to omission.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with 7 params and no output schema, the description omits overwrite semantics, auth requirements, and response behavior. An agent could call it, but not safely without reading the schema closely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with 7 well-documented parameters, so the schema carries the meaning. The description adds no parameter detail beyond 'in a range', which is baseline 3 for fully documented schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (sets values in a range of a spreadsheet), which an agent can distinguish from gsheets_spreadsheets_values_get by direction of data flow. However it offers no explicit sibling differentiation and largely restates the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the sibling gsheets_spreadsheets_values_get or when reading vs writing applies. The agent must infer context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linear_customer_createB
Create a customer. Only name is required.
| Name | Required | Description | Default |
|---|---|---|---|
| variables | Yes | The customer to create. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so the basic write and external-side-effect profile is covered. The description adds nothing beyond that: it does not describe permissions, idempotency, what record is created, rate limits, or any side effect beyond the trivial 'create' verb. With annotations present this is a low-value addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the only required-parameter note. There is no wasted text, and the information is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a rich nested schema for optional customer fields and an open-world write annotation, but the description is minimal. It gives the core action and required field, but does not clarify the Linear context, what happens to omitted optional fields (schema covers defaults), or what is returned. It is minimally adequate but leaves gaps for a creation tool with many optional properties.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and every parameter is documented in the schema. The description's only parameter-related statement ('Only name is required') duplicates the schema's required list. Baseline 3 is appropriate when the schema fully carries parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a customer'), so an agent knows what action is performed. However, it does not differentiate this tool from sibling tools like stripe_customers_create or notion_pages_create, which also create customers/pages, leaving the target system to be inferred from the tool name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use guidance, no prerequisites, and no alternatives. It only states a parameter requirement ('Only name is required'), which is not usage guidance. An agent reading this would not know when to prefer this tool over stripe_customers_create or other creation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notion_data_sources_queryB
Get the rows of a data source, optionally filtered and sorted. The filter grammar is one condition per column type, composed with and and or up to two levels deep. Read the schema first if you do not know the column names — a filter naming a column that is not there is a 400.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | The filter, sort and paging options. All optional. | |
| dataSourceId | Yes | The ID of the data source. | |
| filterProperties | No | Property IDs to return on each row, instead of all of them. The cheapest way to keep a wide table's query readable. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description frames the operation as 'Get the rows', i.e., a read, yet the annotations declare readOnlyHint=false, which tells the agent state can be mutated. That is a direct conflict about side effects. The extra context about filter nesting depth and the 400 on unknown columns is useful, but the contradiction in the safety profile dominates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, all front-loaded: purpose first, grammar constraint second, prerequisite/pitfall last. No filler and every sentence contributes information an agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested-query tool the description covers the filter grammar and one failure mode, but says nothing about pagination (pageSize/startCursor are in the schema) or what the response looks like, and there is no output schema to fall back on. It is adequate but leaves meaningful operational gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents dataSourceId, body, and filterProperties fully. The description restates the filter-composition rule (one condition per type, and/or two levels deep) that the schema already documents, adding only the 400-on-unknown-column error behavior. Baseline 3 is appropriate when the schema carries the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Get the rows of a data source', with the qualifier that results can be filtered and sorted. An agent immediately understands this is a read/query operation over tabular Notion data. It does not, however, name or contrast itself against sibling tools such as notion_search, so sibling differentiation is absent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is largely implied by 'optionally filtered and sorted', and it adds a genuine prerequisite: 'Read the schema first if you do not know the column names'. It also warns that a filter referencing a non-existent column yields a 400, which steers the agent toward a correct call. What is missing is explicit when-not-to-use or an alternative tool to prefer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notion_pages_createA
Create a page — as a subpage of another page, or as a row of a database by giving its data_source_id as the parent. Content comes as a markdown string Notion parses into blocks, or from a template: one or the other, never both.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The page to create. | |
| filterProperties | No | Property IDs to return on the page that comes back, instead of all of them. A page that does not have a listed property omits it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, establishing this as an external write. The description adds the meaningful markdown/template mutual exclusion. It does not disclose auth requirements, rate limits, or the allowAsync async-202 behavior, but with annotations carrying the safety profile a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with the verb and the two modes, with the exclusivity constraint phrased crisply. No wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex creation tool with a fully documented schema, the description covers parent selection and content sourcing adequately. It omits the async task path (allowAsync → 202) and return shape, but with no output schema and 100% schema coverage these are minor gaps, and annotations cover the write semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value: it clarifies that `parent` takes `data_source_id` to make a database row and that template vs. markdown are mutually exclusive — beyond the schema's raw field docs. It doesn't add detail on `properties` or `filterProperties`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a page') and immediately delineates the two creation modes — subpage vs. database row — which is exactly what separates it from siblings like notion_pages_update. An agent can identify the tool's scope without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance: use a page parent for a subpage, `data_source_id` for a database row, and content comes from `markdown` OR a template, 'never both.' The mutual-exclusion rule is explicit. However, it names no sibling alternatives (e.g., update vs. create routing) and states no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notion_pages_updateA
Update a page's property values, icon, cover, or trash state. Properties not named are left alone, and a value set to null is cleared. This cannot move a page — use pages_move.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The fields to change. | |
| pageId | Yes | The ID of the page. | |
| filterProperties | No | Property IDs to return on the page that comes back, instead of all of them. A page that does not have a listed property omits it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true, so the description usefully adds partial-update semantics (unnamed properties untouched, null clears a value). However, it omits the destructive `eraseContent` behavior (deletes every block on the page) and the `isArchived`/`isLocked` nuances, which matter a lot for a write tool with no annotation detail beyond the read-only flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core capability, then the partial-update rule, then the routing exclusion. Every sentence carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large, deeply nested mutation schema with no output schema and minimal annotations, the description covers the headline operations and the move exclusion, but leaves out destructive behaviors (`eraseContent`), the template/lock fields, and the `filterProperties` return-shaping parameter. An agent can call it, but not without risking unintended destruction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters in depth. The description reinforces the properties-clearing semantics ('a value set to null is cleared'), which the schema also states, and says nothing about `filterProperties`, so it adds little beyond structured data. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (page) plus the exact facets it can change (properties, icon, cover, trash state). It also explicitly carves out what it does NOT do and names the sibling (`pages_move`) that handles it, so an agent can distinguish it from related tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear exclusion: it cannot move a page, so use `pages_move` instead. It also implies the partial-update use case by saying unnamed properties are left alone. It does not, however, contrast with `notion_pages_create` or say when a caller should prefer this over other page operations, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notion_searchA
Search the titles of pages and data sources shared with this integration. Titles only — it does not search page content — and it returns nothing that has not been shared. With no query it lists everything visible, which is how you find the IDs to start from.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | The query and its options. All optional. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description frames the tool as a pure read ('Search', 'returns nothing that has not been shared', 'lists everything visible'), while the annotations declare readOnlyHint=false, i.e. the call may modify state. This is the same conflict pattern as a 'create' description paired with readOnlyHint=true, so the text and structured metadata contradict each other.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the verb and resource, then the two key constraints (titles only, shared-only), then the discovery behavior. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because the schema is fully documented and there is no output schema, the description need not explain parameters or return shape; it correctly focuses on scope and the search surface. It is nearly complete, though it says nothing about pagination limits or the read-only question raised by the annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains query, filter, pageSize, startCursor and sort. The description's 'With no query it lists everything visible' restates the query parameter's own schema text rather than adding syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource+scope: 'Search the titles of pages and data sources shared with this integration.' The clause 'Titles only — it does not search page content' pins the exact search surface, so an agent can distinguish it from content-oriented or data-source-querying siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage pattern: 'With no query it lists everything visible, which is how you find the IDs to start from' — an explicit discovery workflow. It does not name an alternative tool or state when NOT to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_chat_post_messageA
Send a message to a Slack channel, private group, or DM. Provide text for a plain message; set thread_ts to reply inside an existing thread.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | The main body text of the message. Required unless blocks or attachments are provided. Used as the fallback string for notifications when blocks are provided, so it is worth setting even then. | |
| parse | No | Change how messages are treated. Accepts 'none' or 'full'. | |
| blocks | No | A JSON-based array of structured Block Kit blocks. | |
| mrkdwn | No | Disable Slack markup parsing by setting to false. Defaults to true. | |
| channel | Yes | An encoded ID or channel name that represents a channel, private group, or IM channel to send the message to. Prefer the encoded ID (e.g. 'C123ABC456'). | |
| iconUrl | No | URL to an image to use as the icon for this message. Requires the chat:write.customize scope. | |
| metadata | No | Application-specific metadata to attach to the message. | |
| threadTs | No | Provide another message's 'ts' value to make this message a reply in that thread. Avoid using a reply's ts value; use the parent's. | |
| username | No | Set the bot's user name. Requires the chat:write.customize scope. | |
| iconEmoji | No | Emoji to use as the icon for this message, e.g. ':chart_with_upwards_trend:'. Requires the chat:write.customize scope. | |
| linkNames | No | Find and link user groups. | |
| attachments | No | A JSON-based array of structured attachments. | |
| unfurlLinks | No | Pass true to enable unfurling of primarily text-based content. | |
| unfurlMedia | No | Pass false to disable unfurling of media content. | |
| markdownText | No | Accepts message text formatted in markdown. Limit this field to 12,000 characters. Cannot be used together with blocks or text. | |
| replyBroadcast | No | Used in conjunction with thread_ts and indicates whether the reply should be made visible to everyone in the channel. Defaults to false. | |
| unfurlAppLinks | No | Pass true to enable unfurling of links to installed apps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external nature is covered. The description adds the threading behavior, but does not disclose required scopes, rate limits, message-size limits, or what a successful send returns — and most scope info already lives in the schema parameter descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and then the two most important parameter behaviors. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter tool with no output schema, the description is thin but the schema carries full parameter documentation, so an agent can call it correctly. Missing behavioral context (rate limits, required scopes, response shape) keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 17 parameters are documented in the schema itself. The description's notes on `text` and `thread_ts` largely restate what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Send') and resource ('a message to a Slack channel, private group, or DM'), making the action and destination unambiguous. An agent can distinguish this from siblings like slack_conversations_create without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives implied usage for `text` and `thread_ts`, but there is no explicit when-to-use vs. when-not, no mention of prerequisites (e.g. chat:write scope), and no reference to alternative messaging tools such as gmail_messages_send. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_conversations_createB
Create a channel. Slack refuses a name that is already taken.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The channel name, lowercase, without spaces or periods and at most 80 characters. Slack answers `name_taken` if it already exists. | |
| isPrivate | No | Create a private channel rather than a public one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so mutation and external-effect semantics are covered. The description adds a useful behavioral note about name collisions, but omits visibility defaults, permissions, and rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core action front-loaded and no filler. It is tightly sized for a two-parameter tool, though the second sentence restates the schema's name_taken note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation with no output schema and fully documented parameters, the description is adequate but thin on visibility semantics and confirmation behavior. Nothing critical is missing, but an agent gets just the minimum.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema, including the lowercase/80-char rules and the name_taken behavior. The description adds no syntax or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('Create a channel') that an agent can distinguish from siblings like slack_chat_post_message. It does not explicitly name a sibling or scope, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites (auth scopes, workspace), and no routing to alternatives. The agent must infer that this is for creating rather than posting to a channel.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_checkout_sessions_createA
Create a Checkout Session and get a hosted payment URL. Pass line_items with price IDs and quantities, mode='payment' for one-time or mode='subscription' for recurring.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | The mode of the Checkout Session. Pass 'subscription' if the session includes at least one recurring item, 'payment' for one-time payments, 'setup' to save payment details for later. | |
| uiMode | No | The UI mode of the session. Defaults to 'hosted_page'. | |
| currency | No | Three-letter ISO currency code, in lowercase. Required in 'setup' mode when payment_method_types is not set. | |
| customer | No | ID of an existing Customer, if one exists. | |
| metadata | No | Set of key-value pairs attached to the session. | |
| cancelUrl | No | If set, Checkout displays a back button and sends customers here if they cancel. | |
| expiresAt | No | The Epoch time in seconds at which the session expires. Between 30 minutes and 24 hours after creation; defaults to 24 hours. Seconds, not milliseconds: 1700000000, not 1700000000000. | |
| lineItems | No | A list of items the customer is purchasing. Required for 'payment' and 'subscription' mode. | |
| returnUrl | No | Where to send the customer after they authenticate or cancel. Required if ui_mode is 'embedded_page', 'elements' or 'form' with redirect-based methods. | |
| successUrl | No | The URL to send customers to when payment or setup is complete. Required for the default hosted flow, and not allowed if ui_mode is 'embedded_page', 'elements' or 'form'. | |
| customerEmail | No | Prefills the customer's email. If not provided, customers are asked to enter it. | |
| clientReferenceId | No | A unique string to reference the Checkout Session — a customer ID, a cart ID — used to reconcile the session with your own systems. | |
| allowPromotionCodes | No | Enables user-redeemable promotion codes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the mutation/create nature is covered. The description usefully discloses the return value (a hosted payment URL), which matters since no output schema exists, but it omits side effects, session lifecycle, or what happens on expiry/cancellation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the action/outcome is front-loaded and the key parameters follow. Nothing is redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter, high-complexity tool the description is thin, though the rich schema (100% coverage) carries the parameter burden and the hosted-URL outcome is stated. It is sufficient to call correctly but leaves session behavior and the payment-vs-setup distinction under-explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 13 parameters are already documented, and the description only echoes line_items and mode. It adds no syntax or constraints beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Create a Checkout Session') and even names the outcome ('get a hosted payment URL'), so the agent immediately understands what it produces. It does not, however, distinguish this tool from neighboring Stripe creation tools such as stripe_invoices_create or stripe_subscriptions_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete invocation guidance ('Pass line_items with price IDs and quantities', mode='payment' vs mode='subscription'), which implies usage. But it never says when to prefer this over alternatives like stripe_invoices_create, nor any prerequisites such as existing customers or prices.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_customers_createC
Create a customer.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | The customer's full name or business name. | |
| No | Customer's email address. Displayed alongside the customer in the dashboard and useful for searching and tracking. | ||
| phone | No | The customer's phone number. | |
| address | No | The customer's address. Required if calculating taxes. | |
| balance | No | An integer amount in the smallest currency unit representing the customer's starting balance. A negative amount is a credit. Cents, not dollars: $15.00 is 1500, and 15 is a fifteen-cent balance. Multiply a decimal amount by 100. | |
| metadata | No | Set of key-value pairs attached to the object, for storing additional information in a structured format. | |
| taxExempt | No | The customer's tax exemption. One of 'none', 'exempt', or 'reverse'. | |
| description | No | An arbitrary string attached to the customer. Displayed alongside the customer in the dashboard. | |
| paymentMethod | No | The ID of the PaymentMethod to attach to the customer. Accepted only when creating; use the Stripe dashboard or the PaymentMethods API to change one later. | |
| preferredLocales | No | Customer's preferred languages, ordered by preference. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true, so the description must carry the rest. It says nothing about idempotency, whether the call is reversible, what permissions or API key scope are needed, or that a customer ID is returned and needed downstream (e.g., for checkout sessions or invoices).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero waste. It is efficient, though its brevity reflects under-specification rather than disciplined conciseness, so it does not reach a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation with no output schema and only two coarse annotations, the description omits anything an agent needs beyond the schema: no return value, no error/idempotency behavior, no note that this is the prerequisite step for invoices, subscriptions, or checkout sessions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (balance units, taxExempt enum values, paymentMethod create-only behavior) is already documented in the schema. The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a customer'), so the operation is unambiguous. However, it does nothing to distinguish this from the sibling linear_customer_create or to indicate which platform's customer model is being created, which a 5 would require.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives such as stripe_customers_retrieve or linear_customer_create. The agent must infer that this is the Stripe-side entry point for creating a customer record.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_customers_retrieveBRead-onlyIdempotent
Retrieve a single customer by ID.
| Name | Required | Description | Default |
|---|---|---|---|
| customer | Yes | The identifier of the customer, e.g. 'cus_NffrFeUfNV2Hib'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, covering the safety profile. The description adds no behavioral context beyond what the annotations and the verb 'Retrieve' already imply, such as authentication needs, rate limits, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for a simple retrieval tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple read operation, full schema coverage, and annotations that cover safety traits, the description is almost complete. It does not describe the returned customer object, which is a minor gap in the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented with an example. The description adds only 'by ID', which is already conveyed by the schema's parameter name and description, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Retrieve'), resource ('customer'), and scope ('single ... by ID'). It clearly distinguishes retrieval from creation, but does not explicitly name the sibling tool (stripe_customers_create) or otherwise differentiate itself from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description implies a straightforward lookup, but gives no context about when it is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_invoice_items_createA
Add a line to an invoice. Name an existing price as pricing.price, or give an amount directly. Without invoice the line waits for the customer's next subscription invoice.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | The amount in the smallest currency unit. Negative reduces what the invoice is due. Cents, not dollars: $15.00 is 1500, and 15 puts fifteen cents on the invoice. Multiply a decimal amount by 100. | |
| period | No | The service period this line covers. | |
| invoice | No | The draft invoice to add this line to, at most 250 lines. Absent, the line waits for the customer's next subscription invoice and a standalone invoice will not collect it on its own. | |
| pricing | No | An existing price to bill, named as `pricing.price`. | |
| taxCode | No | The tax code for what is being billed. | |
| currency | No | Three-letter lowercase ISO currency code. | |
| customer | Yes | The ID of the customer to bill. | |
| metadata | No | Set of key-value pairs attached to the line. | |
| quantity | No | How many units this line bills. | |
| discounts | No | Coupons or promotion codes applying to this line alone. | |
| priceData | No | A price created inline for this line. | |
| description | No | What this line says on the invoice. | |
| taxBehavior | No | Whether the amount includes tax. Once set to inclusive or exclusive it cannot be changed. | |
| discountable | No | Whether invoice-level discounts apply to this line. | |
| subscription | No | Bill this line on that subscription's invoices only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations disclose mutation (readOnlyHint=false) and open-world reach, so the safety profile is covered. The description usefully adds lifecycle behavior — that a line without `invoice` defers to the customer's next subscription invoice — which is non-obvious and not captured by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the core action front-loaded and the deferred-invoice caveat last. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter mutation tool with full schema coverage, annotations for the mutation profile, and no output schema, the description covers purpose, the two pricing paths, and the missing-invoice behavior. The required `customer` parameter is left to the schema, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents all 15 parameters in depth (units, formats, limits). The description's `pricing.price` vs `amount` note restates what the schema says, adding little beyond it, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Add a line to an invoice" gives a specific verb and resource that clearly maps to creating a Stripe invoice item, and the pricing/amount guidance sharpens it. It does not explicitly name or rule out sibling tools such as stripe_invoices_create, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states two concrete ways to supply the line's value (name an existing price via `pricing.price`, or pass `amount` directly) and explains the conditional behavior when `invoice` is omitted. It gives clear operating context but no explicit when-to-use-this-vs-an-alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_invoices_createA
Create a draft invoice for a customer. It bills nothing until finalised, so this is the safe half of invoicing.
| Name | Required | Description | Default |
|---|---|---|---|
| footer | No | Footer text displayed on the invoice. | |
| dueDate | No | Unix timestamp when payment is due. Valid only when `collection_method` is 'send_invoice'. Seconds, not milliseconds: 1700000000, not 1700000000000. | |
| currency | No | Three-letter lowercase ISO currency code. Absent, the customer's currency. | |
| customer | Yes | The ID of the customer to bill. | |
| metadata | No | Set of key-value pairs attached to the invoice. | |
| discounts | No | Coupons and promotion codes to apply. Absent, the invoice inherits the customer's discount. | |
| onBehalfOf | No | The connected account the funds are intended for. Its branding and support information appear on the invoice. | |
| autoAdvance | No | Whether Stripe collects the invoice automatically. False leaves the invoice where it is until you act on it. | |
| description | No | An arbitrary string attached to the invoice. Shown as the memo in the Dashboard. | |
| daysUntilDue | No | Days until the invoice is due. Valid only when `collection_method` is 'send_invoice'. | |
| subscription | No | Bill this subscription. The invoice then includes only that subscription's pending items, and its billing cycle is untouched. | |
| collectionMethod | No | How to collect payment. Absent, Stripe charges the customer's default payment method. | |
| statementDescriptor | No | What the customer sees on their card statement. Must contain at least one letter. | |
| defaultPaymentMethod | No | The payment method to charge. It must belong to the invoice's customer. | |
| pendingInvoiceItemsBehavior | No | Whether to pull the customer's pending invoice items onto this invoice. Absent, Stripe excludes them and the draft is empty. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag a mutation (readOnlyHint false) and open-world scope; the description adds genuinely new behavioral context: the invoice does not bill until finalised, so creation itself is side-effect-free with respect to payment. It does not cover auth/permission needs or that a draft is empty without line items.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the key behavioral reassurance. Nothing is wasted and nothing important is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter mutation with no output schema, the description is thin: it omits the workflow detail that the draft is empty unless invoice items are added (via stripe_invoice_items_create) and says nothing about finalization as the follow-up step. The rich schema compensates on parameters, but the end-to-end workflow is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 15 parameters are already documented in the schema. The description adds no additional parameter meaning (e.g., defaults, interactions), which matches the baseline of 3 when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a draft invoice for a customer.' The phrase 'the safe half of invoicing' nicely frames it as distinct from a finalization step, though it never names the sibling (stripe_invoices_finalize) explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when this is appropriate ('it bills nothing until finalised') which nudges the agent toward creating a draft before finalizing, but it never states prerequisites, when-not-to-use, or the alternative tool by name. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_invoices_finalizeA
Finalise a draft invoice, making it open and payable. The amounts stop being editable at this point.
| Name | Required | Description | Default |
|---|---|---|---|
| invoice | Yes | The ID of the invoice. | |
| autoAdvance | No | Whether Stripe collects the invoice automatically after finalising. False leaves it open until you act on it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so the description needn't restate that this is a mutating operation. It adds genuine value beyond them by disclosing the irreversible consequence: 'The amounts stop being editable at this point.' It does not mention permission requirements or failure modes, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and effect. Every clause carries information; nothing is redundant with the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema, the description covers the essential state transition and its consequences, and the schema covers both parameters including defaults. Error behaviour and what the response returns are not addressed, but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, including a clear explanation of autoAdvance's default behaviour, so the schema already carries the parameter burden. The description adds no parameter-level detail beyond it, which is the baseline 3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Finalise') plus resource ('draft invoice') and an explicit statement of the resulting state ('making it open and payable'). This clearly distinguishes it from the sibling stripe_invoices_create, which creates rather than finalises an invoice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'a draft invoice' — the agent can infer this applies only to drafts. However, there is no explicit when-to-use guidance, no mention of prerequisites, and no routing to alternatives such as stripe_invoices_create for a non-draft case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_prices_listARead-onlyIdempotent
List prices. Filter by product to find what a product costs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | A limit on the number of objects to be returned, between 1 and 100. Defaults to 10. | |
| active | No | Only return prices that are active or inactive. | |
| product | No | Only return prices for the given product. | |
| currency | No | Only return prices in this three-letter ISO currency code. | |
| endingBefore | No | A cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after. | |
| startingAfter | No | A cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is fully covered by structured data. The description adds only the product-filter use case and discloses nothing about pagination behavior, defaults, or the shape of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with purpose and followed by the practical filter hint. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema and full parameter documentation, the essentials are reachable, but the description omits pagination guidance despite two cursor parameters and says nothing about defaults (e.g., limit 10, active filter). Adequate but with a clear gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (limit, active, product, currency, endingBefore, startingAfter) is already documented in the schema. The description only restates the product filter, adding no syntax or semantics beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('prices'), which is clear enough to distinguish it from write-oriented siblings like stripe_customers_create. However, it does not differentiate itself from other stripe list tools such as stripe_subscriptions_list beyond the resource name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It offers one implied usage pattern ('Filter by product to find what a product costs'), which gives a hint of intent, but there is no when-to-use/when-not guidance and no mention of alternatives or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_subscriptions_listBRead-onlyIdempotent
List subscriptions. Filter by customer or status.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | A limit on the number of objects to be returned, between 1 and 100. Defaults to 10. | |
| price | No | Filter for subscriptions that contain this recurring price ID. | |
| status | No | The status of the subscriptions to retrieve. Pass 'all' to return subscriptions of all statuses. | |
| customer | No | The ID of the customer whose subscriptions will be retrieved. | |
| endingBefore | No | A cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after. | |
| startingAfter | No | A cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds no behavioral context beyond that — nothing about the default status set returned, pagination cursor semantics, or rate/limit behavior — so it earns little credit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded before the filtering hint. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter list tool with no output schema and no required params, the description is minimal but workable since the schema carries full parameter documentation. It omits any note on pagination, default page size, or what a subscription object contains, which an agent would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including limit, price, and both cursors is already documented in the schema. The description only echoes customer and status, adding no syntax or format detail beyond the structured fields; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('subscriptions'), which cleanly separates it from siblings like stripe_prices_list and stripe_customers_retrieve. It does not, however, explicitly contrast itself with any sibling or state the scope of what is listed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence implies the two main filtering paths (customer, status), giving implied usage context, but there is no explicit when-to-use guidance, no mention of alternatives for narrower queries, and no note on default result behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tavily_research_createA
Create an async research task that searches, analyzes sources, and generates a cited report. Poll results with research_get.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | Attach up to 5 files as additional sources. Each file may be at most 80,000 words; combined total at most 80,000 words. | |
| input | Yes | Research task or question. | |
| model | No | Research agent model tier. The server applies `auto` when this is absent. | |
| outputLength | No | Target response size. The server applies `standard` when this is absent. | |
| outputSchema | No | JSON Schema defining structured output shape. | |
| citationFormat | No | Citation format in the report. The server applies `numbered` when this is absent. | |
| excludeDomains | No | Hard blocklist (max 20). Downward subdomain matching only. | |
| includeDomains | No | Soft source preference (max 20). Host-based subdomain matching. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so the safety profile is partly covered. The description adds important behavioral context beyond annotations: the task is asynchronous and results must be polled via research_get. It does not cover potential costs, rate limits, or task persistence, but the async/polling disclosure is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and immediately followed by the essential polling instruction. No filler or repetition; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 8-parameter tool with no output schema, the description covers the key behavioral facts: it creates an async task that produces a cited report, and results are retrieved with research_get. It omits details about return shape or parameter nuances, but those are either in the schema or in the sibling polling tool, making the description largely complete for its purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 8 parameters are fully documented in the schema. The description adds no parameter-specific meaning (e.g., it does not explain input, files, or model tiers). Baseline 3 is appropriate when the schema already carries the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and resource ('async research task'), and explains the action chain: searches, analyzes sources, and generates a cited report. It differentiates from the sibling tavily_research_get by pointing to polling, but does not explicitly distinguish itself from tavily_search, leaving a small gap in sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through 'Create an async research task...' and gives the follow-up action 'Poll results with research_get.' However, it does not state when to choose this over alternatives like tavily_search, nor any preconditions or exclusions. Usage is only partially implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tavily_research_getARead-onlyIdempotent
Retrieve the status and results of a research task by request_id. HTTP 202 means still running; poll until HTTP 200.
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | Yes | Research task UUID returned by `research_create`. | |
| includeUsage | No | Include credit usage in the response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, open-world, and idempotent characteristics. The description adds useful behavioral detail beyond those annotations by explaining the HTTP 202 in-progress status and the need to poll until HTTP 200.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with no wasted words. The purpose is front-loaded, followed immediately by the key polling behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter polling tool with full schema coverage and no output schema, the description covers the essential purpose and polling semantics. It does not describe the shape of returned results, but with no output schema that omission is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both requestId and includeUsage are already documented in the input schema. The description mentions request_id but adds no syntax, format, or usage detail beyond what the schema provides, making this baseline-level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Retrieve') and resource ('status and results of a research task') with the lookup key ('request_id'). It is clear enough to distinguish from generic search tools, though it does not explicitly name the sibling tavily_research_create as the origin of the request_id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational guidance for polling: HTTP 202 means still running, and the agent should poll until HTTP 200. It does not explicitly say when not to use this tool or name alternatives, but the intended context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tavily_searchA
Execute a real-time web search optimized for AI agents. Use when sources are unknown or current web context is needed. Prefer search_depth advanced with chunks_per_source 3 for stronger evidence per source.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query to execute. | |
| topic | No | Search category. The server applies `general` when this is absent. `news` automatically enables `include_published_date`. | |
| country | No | Boost results from a country using Tavily's lowercase English country name (for example `united states`). Available only when `topic` is `general`. | |
| endDate | No | Return results before this date (`YYYY-MM-DD`). | |
| language | No | Boost or filter results by language — an ISO 639-1 code (for example `en`, `fr`, `zh-cn`) or English language name (for example `english`, `french`). | |
| startDate | No | Return results after this date (`YYYY-MM-DD`). | |
| timeRange | No | Filter by publish or last-updated date window. | |
| exactMatch | No | Return only results containing the exact quoted phrase(s) in the query. | |
| maxResults | No | Maximum search results to return. The server applies 10 when this is absent. | |
| safeSearch | No | Filter adult or unsafe content. Not supported when `search_depth` is `fast` or `ultra-fast`. | |
| searchDepth | No | Latency/relevance tradeoff. The server applies `basic` when this is absent. `advanced` costs 2 credits; `basic`, `fast` and `ultra-fast` cost 1 credit. | |
| includeUsage | No | Include credit usage in the response. | |
| includeAnswer | No | Include an LLM-generated answer. `true` or `basic` returns a quick answer; `advanced` returns a detailed answer. The server applies `false` when this is absent. | |
| includeImages | No | Include query-related images and per-result `images`. | |
| autoParameters | No | Let Tavily configure parameters from the query. Explicit values override auto-selected ones. `include_answer`, `include_raw_content` and `max_results` must always be set manually when using this. | |
| excludeDomains | No | Domains to exclude (max 150). | |
| includeDomains | No | Domains to include (max 300). | |
| includeFavicon | No | Include a favicon URL per result. | |
| chunksPerSource | No | Maximum relevant chunks per source in each result's `content`. The server applies 3 when this is absent. Available only when `search_depth` is `advanced`, `basic` or `fast`. Each chunk is at most 500 characters and joined with `[...]`. | |
| filterByLanguage | No | Strictly filter out non-matching languages. Requires `language`. | |
| includeRawContent | No | Include cleaned page content per result. `true` or `markdown` returns markdown; `text` returns plain text and may increase latency. The server applies `false` when this is absent. | |
| includeDomainsMode | No | How `include_domains` is applied. Requires `include_domains` to be set. | |
| includePublishedDate | No | Include `published_date` on each result. Beta feature. Automatically enabled when `topic` is `news`. | |
| filterByPublishedDate | No | Remove results outside the date window or with no detectable date. Also enables `include_published_date`. | |
| includeImageDescriptions | No | Add descriptive text per image when `include_images` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, which is an unusual pairing for a read-only search and is left unexplained. The description adds a config recommendation (search_depth advanced, chunks_per_source 3) but doesn't clarify the credit costs or why a read-only operation isn't marked read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it does, when to use it, and an actionable configuration tip. Front-loaded and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large but fully documented 25-parameter schema with no output schema, the description covers the essential purpose, usage trigger, and a key tuning recommendation. It's adequate, though the odd readOnlyHint=false on a search tool could have been addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 25 parameters is already documented in the schema. The description names two parameters and their preferred values without adding semantics beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Execute a real-time web search') and scopes it ('optimized for AI agents'). It distinguishes itself from tavily_research_create by being a real-time search vs. a research task, though it never explicitly names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage trigger: 'Use when sources are unknown or current web context is needed.' This tells the agent when to reach for it over knowledge-only answers, though it doesn't name alternatives like firecrawl_scrape or tavily_research_create.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
35 tool updates
v0.1.0- First observed
connect - First observed
connection_status - First observed
firecrawl_extract - First observed
firecrawl_scrape - First observed
gcalendar_events_insert - First observed
gcalendar_events_list - First observed
gdocs_documents_get - First observed
gdrive_files_export - First observed
gdrive_files_list - First observed
gforms_forms_responses_list - First observed
gmail_drafts_create - First observed
gmail_messages_send - First observed
gmail_threads_list - First observed
granola_notes_get - First observed
granola_notes_list - First observed
gsheets_spreadsheets_values_get - First observed
gsheets_spreadsheets_values_update - First observed
linear_customer_create - First observed
notion_data_sources_query - First observed
notion_pages_create - First observed
notion_pages_update - First observed
notion_search - First observed
slack_chat_post_message - First observed
slack_conversations_create - First observed
stripe_checkout_sessions_create - First observed
stripe_customers_create - First observed
stripe_customers_retrieve - First observed
stripe_invoice_items_create - First observed
stripe_invoices_create - First observed
stripe_invoices_finalize - First observed
stripe_prices_list - First observed
stripe_subscriptions_list - First observed
tavily_research_create - First observed
tavily_research_get - First observed
tavily_search
TDQS
Scored across 35 tools
Most tools are clearly separated by app prefix and action, e.g. gmail_messages_send vs gmail_drafts_create. Some overlap remains among tavily_search, tavily_research_create, firecrawl_scrape, and firecrawl_extract, which could be confused for general web research tasks. Overall the set is mostly distinct but not perfectly unambiguous.
Tool names follow a highly predictable snake_case pattern, usually app/service/resource/action, such as notion_pages_create, stripe_invoices_finalize, and gcalendar_events_list. Minor deviations like connect and connection_status are meta-tools and do not break the pattern. The naming convention is consistent throughout.
With 35 tools, this is well above the typical 3-15 well-scoped range and exceeds the 25+ threshold for being too many. Although the server spans many integrations, the surface feels heavy and likely includes tools that could be consolidated or conditionally exposed. The count is not well-scoped for a single sales-prep agent.
Coverage is broad across apps but shallow and has significant gaps: Gmail can list threads but not read message bodies, Calendar can list/insert but not update/delete, and Linear only creates customers. granola_notes_get even references a notes_transcript_get tool that is not present, creating a dead end. These gaps would likely cause agent failures in real sales-prep workflows.
Maintenance
Related MCP Connectors
- mcp-serverOAuthcom.make
Give your AI agents the tools to build, manage, and run automation workflows.
CRM, inbox, meetings, revenue intelligence, analytics and automations for AI agents.
1Run B2B outreach from your AI agent: 250+ tools for campaigns, leads, LinkedIn and email workflows.
Agent-native CRM. 25 tools — contacts, deals, sequences, enrichment waterfall, audit log.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables agents to find leads, read buying signals, write and send email and LinkedIn outreach, handle replies, and report results, with built-in safety gates, approvals, limits, and compliance controls.1Apache 2.0
- AlicenseBqualityBmaintenanceRuns 20 slash-command workflows across Google Calendar, Gmail, Linear, Slack, Granola, Google Docs, GitHub, Stripe, Notion, Google Forms, Sheets and Drive to produce morning briefs, meeting prep, action items assigned to owners and weekly updates. Reads proceed without asking, while anything that creates, sends, changes or deletes is shown for approval first.47Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables freelancers and agencies to run eight back-office workflows across Stripe, Google Drive, Linear, Google Calendar, Gmail, GitHub, Google Docs, Granola, Google Sheets and Firecrawl, covering client onboarding, invoices from calendar and commits, status reports, scope-creep detection, site audits and overdue invoice chasers. Reads run freely, while anything that creates, sends, changes or deletes is shown for approval first, with credentials kept in your own OS keychain and no proxying through any third-party server.Apache 2.0
- AlicenseBqualityBmaintenanceRuns 13 agent slash-command workflows across Linear, GitHub, Google Docs, Slack, Firecrawl, Sheets, Tavily, Forms, Calendar, Gmail, Stripe, Granola and Notion — turning shipped features into blog and social drafts and handling competitor pricing, mention monitoring, SEO gaps, webinars, newsletters and launch-day tracking. Every credential is your own, kept in the OS keychain or client config, and anything that writes, sends or deletes is shown for approval first.29Apache 2.0