Marketing Ops
Reads repository data such as issues, pull requests, commits and releases to see what shipped, and compares public docs pages against the code they describe. Feeds blog/social draft generation, weekly newsletter content and docs-freshness flagging.
Sends and drafts email such as webinar invites and reminders, case-study permission asks, event follow-ups and the weekly newsletter.
Browser sign-in over your own OAuth client gives the server access to your Google Workspace apps (Docs, Sheets, Forms, Calendar and Gmail) without a per-app key.
Creates and manages calendar events for webinar scheduling and reminders.
Creates and updates Google Docs with generated content: blog drafts, social copy, case studies, insight memos and newsletter drafts.
Reads form responses, e.g. webinar registrations, survey answers and event lead scans, as the input to downstream workflows.
Writes and maintains spreadsheets such as competitor pricing history, prioritised SEO content gaps and tagged survey response tables.
Reads issues and shipped project data, and creates issues from SEO content gaps, site QA crawl findings and docs freshness problems.
Reads and writes Notion pages and databases for newsletter source material, enriched event leads, brand mention tracking and publishing blog posts and social snippet queues.
Posts alerts and updates to channels, e.g. competitor price changes, brand mentions, site QA issues, launch-day signup/bug pulses and insight memo summaries.
Reads customer and subscription data to identify longest-paying happy customers for case studies and to post hourly signup numbers during a launch.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Marketing Opsturn this week's shipped features into blog and social drafts"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Marketing Ops
Shipped work becomes blog drafts; competitor pricing, mentions and webinars run on schedule.
An MCP server with 13 workflows across Linear, GitHub, Google Docs, Slack, Firecrawl, Google Sheets, Tavily, Google Forms, Google Calendar, Gmail, Stripe, Granola and Notion. Each workflow is a prompt your agent runs as a slash command, over the 27 tools it needs and no others.
uv tool install https://github.com/r28ai/marketing-ops-mcp/releases/download/v0.1.0/marketing_ops_mcp-0.1.0-py3-none-any.whl
claude mcp add marketing -- marketing-ops-mcpIt installs with uv from this repository's release, with no git and nothing to build; nothing but Charter and the libraries it uses comes from PyPI. To update, run the install line from the latest release. If a desktop app cannot find marketing-ops-mcp, give it the full path from which marketing-ops-mcp (where marketing-ops-mcp on Windows).
Then ask your agent to connect your apps, or run /mcp__marketing__setup.
Connect your apps
Ask the agent to connect one ("connect Linear"). It tells you where to get that app's key and the command that stores it, and the next call works, with no restart. The agent never asks for a key in the chat.
Or connect everything this server uses from a terminal:
marketing-ops-mcp login # each app in turn
marketing-ops-mcp login linear # just one
marketing-ops-mcp status # what is connectedTokens and keys go to your operating system's keychain (macOS Keychain, Windows Credential Manager, the Secret Service on Linux), and are checked with one read-only call to the app's own API before they are kept. Every key, token and OAuth client is yours: we register no app with any of these services, and nothing passes through a server of ours, because there isn't one.
App | How it connects | Or set |
Linear | Your own key (get one), entered once. |
|
GitHub | Your own key (get one), entered once. |
|
Browser sign-in, over your own OAuth client (make one). |
| |
Slack | Your own key (get one), entered once. A bot token from your own Slack app, which the guide sets up in about three minutes. |
|
Firecrawl | Your own key (get one), entered once. |
|
Tavily | Your own key (get one), entered once. |
|
Stripe | Your own key (get one), entered once. |
|
Granola | Your own key (get one), entered once. In the Granola app: Settings → Connectors → API keys. Business plan or above. |
|
Notion | Your own key (get one), entered once. Then share the pages it should see with the integration. |
|
A variable set in your client's config always wins over the keychain.
Related MCP server: Freelancer Ops MCP
Workflows
Workflow | What you get | Apps |
What shipped → blog and social drafts | Marketing hears about features from the tracker, not from customers. | Linear, GitHub, Google Docs, Slack |
Competitor pricing tracker | Price and packaging changes, dated, in a sheet, with an alert when one moves. | Firecrawl, Google Sheets, Slack |
SEO content gap | Topics competitors rank for that you have no page on, as a prioritised backlog. | Firecrawl, Tavily, Google Sheets, Linear |
Webinar ops end to end | Registration, invite and reminders with no webinar tool. | Google Forms, Google Calendar, Gmail |
Customer story pipeline | Your longest-paying happy customers become case-study drafts and a permission ask. | Stripe, Granola, Google Docs, Gmail |
Weekly newsletter draft | Product news, posts and releases gathered into one draft every Friday. | Linear, GitHub, Notion, Google Docs, Gmail |
Event leads → enriched follow-ups | Booth scans become enriched leads with a follow-up waiting, same day. | Google Forms, Firecrawl, Notion, Gmail |
Brand mention monitor | Every mention worth replying to shows up with the context to reply. | Tavily, Firecrawl, Notion, Slack |
Webinar transcript → blog post | An hour of talk becomes a post and a queue of social snippets. | Granola, Google Docs, Notion |
Site QA crawl | Broken links, empty pages and stale pricing get issues before a prospect finds them. | Firecrawl, Linear, Slack |
Launch-day pulse | Hourly signups and fresh bugs posted to the launch channel without anyone refreshing dashboards. | Stripe, Linear, Slack |
Survey → insight memo | Raw responses become a tagged sheet and a one-page memo. | Google Forms, Google Sheets, Google Docs, Slack |
Docs site freshness check | Public docs pages older than the code they describe get flagged. | Firecrawl, GitHub, Linear |
Every prompt takes one optional argument, details: the repo, team, channel, customer or date range you mean, so the agent does not have to ask. In Claude Code, put it in quotes, or only its first word arrives:
/mcp__marketing__what_shipped_to_blog_and_social_drafts "everything shipped since the 1st"Reads run without asking. Before anything that creates, sends, changes or deletes, the prompt tells the agent to show you the call and wait.
4 of the 13 workflows need no Google or Granola credential.
Other clients
Claude Desktop: install uv if you have not, since Claude Desktop starts the server with it, then open the .mcpb from the latest release. Claude asks for any keys in its own settings and keeps them in your keychain. The first start takes a few seconds longer, while uv installs it.
VS Code (.vscode/mcp.json): VS Code asks for each key the first time the server starts and stores it securely. Leave out any you stored with login.
{
"inputs": [
{
"type": "promptString",
"id": "linear-api-key",
"description": "Linear: Personal API key",
"password": true
},
{
"type": "promptString",
"id": "github-token",
"description": "GitHub: Personal access token",
"password": true
},
{
"type": "promptString",
"id": "google-client-secret",
"description": "Google: OAuth client secret",
"password": true
},
{
"type": "promptString",
"id": "slack-bot-token",
"description": "Slack: Bot token (xoxb-\u2026)",
"password": true
},
{
"type": "promptString",
"id": "firecrawl-api-key",
"description": "Firecrawl: API key",
"password": true
},
{
"type": "promptString",
"id": "tavily-api-key",
"description": "Tavily: API key",
"password": true
},
{
"type": "promptString",
"id": "stripe-api-key",
"description": "Stripe: Secret or restricted key",
"password": true
},
{
"type": "promptString",
"id": "granola-api-key",
"description": "Granola: API key",
"password": true
},
{
"type": "promptString",
"id": "notion-api-key",
"description": "Notion: Integration secret (ntn_\u2026)",
"password": true
}
],
"servers": {
"marketing": {
"type": "stdio",
"command": "marketing-ops-mcp",
"env": {
"LINEAR_API_KEY": "${input:linear-api-key}",
"GITHUB_TOKEN": "${input:github-token}",
"GOOGLE_CLIENT_SECRET": "${input:google-client-secret}",
"SLACK_BOT_TOKEN": "${input:slack-bot-token}",
"FIRECRAWL_API_KEY": "${input:firecrawl-api-key}",
"TAVILY_API_KEY": "${input:tavily-api-key}",
"STRIPE_API_KEY": "${input:stripe-api-key}",
"GRANOLA_API_KEY": "${input:granola-api-key}",
"NOTION_API_KEY": "${input:notion-api-key}",
"GOOGLE_CLIENT_ID": ""
}
}
}
}Cursor (.cursor/mcp.json) starts it the same way:
{
"mcpServers": {
"marketing": {
"command": "marketing-ops-mcp"
}
}
}Codex (~/.codex/config.toml) starts a turn without waiting for a server unless it is required, and then the agent has none of its tools. required = true makes the session wait for it, and startup_readiness = "catalog" waits for its tool list rather than just its connection:
[mcp_servers.marketing]
command = "marketing-ops-mcp"
required = true
startup_readiness = "catalog"
startup_timeout_sec = 30Name the server marketing. A host builds each tool's name from that key, and a longer one can push a tool past the 64 characters a function name allows.
Built with Charter
Every tool here is a Charter declaration: a Pydantic schema saying where each field goes on the wire. Charter's runtime builds the request, attaches and refreshes the credential, and trims the response before the model reads it. It runs in your process, with no proxy and no telemetry.
The 27 tool schemas come to 56,044 tokens.
The same tools work in your own agent, without MCP:
from charter.adapters.openai import to_openai_tools
from charter_packs_mcp import FAMILIES
tools = FAMILIES["marketing"].tools()
definitions = to_openai_tools(tools) # or charter.adapters.langchainNeed an API that isn't here? Write a pack: your coding agent writes the declarations, and Charter's conformance suite checks them.
Linear:
linear_issues_list,linear_issue_createGitHub:
github_releases_list,github_repos_list_commitsGoogle Docs:
gdocs_documents_createSlack:
slack_chat_post_message,slack_chat_schedule_messageFirecrawl:
firecrawl_monitor_create,firecrawl_extract,firecrawl_map,firecrawl_scrape,firecrawl_crawl,firecrawl_crawl_statusGoogle Sheets:
gsheets_spreadsheets_values_append,gsheets_spreadsheets_values_updateTavily:
tavily_searchGoogle Forms:
gforms_forms_create,gforms_forms_responses_listGoogle Calendar:
gcalendar_events_insertGmail:
gmail_messages_send,gmail_drafts_createStripe:
stripe_customers_list,stripe_checkout_sessions_listGranola:
granola_notes_list,granola_notes_transcript_getNotion:
notion_data_sources_query,notion_pages_create
License
Apache 2.0.
Available Tools
29 toolsconnectA
Connect one app this server uses. For an app that issues keys, says where to get one and the terminal command that stores it. For Google, once the user's own OAuth client is set, starts the browser sign-in and returns at once: the user approves in the browser and the next call works. To see which apps are connected, call connection_status. Never ask the user for a key in the chat.
| Name | Required | Description | Default |
|---|---|---|---|
| app | Yes | The app to connect. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false and openWorldHint=false; the description goes well beyond them by describing the key-provisioning flow, the fact that Google sign-in returns immediately while approval happens in the browser, and that the next call will succeed. This is exactly the behavioral context an agent needs for an auth/setup mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and packs three useful, non-redundant pieces of information into a short block. The Google sentence is somewhat dense and reads as a run-on, costing a point, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating return behavior ('returns at once: the user approves in the browser and the next call works'). Combined with the alternative tool pointer and the anti-pattern warning, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single `app` parameter has a full enum, so the schema already carries the parameter meaning. The description adds no format or syntax detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (connect) and resource (one app this server uses), and the enum in the schema pins the exact set of apps. It is trivially distinguishable from the sibling app-operation tools and from the read-side `connection_status`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'To see which apps are connected, call `connection_status`', and gives a hard rule ('Never ask the user for a key in the chat'). Per-app guidance (key-issuing apps vs Google) tells the agent when different sub-flows apply.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connection_statusARead-only
See which apps this server is connected to, and how to connect each one that is not. Changes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered by structured data. 'Changes nothing' restates the readOnly hint rather than adding new behavior; the only incremental value is noting that connect instructions are returned for unconnected apps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the primary purpose and appends the secondary benefit with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-param, read-only status tool with no output schema, the description covers both what is inspected and the shape of the useful payload (connect guidance). Return format details are absent but minimal given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly implies no input is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('See') and resource ('which apps this server is connected to'), and adds the secondary payload of connect instructions for missing apps. This distinguishes it from the sibling 'connect' tool, which performs the connection rather than reporting status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'how to connect each one that is not' implies this tool is the discovery step before using 'connect', but the sibling is never named and there is no explicit when-to-use/when-not statement. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_crawlC
Recursively crawl a website and scrape each discovered page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The base URL to start crawling from | |
| delay | No | Delay in seconds between scrapes. Setting this forces concurrency to 1. | |
| limit | No | Maximum number of pages to crawl. The server applies 10000 when this is absent. | |
| prompt | No | Natural language prompt to generate crawler options from. | |
| sitemap | No | Sitemap mode when crawling. The server applies 'include' when this is absent. | |
| webhook | No | Webhook specification for crawl lifecycle events. | |
| excludePaths | No | URL pathname regex patterns that exclude matching URLs from the crawl. | |
| includePaths | No | URL pathname regex patterns that include matching URLs in the crawl. | |
| scrapeOptions | No | Options applied when scraping each crawled page. | |
| maxConcurrency | No | Maximum number of concurrent scrapes for this crawl. | |
| regexOnFullURL | No | Match includePaths and excludePaths against the full URL instead of just the pathname. | |
| allowSubdomains | No | Allow the crawler to follow links to subdomains of the main domain. | |
| ignoreRobotsTxt | No | Ignore the website's robots.txt rules. Enterprise only. | |
| robotsUserAgent | No | Custom User-Agent string for robots.txt evaluation. Enterprise only. | |
| crawlEntireDomain | No | Allow the crawler to follow internal links to sibling or parent URLs, not just child paths. | |
| maxDiscoveryDepth | No | Maximum depth to crawl based on discovery order. | |
| zeroDataRetention | No | If true, this will enable zero data retention for this crawl. | |
| allowExternalLinks | No | Allow the crawler to follow links to external websites (one hop only). | |
| ignoreQueryParameters | No | Do not re-scrape the same path with different query parameters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the description need only add context, and it adds none. It omits that this kicks off an asynchronous crawl job (a firecrawl_crawl_status sibling exists for polling), that pages/credits are consumed, and that a limit defaults to 10000.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. It is efficient, though its brevity is partly under-specification rather than true economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 19-parameter, open-world, non-read-only crawl tool with no output schema, the description is far too thin. It says nothing about asynchronous job behavior, polling via crawl_status, scope controls (allowSubdomains, includePaths), or cost, all of which an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the 19 parameters are fully documented in the schema. The description contributes no additional parameter meaning, which is the baseline-3 case when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'recursively crawl a website and scrape each discovered page.' The word 'recursively' and 'each discovered page' implicitly distinguish it from the single-page firecrawl_scrape sibling, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use guidance, no conditions selecting it over firecrawl_scrape, firecrawl_map, or firecrawl_extract, and no prerequisites. The agent must infer usage purely from the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_crawl_statusBRead-onlyIdempotent
Get the status and results of a crawl job.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The ID of the crawl job |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, covering the safety profile. The description adds that it returns both status and results (not just status), which is useful context, but it omits polling guidance, rate limits, and result format details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero wasted words. It is appropriately sized for the simplicity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only status tool with rich annotations and no output schema, the description is minimally adequate. It does not explain the async/polling nature of crawl jobs or what 'results' entail, leaving clear gaps for an agent that may need to know whether to call repeatedly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single parameter 'id' is fully documented in the schema. The description adds no parameter-specific meaning beyond what the schema provides, which is the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (status and results of a crawl job), which clearly distinguishes it from the sibling firecrawl_crawl that likely starts a crawl. However, it does not explicitly name the alternative or state the relationship, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description only says what the tool does; it gives no guidance on when to use it versus alternatives, no prerequisites, and no mention of polling behavior. Implied usage is weak because no context is provided about the crawl lifecycle.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_extractA
Extract structured data from one or more URLs using an LLM. Poll results with extract_status.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | The URLs to extract data from. URLs should be in glob format. | |
| prompt | No | Prompt to guide the extraction process. | |
| schema | No | Schema to define the structure of the extracted data. Must conform to JSON Schema. | |
| showSources | No | When true, the sources used to extract the data will be included in the response as `sources`. | |
| ignoreSitemap | No | When true, sitemap.xml files will be ignored during website scanning. | |
| scrapeOptions | No | Options applied when scraping pages for extraction. | |
| enableWebSearch | No | When true, the extraction will use web search to find additional data. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| ignoreInvalidURLs | No | If invalid URLs are specified, they are ignored and returned in invalidURLs instead of failing the request. The server applies true when this is absent. | |
| includeSubdomains | No | When true, subdomains of the provided URLs will also be scanned. The server applies true when this is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, openWorldHint=true) are minimal and do not reveal the async job lifecycle, so the description correctly surfaces the most important behavioral trait: results must be polled via extract_status. It stops there, omitting what the initial call returns (a job id?), failure behavior, and whether web-search/credit costing applies. Decent added context over annotations, but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The core purpose is front-loaded and the polling instruction follows immediately. Nothing to trim and nothing misplaced.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, no output schema, and an async job model, the description covers the poll step but not the full lifecycle: it does not explain what the initial invocation returns or how to correlate extract_status results back to it. For a complex async tool with no output schema to fall back on, this is a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 10 parameters (urls, prompt, schema, scrapeOptions, enableWebSearch, etc.) are already documented in the schema. The description adds no parameter-level detail (e.g., the glob URL format or JSON Schema requirement) beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Extract structured data from one or more URLs') and adds the key mechanism ('using an LLM'), so the agent understands this is LLM-driven structured extraction rather than raw scraping. It does not, however, distinguish itself from the very close sibling firecrawl_scrape, leaving the boundary between the two to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence ('Poll results with extract_status') is genuine usage guidance: it tells the agent this is an async operation that must be polled. But it gives no guidance on when to choose this tool over firecrawl_scrape or other extraction options, and does not name any exclusion criteria. Usage is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_mapB
Discover and list all URLs of a website starting from a base URL. Use when the user wants to explore a site's structure or find specific pages.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The base URL to start crawling from. | |
| limit | No | Maximum number of links to return. The server applies 5000 when this is absent. | |
| search | No | Specify a search query to order the results by relevance. Example: 'blog' will return URLs that contain the word 'blog' in the URL ordered by relevance. | |
| sitemap | No | Sitemap mode when mapping. If you set it to `skip`, the sitemap won't be used to find URLs. If you set it to `only`, only URLs that are in the sitemap will be returned. The server applies `include` when this is absent. | |
| timeout | No | Timeout in milliseconds. There is no timeout by default. Milliseconds, not seconds: 30000 is thirty seconds, and 30 fails instantly. | |
| location | No | Location settings for the request. | |
| ignoreCache | No | Bypass the sitemap cache to retrieve fresh URLs. Sitemap data is cached for up to 7 days; use this parameter when your sitemap has been recently updated. | |
| auditMetadata | No | User attribution included with SIEM logging events when SIEM is enabled. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| includeSubdomains | No | Include subdomains of the website. The server applies true when this is absent. | |
| ignoreQueryParameters | No | Do not return URLs with query parameters. The server applies true when this is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint=true, and the schema itself discloses the practically important behaviors (7-day sitemap cache, ignoreCache to bypass it, timeout in milliseconds, default limit of 5000). The description adds no behavioral context beyond 'list all URLs' and is silent on cost, latency, rate limits, or whether the result is bare URLs or URL+metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the core capability is front-loaded before the usage hint. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 parameters, no output schema, and no annotations describing the return value, the definition is serviceable but thin at the top level. The schema compensates heavily for parameter behavior, yet nothing tells an agent what a successful map actually returns (URL list shape, ordering, whether it is paginated).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all 11 parameters carry their own prose, including defaults the server applies when omitted. Baseline 3 applies; the description contributes no additional semantic guidance about any parameter such as search, sitemap mode, or limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Discover and list all URLs of a website starting from a base URL.' That is unambiguous on its own. However, it never distinguishes itself from firecrawl_crawl or firecrawl_scrape, which are adjacent in the sibling list and do partially overlapping work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use when the user wants to explore a site's structure or find specific pages' gives a usage context, so it is more than implied. But it names no alternative and offers no exclusions, which is a real gap given firecrawl_crawl (get URLs plus page content) and firecrawl_scrape (single page) sit right beside it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_monitor_createB
Create a scheduled monitor for scrape, crawl, or search targets.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | Plain-language goal used to judge whether changed pages are meaningful. | |
| name | Yes | Monitor name. | |
| targets | Yes | Targets to run on each check. | |
| webhook | No | Webhook destination for monitor events. | |
| schedule | Yes | Schedule for monitor checks. | |
| judgeEnabled | No | Whether to judge changed pages against goal. | |
| notification | No | Notification destinations. | |
| retentionDays | No | How long to retain monitor history. The server applies 30 when this is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so the mutation and external-reach profile is already covered; the description's 'create' is consistent with these and adds no contradiction. It adds only the notion that the created object is scheduled/recurring, but says nothing about lifecycle, cost, or persistence beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, correctly leading with the verb and the resource. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex creation tool with required schedule/targets, nested target union types, webhooks, notifications, retention, and an optional judge/goal mechanism, and no output schema to fall back on. The one-line description does not explain what a monitor does over time or what happens after creation, leaving significant gaps for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 8 parameters (schedule, targets, webhook, judgeEnabled, notification, retentionDays, goal, name) are already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a scheduled monitor') and scopes it to the three target kinds the schema supports (scrape, crawl, search). It implicitly separates this from the one-off firecrawl_scrape/firecrawl_crawl siblings via 'scheduled,' but never names them, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites (e.g. that monitors run repeatedly and incur ongoing credit usage), and no pointer to alternatives like firecrawl_crawl or firecrawl_crawl_status for one-off jobs. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_scrapeB
Scrape a single URL and optionally extract information. Use when the user wants to read or summarize a specific webpage. Supports markdown, HTML, screenshots, and structured JSON extraction.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| proxy | No | Specifies the type of proxy to use. | |
| maxAge | No | Returns a cached version of the page if it is younger than this age in milliseconds. The server applies 172800000 (2 days) when this is absent. | |
| minAge | No | When set, the request only checks the cache and never triggers a fresh scrape. | |
| mobile | No | Emulate scraping from a mobile device. | |
| actions | No | Actions to perform on the page before grabbing the content. | |
| formats | No | Output formats to include in the response. Strings or objects. The server applies markdown when this is absent. | |
| headers | No | Headers to send with the request. | |
| parsers | No | Controls how files are processed during scraping. | |
| profile | No | Persistent browser storage across scrape and interact sessions. | |
| timeout | No | Timeout in milliseconds. The server applies 60000 when this is absent. | |
| waitFor | No | Specify a delay in milliseconds before fetching the content. The server applies 0 when this is absent. | |
| blockAds | No | Enables ad-blocking and cookie popup blocking. | |
| location | No | Location settings for the request. | |
| lockdown | No | Serve from cache only and never make an outbound request. On miss, returns 404 SCRAPE_LOCKDOWN_CACHE_MISS. | |
| redactPII | No | Redact personally identifiable information from returned markdown. Pass true for defaults, or an object to tune it. | |
| excludeTags | No | Tags to exclude from the output. | |
| includeTags | No | Tags to include in the output. | |
| storeInCache | No | If true, the page will be stored in the Firecrawl index and cache. | |
| auditMetadata | No | User attribution included with SIEM logging events when SIEM is enabled. | |
| onlyMainContent | No | Only return the main content of the page excluding headers, navs, footers, etc. The server applies true when this is absent. | |
| onlyCleanContent | No | Beta. LLM pass over markdown to remove residual boilerplate that onlyMainContent can miss. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| zeroDataRetention | No | If true, this will enable zero data retention for this scrape. To enable this feature, please contact help@firecrawl.dev | |
| removeBase64Images | No | Removes all base64 images from the markdown output. | |
| skipTlsVerification | No | Skip TLS certificate verification when making requests. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the agent knows this touches external state. The description adds the supported output formats, which is useful, but it omits behavior implied by the schema — that actions (click/write/executeJavascript) mutate the page, that storeInCache writes to an external index, and that some features cost credits. Nothing contradicts the annotations, but the added behavioral detail is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core purpose front-loaded and no filler. The trailing format list is somewhat redundant with the schema's formats enum, which keeps it short of ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a 26-parameter tool with no output schema and only minimal annotations, yet the description is three sentences long. It omits cost/credit implications, caching/lockdown semantics, the relationship to firecrawl_extract, and any hint of what the response looks like, so an agent invoking it correctly still depends almost entirely on reading the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 26 parameters are already documented in-schema and the baseline is 3. The description echoes the formats dimension ('markdown, HTML, screenshots, structured JSON extraction') but adds no format syntax, precedence, or interaction detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Scrape a single URL') plus an optional outgrowth ('optionally extract information'), so an agent can tell it is a per-URL content fetcher. The phrase 'a single URL' gestures at the multi-URL alternative but never names firecrawl_extract, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one clear trigger ('Use when the user wants to read or summarize a specific webpage'), which is real usage guidance. But it never states when NOT to use it, nor does it point to firecrawl_extract for bulk/structured extraction, so the routing decision against the closest sibling is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gcalendar_events_insertC
Create a calendar event; returns details of the event.
| Name | Required | Description | Default |
|---|---|---|---|
| event | Yes | The event to insert. | |
| calendarId | Yes | Calendar identifier. To retrieve calendar IDs call the calendarList.list method. If you want to access the primary calendar of the currently logged in user, use the "primary" keyword. | |
| sendUpdates | No | Guests who should receive notifications about the change. Acceptable values are: "all" (notifications are sent to all guests), "externalOnly" (notifications are sent to non-Google Calendar guests only), "none" (no notifications are sent; for calendar migration tasks, consider using the Events.import method instead). | |
| maxAttendees | No | The maximum number of attendees to include in the response. If there are more than the specified number of attendees, only the participant is returned. Optional. | |
| eventLabelVersion | No | Version number of the event label feature supported by the API client. Version 0 assumes no event label support and processes the colorId field for color management. Version 1 enables support for event labels, and processes the eventLabelId in the event's body. In this case, the colorId field is ignored. The default is 0. Acceptable values are 0 to 1, inclusive. | |
| supportsAttachments | No | Whether API client performing operation supports event attachments. Optional. The default is False. | |
| conferenceDataVersion | No | Version number of conference data supported by the API client. Version 0 assumes no conference data support and ignores conference data in the event's body. Version 1 enables support for copying of ConferenceData as well as for creating new conferences using the createRequest field of conferenceData. The default is 0. Acceptable values are 0 to 1, inclusive. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false (a mutation) and openWorldHint=true. The description adds only 'returns details of the event' and omits meaningful behavioral context for a mutation tool of this complexity: guest notifications, the sendUpdates/conferenceDataVersion/supportsAttachments side-effect flags, permission requirements, and attendee-invitation behavior are all undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. It is efficient, though the trailing return-value clause could arguably be dropped since it is the only content beyond the verb.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the brief note about returned details is helpful, but for a mutation tool with 7 parameters and rich side-effect potential (invitations, notifications, conference generation) the description is thin. It is minimally adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema fully documents all 7 parameters including sendUpdates, conferenceDataVersion, and maxAttendees. Per the rubric, that establishes a baseline of 3; the description adds nothing beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a calendar event'), which clearly distinguishes it from the sibling gcalendar_events_list. It does not, however, differentiate itself from similar write operations or mention scope (which calendar), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use context, prerequisites, or alternatives. An agent must infer that this is the creation counterpart to gcalendar_events_list without any explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gdocs_documents_createA
Create a blank document with a title. Only the title is honoured — the document is created empty. To add content, call this and then documents_batch_update with the returned documentId.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | The document to create. Only the title is honoured. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries the behavioral load and does well: it discloses the critical constraint that only the title is honoured and the document is created empty. It also notes the documentId is returned, which matters with no output schema. It stops short of noting auth/permission or quota behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, zero padding, and the most important constraint (empty document, title-only) is front-loaded before the follow-up instructions. Every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by mentioning the returned documentId and the batch_update follow-up, which is what an agent needs to chain calls. For a single-param mutation tool with annotations covering the safety profile, this is nearly complete; only auth/error behavior is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the title field already states it is the only honoured field, so the description largely restates structured data. Baseline 3 applies; the emphasis on the empty-document effect is useful but not new information beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (document), and immediately disambiguates the scope: it produces a blank/empty document, not a content-bearing one. This is a distinct action an agent can tell apart from content-writing operations. The mention of documents_batch_update further fixes its place in the workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear sequencing guidance: to add content, create first and then call documents_batch_update with the returned documentId. It does not state explicit exclusions or alternatives (e.g., when to use a copy/template flow instead), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gforms_forms_createA
Create a new form from a title. Only the title and the document title are honoured: the form is created with no description, no items and default settings. To add questions, call this and then forms_batch_update with the returned formId. Pass unpublished=true to create a form that does not yet accept responses.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The form to create. | |
| unpublished | No | Optional. Whether the form is unpublished. If set to `true`, the form doesn't accept responses. If set to `false` or unset, the form is published and accepts responses. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give readOnlyHint=false and openWorldHint=true, but the description adds real behavioral context: only title and documentTitle survive, the form starts with no description/items and default settings, and unpublished controls whether responses are accepted. That is exactly the kind of side-effect disclosure a mutation tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope, then the chaining instruction, then the flag. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers creation scope, the chaining workflow, and the unpublished flag; mentioning the returned formId partially compensates for the absent output schema. Minor gaps remain (e.g., auth or quota requirements), but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents body.info and unpublished, including the unpublished response-acceptance semantics. The description's 'only title and documentTitle are honoured' mirrors the schema's own wording on body.info, adding little beyond it, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new form from a title') and immediately scopes what the created object contains. It is clearly distinguishable from gdocs_documents_create and other create-style siblings by naming the Google Forms resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent the follow-up path for adding questions: 'call this and then forms_batch_update with the returned formId', naming the sibling tool and the condition. It also explains the unpublished=true branch, so the agent knows when to use that flag.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gforms_forms_responses_listARead-onlyIdempotent
List a form's submitted responses, newest page first, up to 5000 per page. The only supported filter is on submission time: pass filter='timestamp >= 2026-01-01T00:00:00Z' to read what has arrived since a point in time. Answers come back keyed by questionId.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Which form responses to return. Currently, the only supported filters are: `timestamp > N` which means to get all form responses submitted after (but not at) timestamp N, and `timestamp >= N` which means to get all form responses submitted at and after timestamp N. For both supported filters, timestamp must be formatted in RFC3339 UTC "Zulu" format. Examples: "2014-10-02T15:01:23Z" and "2014-10-02T15:01:23.045123456Z". The whole filter is one string, operator included: 'timestamp >= 2014-10-02T15:01:23Z'. There is no other filterable field — a question, an email or a score cannot be filtered here. | |
| formId | Yes | Required. ID of the Form whose responses to list. | |
| pageSize | No | The maximum number of responses to return. The service may return fewer than this value. If unspecified or zero, at most 5000 responses are returned. | |
| pageToken | No | A page token returned by a previous list response. If this field is set, the form and the values of the filter must be the same as for the original request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds behavior the annotations do not: newest-first ordering, the 5000-per-page ceiling, and that answers are keyed by questionId. It omits pagination semantics (that pageToken must repeat the same filter), which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the resource and ordering, then the filter rule, then the return shape. Every sentence carries information an agent needs and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description partially compensates by noting answers are keyed by questionId, plus it gives ordering and page limits. It stops short of describing the response envelope (e.g. nextPageToken, responseId) or how to continue paging, which is the remaining gap for a paginated list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including the full filter grammar, RFC3339 format and the pageToken consistency rule. The description's filter example and questionId note mostly restate that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List a form's submitted responses') plus ordering ('newest page first') and a hard cap ('up to 5000 per page'). No other Google Forms tool exists among the siblings, so there is nothing to disambiguate against, and the purpose is unambiguous on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operative constraint clearly: the only supported filter is on submission time, with a worked example that shows what the filter buys you ('read what has arrived since a point in time'). It does not state when *not* to use it or how pagination resumes, but for a single-purpose list tool the context given is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_releases_listARead-onlyIdempotent
List a repository's releases, newest first. Tags that were never made into releases do not appear here — repos_list_tags has those. Drafts are visible only with push access.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | The page number of the results to fetch. Defaults to 1. | |
| repo | Yes | The name of the repository without the `.git` extension. The name is not case sensitive. | |
| owner | Yes | The account owner of the repository. The name is not case sensitive. | |
| perPage | No | The number of results per page (max 100). Defaults to 30. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds real value beyond that: newest-first ordering and the auth nuance that drafts require push access, which an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and ordering, followed by the disambiguation and auth caveat. No filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with full schema coverage and no output schema, the description covers ordering, sibling routing and the draft-access caveat. Pagination behavior is left entirely to the schema, which is a minor gap but not a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so page, perPage, owner and repo are fully documented by the schema itself. The description adds no parameter-level detail such as pagination limits or format hints, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (a repository's releases) plus the ordering guarantee (newest first). It also explicitly distinguishes itself from the sibling repos_list_tags, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative tool (repos_list_tags) and the exact condition that selects it: tags never made into releases do not appear here. It also states the draft-visibility precondition, giving clear when-to-use and when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
github_repos_list_commitsARead-onlyIdempotent
List commits, newest first. Narrow with sha for a branch, path for one file's history, or since/until for a window.
| Name | Required | Description | Default |
|---|---|---|---|
| sha | No | SHA or branch to start listing commits from. Defaults to the repository's default branch. | |
| page | No | The page number of the results to fetch. Defaults to 1. | |
| path | No | Only commits containing this file path will be returned. | |
| repo | Yes | The name of the repository without the `.git` extension. The name is not case sensitive. | |
| owner | Yes | The account owner of the repository. The name is not case sensitive. | |
| since | No | Only commits after this date will be returned, in ISO 8601 format: `YYYY-MM-DDTHH:MM:SSZ`. | |
| until | No | Only commits before this date will be returned, in ISO 8601 format: `YYYY-MM-DDTHH:MM:SSZ`. | |
| author | No | GitHub username or email address to filter commits by author. | |
| perPage | No | The number of results per page (max 100). Defaults to 30. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is fully covered structurally. The description adds the 'newest first' ordering trait, but says nothing about pagination behavior, result caps, or rate limits that the annotations don't cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the ordering behavior is front-loaded and the filtering options follow compactly. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter read tool with full schema coverage, no output schema, and annotations covering safety, the description supplies the key behavioral trait (ordering) and a map of the main filters. Pagination via page/perPage is only in the schema, a minor omission rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented in the schema (including defaults for page/perPage and ISO 8601 format for since/until). The description only restates sha/path/since/until usage without adding format or edge-case detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List commits') plus an ordering guarantee ('newest first'), which is more than the name alone conveys. There is no sibling commit-listing tool to differentiate against, so the lack of explicit contrast is not a gap here.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent how to narrow results: `sha` for a branch, `path` for one file's history, `since`/`until` for a time window. This is clear usage guidance, though it never states when-not to use the tool or names an alternative tool for other commit-listing needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_drafts_createC
Save an email draft to Gmail.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The draft to create. | |
| userId | No | The user's email address. The special value 'me' can be used to indicate the authenticated user. | me |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/non-read nature is covered structurally. The description adds nothing beyond that: it does not say the draft is not sent, whether creation is idempotent, or what auth is required, so the behavioral burden is largely unmet.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is well structured. It is arguably under-specified rather than bloated, but as a size/structure judgment it is tight and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with only minimal annotations and no output schema, the description should at least clarify that it saves rather than sends and hint at the returned draft. The rich input schema compensates for parameters, but the core behavioral distinction is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the nested Message/Draft/EmailContent fields are richly documented (threadId rules, bodyHtml multipart behavior, in_reply_to threading). The description adds no parameter meaning at all, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Save an email draft') and destination (Gmail), which is enough to know it creates a draft rather than sending. However, it does not distinguish itself from the adjacent gmail_messages_send sibling, so an agent gets no explicit routing cue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the alternative tool (gmail_messages_send) for actually delivering mail. The agent must infer that this only persists a draft.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_messages_sendC
Send an email via the Gmail API.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The email message data. | |
| userId | No | The user's email address. The special value 'me' can be used to indicate the authenticated user. | me |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the safety profile is partially covered. The description adds no behavioral traits beyond that, such as immediate sending, authentication requirements, irreversibility, or threading behavior; it essentially restates the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded sentence with no wasted words. While it is sparse, the structure is efficient and the core action is stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with open-world annotations and no output schema, the description is incomplete. It does not clarify that the message is sent immediately rather than saved as a draft, nor does it describe return behavior or error handling, leaving important context gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the nested EmailContent fields are fully documented. The description adds no parameter-level meaning beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Send) and resource (email via Gmail API), which is clearer than a generic action. However, it does not differentiate from sibling tools like gmail_drafts_create or slack_chat_post_message, so it misses the top mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It simply states the action without any contextual routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
granola_notes_listARead-onlyIdempotent
List meeting notes, filtered by when they were created or last updated and optionally narrowed to one folder and its subfolders. Returns each note's id, title, owner and timestamps, not its content. Fetch that with notes_get. Only notes that already have a generated AI summary appear here.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | The cursor to continue from | |
| folderId | No | Return notes in this folder and any of its child folders. Use the list folders endpoint to discover folder IDs. | |
| pageSize | No | Maximum number of notes to return per page. The server returns 10 when this is absent. | |
| createdAfter | No | Return notes created after this date. A date (`2026-01-27`) or a date-time (`2026-01-27T15:30:00Z`). | |
| updatedAfter | No | Return notes updated after this date. A date (`2026-01-27`) or a date-time (`2026-01-27T15:30:00Z`). | |
| createdBefore | No | Return notes created before this date. A date (`2026-01-27`) or a date-time (`2026-01-27T15:30:00Z`). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld), so the bar is lower, and the description adds real behavioral context: the response is metadata-only (id, title, owner, timestamps) and only AI-summarized notes are returned. It does not mention pagination or cursor behavior, which is a notable omission for a list tool returning partial pages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what is listed and filtered, then the return shape, then the sibling routing. Every sentence carries distinct information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields and the AI-summary precondition, which is the key thing an agent needs to know before calling. Pagination behavior is left entirely to the cursor parameter's schema description, a minor gap for a paged list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including folder recursion, pageSize default of 10, cursor, and date formats. The description only restates the existence of the time and folder filters, adding essentially nothing beyond the schema, which is the baseline-3 case for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list meeting notes) plus the two filtering axes (creation/update time, folder scope) and explicitly distinguishes the resource from its sibling by noting that content is fetched with notes_get. An agent can differentiate it from granola_notes_get without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to the right sibling for content ('Fetch that with notes_get') and discloses a decisive selection condition ('Only notes that already have a generated AI summary appear here'). It stops short of explicit when-not-to-use guidance, such as what to do if the summary requirement excludes a desired note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
granola_notes_transcript_getARead-onlyIdempotent
Read a meeting transcript one page at a time. Use this for any long transcript, and whenever notes_get answers 413 TRANSCRIPT_TOO_LARGE. Each item is one line of speech with who said it and when.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | The opaque cursor returned by the previous page | |
| noteId | Yes | The ID of the note, as returned by the list endpoint — a `not_` prefix followed by fourteen alphanumeric characters. The UUID in a Granola web app URL is a different identifier and is not accepted here. | |
| pageSize | No | The maximum number of transcript items to return. The server returns 50 when this is absent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so safety is covered. The description adds genuinely new behavioral context: the tool returns paginated results and each item is a line of speech with speaker and timestamp. It stops short of describing cursor exhaustion/termination semantics, which is the main remaining behavioral unknown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste, front-loaded with what the tool does before the when-to-use rule and the return shape. Nothing is padded or repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully describes the shape of each returned item (speaker + timestamp line). Combined with the schema's parameter docs, an agent has enough to call it correctly; only pagination termination behavior is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so cursor, noteId and pageSize are already fully documented in the schema (including the default page size of 50 and the not_ ID format). The description only implies paging and adds no syntax or format detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (read) and resource (meeting transcript), plus the paging scope ('one page at a time'). It also implicitly distinguishes itself from the sibling notes_get by being the paginated path for the same resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing is given: 'Use this for any long transcript, and whenever notes_get answers 413 TRANSCRIPT_TOO_LARGE.' Both the trigger condition and the alternative tool are named, so the agent does not have to infer when to switch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gsheets_spreadsheets_values_appendC
Appends values to a spreadsheet.
| Name | Required | Description | Default |
|---|---|---|---|
| range | Yes | The A1 notation of a range to search for a logical table of data. Values are appended after the last row of the table. | |
| valueRange | Yes | The request body contains an instance of ValueRange. | |
| spreadsheetId | Yes | The ID of the spreadsheet to update. | |
| insertDataOption | No | How the input data should be inserted. | |
| valueInputOption | Yes | How the input data should be interpreted. | |
| includeValuesInResponse | No | Determines if the update response should include the values of the cells that were appended. By default, responses do not include the updated values. | |
| responseValueRenderOption | No | Determines how values in the response should be rendered. The default render option is FORMATTED_VALUE. | |
| responseDateTimeRenderOption | No | Determines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=false and openWorldHint=true, indicating a mutating, external operation. The description adds nothing beyond this – it doesn't disclose how appended data interacts with existing tables, whether headers are auto-detected, or rate-limit/permission requirements. For a mutation tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single, front-loaded sentence that wastes no words. However, the extreme brevity contributes to gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, nested value types, mutation behavior), the description is critically underspecified. It omits key behavioral details like how the append range is determined, interaction with existing data, and available options (insertDataOption, valueInputOption). No output schema exists to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds no parameter-level detail beyond what's already provided. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (appends) and resource (values to a spreadsheet), which is clear but does not differentiate from the sibling gsheets_spreadsheets_values_update or clarify the 'logical table' append behavior. It's clear but lacks sibling differentiation within the gsheets tool family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives like gsheets_spreadsheets_values_update, and no mention of required preconditions or idempotency considerations. The description is silent on usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gsheets_spreadsheets_values_updateC
Sets values in a range of a spreadsheet.
| Name | Required | Description | Default |
|---|---|---|---|
| range | Yes | The A1 notation of the values to update. | |
| valueRange | Yes | The request body contains an instance of ValueRange. | |
| spreadsheetId | Yes | The ID of the spreadsheet to update. | |
| valueInputOption | Yes | How the input data should be interpreted. | |
| includeValuesInResponse | No | Determines if the update response should include the values of the cells that were updated. By default, responses do not include the updated values. If the range to write was larger than the range actually written, the response includes all values in the requested range (excluding trailing empty rows and columns). | |
| responseValueRenderOption | No | Determines how values in the response should be rendered. The default render option is FORMATTED_VALUE. | |
| responseDateTimeRenderOption | No | Determines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, so the write nature is known. But the description adds nothing beyond the name – it does not disclose that existing cell values are overwritten, how valueInputOption affects interpretation, or what the update returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is tight. But for a 7-parameter mutation tool it is under-specified rather than appropriately sized; conciseness here is closer to omission.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with 7 params and no output schema, the description omits overwrite semantics, auth requirements, and response behavior. An agent could call it, but not safely without reading the schema closely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with 7 well-documented parameters, so the schema carries the meaning. The description adds no parameter detail beyond 'in a range', which is baseline 3 for fully documented schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (sets values in a range of a spreadsheet), which an agent can distinguish from gsheets_spreadsheets_values_get by direction of data flow. However it offers no explicit sibling differentiation and largely restates the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the sibling gsheets_spreadsheets_values_get or when reading vs writing applies. The agent must infer context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linear_issue_createA
Create an issue. team_id and title are required; everything else is optional. The UUIDs for team, assignee, state and labels come from teams_list, users_list and workflow_states_list — Linear does not accept names here. Set parent_id to create a sub-issue.
| Name | Required | Description | Default |
|---|---|---|---|
| variables | Yes | The issue to create. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the mutation/external-scope profile is covered. The description adds genuinely useful behavior: Linear rejects names and only accepts UUIDs resolved via teams_list/users_list/workflow_states_list, and parent_id turns this into a sub-issue creation. It stops short of noting side effects like notifications or returned identity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and the required fields, followed by the ID-resolution constraint and the sub-issue tip. No filler, though the UUID guidance partially duplicates the schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool whose one parameter is a deeply nested object fully documented by the schema, plus annotations covering the safety profile, the description covers the essentials an agent needs. It lacks any mention of what creation returns, which is a minor gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every field including the resolver-tool hints and the sub-issue semantics. The description's parameter notes (team_id/title required, everything else optional) largely restate the schema rather than adding new meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create an issue'), which is instantly distinguishable from the list-oriented sibling linear_issues_list. It does not explicitly name a sibling it is not, so it falls short of a 5, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives practical invocation guidance (required vs optional fields, which resolver tools supply UUIDs, how to make a sub-issue) but never states when to reach for this tool versus alternatives such as linear_issues_list, nor any exclusion or prerequisite conditions. Usage is implied rather than framed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linear_issues_listARead-onlyIdempotent
List issues, optionally filtered. Conditions on one filter object combine with AND. To find a team's open work, filter on team.key and state.type.
| Name | Required | Description | Default |
|---|---|---|---|
| variables | No | Paging and filtering. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered by structured data. The description's contribution is the AND-combination rule for filters, but that same rule is already documented in the schema's filter field, so added value is minimal beyond confirming the semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the core purpose front-loaded and the practical example last. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with a fully documented filter schema and no output schema, the description covers purpose, filter semantics, and a usage pattern. Slightly short on pagination/ordering behavior, though that is documented in the schema's variables object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the nested comparators are fully documented, so the schema does the heavy lifting. The description's filter guidance (AND combination, team.key/state.type paths) largely repeats what the schema already states, making 3 the appropriate baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('issues') with scope note 'optionally filtered'. It distinguishes itself from the sibling write tool linear_issue_create implicitly, but never names a sibling or explicitly rule out other Linear tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when-to-use example: 'To find a team's open work, filter on team.key and state.type.' It provides clear positive usage context but no when-not guidance or named alternative for other listing scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notion_data_sources_queryB
Get the rows of a data source, optionally filtered and sorted. The filter grammar is one condition per column type, composed with and and or up to two levels deep. Read the schema first if you do not know the column names — a filter naming a column that is not there is a 400.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | The filter, sort and paging options. All optional. | |
| dataSourceId | Yes | The ID of the data source. | |
| filterProperties | No | Property IDs to return on each row, instead of all of them. The cheapest way to keep a wide table's query readable. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description frames the operation as 'Get the rows', i.e., a read, yet the annotations declare readOnlyHint=false, which tells the agent state can be mutated. That is a direct conflict about side effects. The extra context about filter nesting depth and the 400 on unknown columns is useful, but the contradiction in the safety profile dominates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, all front-loaded: purpose first, grammar constraint second, prerequisite/pitfall last. No filler and every sentence contributes information an agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested-query tool the description covers the filter grammar and one failure mode, but says nothing about pagination (pageSize/startCursor are in the schema) or what the response looks like, and there is no output schema to fall back on. It is adequate but leaves meaningful operational gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents dataSourceId, body, and filterProperties fully. The description restates the filter-composition rule (one condition per type, and/or two levels deep) that the schema already documents, adding only the 400-on-unknown-column error behavior. Baseline 3 is appropriate when the schema carries the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Get the rows of a data source', with the qualifier that results can be filtered and sorted. An agent immediately understands this is a read/query operation over tabular Notion data. It does not, however, name or contrast itself against sibling tools such as notion_search, so sibling differentiation is absent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is largely implied by 'optionally filtered and sorted', and it adds a genuine prerequisite: 'Read the schema first if you do not know the column names'. It also warns that a filter referencing a non-existent column yields a 400, which steers the agent toward a correct call. What is missing is explicit when-not-to-use or an alternative tool to prefer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notion_pages_createA
Create a page — as a subpage of another page, or as a row of a database by giving its data_source_id as the parent. Content comes as a markdown string Notion parses into blocks, or from a template: one or the other, never both.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The page to create. | |
| filterProperties | No | Property IDs to return on the page that comes back, instead of all of them. A page that does not have a listed property omits it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, establishing this as an external write. The description adds the meaningful markdown/template mutual exclusion. It does not disclose auth requirements, rate limits, or the allowAsync async-202 behavior, but with annotations carrying the safety profile a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with the verb and the two modes, with the exclusivity constraint phrased crisply. No wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex creation tool with a fully documented schema, the description covers parent selection and content sourcing adequately. It omits the async task path (allowAsync → 202) and return shape, but with no output schema and 100% schema coverage these are minor gaps, and annotations cover the write semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value: it clarifies that `parent` takes `data_source_id` to make a database row and that template vs. markdown are mutually exclusive — beyond the schema's raw field docs. It doesn't add detail on `properties` or `filterProperties`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a page') and immediately delineates the two creation modes — subpage vs. database row — which is exactly what separates it from siblings like notion_pages_update. An agent can identify the tool's scope without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance: use a page parent for a subpage, `data_source_id` for a database row, and content comes from `markdown` OR a template, 'never both.' The mutual-exclusion rule is explicit. However, it names no sibling alternatives (e.g., update vs. create routing) and states no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_chat_post_messageA
Send a message to a Slack channel, private group, or DM. Provide text for a plain message; set thread_ts to reply inside an existing thread.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | The main body text of the message. Required unless blocks or attachments are provided. Used as the fallback string for notifications when blocks are provided, so it is worth setting even then. | |
| parse | No | Change how messages are treated. Accepts 'none' or 'full'. | |
| blocks | No | A JSON-based array of structured Block Kit blocks. | |
| mrkdwn | No | Disable Slack markup parsing by setting to false. Defaults to true. | |
| channel | Yes | An encoded ID or channel name that represents a channel, private group, or IM channel to send the message to. Prefer the encoded ID (e.g. 'C123ABC456'). | |
| iconUrl | No | URL to an image to use as the icon for this message. Requires the chat:write.customize scope. | |
| metadata | No | Application-specific metadata to attach to the message. | |
| threadTs | No | Provide another message's 'ts' value to make this message a reply in that thread. Avoid using a reply's ts value; use the parent's. | |
| username | No | Set the bot's user name. Requires the chat:write.customize scope. | |
| iconEmoji | No | Emoji to use as the icon for this message, e.g. ':chart_with_upwards_trend:'. Requires the chat:write.customize scope. | |
| linkNames | No | Find and link user groups. | |
| attachments | No | A JSON-based array of structured attachments. | |
| unfurlLinks | No | Pass true to enable unfurling of primarily text-based content. | |
| unfurlMedia | No | Pass false to disable unfurling of media content. | |
| markdownText | No | Accepts message text formatted in markdown. Limit this field to 12,000 characters. Cannot be used together with blocks or text. | |
| replyBroadcast | No | Used in conjunction with thread_ts and indicates whether the reply should be made visible to everyone in the channel. Defaults to false. | |
| unfurlAppLinks | No | Pass true to enable unfurling of links to installed apps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external nature is covered. The description adds the threading behavior, but does not disclose required scopes, rate limits, message-size limits, or what a successful send returns — and most scope info already lives in the schema parameter descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and then the two most important parameter behaviors. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter tool with no output schema, the description is thin but the schema carries full parameter documentation, so an agent can call it correctly. Missing behavioral context (rate limits, required scopes, response shape) keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 17 parameters are documented in the schema itself. The description's notes on `text` and `thread_ts` largely restate what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Send') and resource ('a message to a Slack channel, private group, or DM'), making the action and destination unambiguous. An agent can distinguish this from siblings like slack_conversations_create without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives implied usage for `text` and `thread_ts`, but there is no explicit when-to-use vs. when-not, no mention of prerequisites (e.g. chat:write scope), and no reference to alternative messaging tools such as gmail_messages_send. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
slack_chat_schedule_messageA
Schedule a message for later, up to 120 days ahead. The returned scheduled_message_id is what cancels it.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The message text, as Slack markup. | |
| postAt | Yes | Unix timestamp representing the future time the message should post to Slack. Seconds, not milliseconds: 1700000000, not 1700000000000. At most 120 days ahead, beyond which Slack answers `time_too_far`, and at most thirty scheduled messages per channel in any five minutes. | |
| channel | Yes | The channel to post in. | |
| threadTs | No | Post into this thread rather than the channel. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries part of the burden and does so usefully: it discloses the 120-day ceiling and, critically, that the returned scheduled_message_id is the cancellation handle. It does not restate the rate-limit or timestamp-unit details, which live in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the scoping constraint front-loaded and the cancellation handle following. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully names the return value (scheduled_message_id) an agent needs for follow-up. Rate limits, timestamp units and the time_too_far failure mode are covered by the schema, so remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents channel, postAt, text and threadTs in depth. The description adds no parameter-level detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Schedule) and resource (a message) plus the horizon constraint (up to 120 days ahead), so an agent immediately knows this defers delivery rather than posting now. It never names slack_chat_post_message, so the distinction from that sibling is left implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for later' implies when to use it versus immediate posting, but there is no explicit when/when-not guidance and no named alternative (slack_chat_post_message). Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_checkout_sessions_listARead-onlyIdempotent
List checkout sessions. Stripe has no payment_status filter, so find paid ones by reading the results.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | A limit on the number of objects to be returned, between 1 and 100. Defaults to 10. | |
| status | No | Only return sessions in this state. 'complete' means finished, not necessarily paid. | |
| created | No | Only return sessions created in this window. | |
| customer | No | Only return sessions for this customer. | |
| paymentLink | No | Only return sessions created by this payment link. | |
| endingBefore | No | A cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after. | |
| subscription | No | Return the session that created this subscription. | |
| paymentIntent | No | Return the session for this PaymentIntent. | |
| startingAfter | No | A cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds a genuine behavioral caveat (no payment_status filter, must inspect results), but that partially overlaps with the schema's own note that 'complete' means finished, not necessarily paid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste; the core purpose is front-loaded and the caveat follows immediately. Nothing could be cut without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only paginated list tool with a fully documented schema and annotations, the essentials are present. However, with no output schema, the description does not indicate the shape of returned sessions or pagination behavior, leaving a visible gap for an agent planning to scan results for paid sessions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 9 parameters, including timestamp units and pagination cursors, so the schema does the heavy lifting. The description adds no parameter-level detail, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List checkout sessions"), which is unambiguous against the sibling stripe_customers_list. It does not explicitly name an alternative or scope, so it stops short of the sibling-differentiation a 5 would require.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence implies a usage pattern — when you want paid sessions, you must read the results because there is no payment_status filter. That is useful implied guidance, but there is no explicit when-to-use/when-not-to-use framing or named alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stripe_customers_listARead-onlyIdempotent
List customers, most recently created first. Filter by email to find one.
| Name | Required | Description | Default |
|---|---|---|---|
| No | A case-sensitive filter on the list based on the customer's email field. The value must be a string. | ||
| limit | No | A limit on the number of objects to be returned, between 1 and 100. Defaults to 10. | |
| endingBefore | No | A cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after. | |
| startingAfter | No | A cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered without description help. The description adds the non-obvious ordering guarantee (most recently created first), which is genuinely useful, but says nothing about rate limits, page size defaults, or result shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the ordering behavior is front-loaded before the filtering hint. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no output schema, the description plus the rich parameter schema and annotations give an agent enough to call it correctly. The main omission is any note about pagination workflow or default result count, though the schema covers the cursor mechanics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (email, limit, endingBefore, startingAfter) are already fully documented in the schema. The description echoes the email filter without adding syntax, case-sensitivity, or pagination semantics beyond what the schema states, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (customers) plus a default sort order, so the agent knows exactly what it retrieves. It does not explicitly distinguish itself from the sibling stripe_checkout_sessions_list, but the resource noun is unambiguous enough to route correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Filter by email to find one' implies the lookup use case, giving some usage context. However, there is no guidance on when to use this versus other Stripe list endpoints, and no mention of pagination workflow for iterating beyond the default limit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tavily_searchA
Execute a real-time web search optimized for AI agents. Use when sources are unknown or current web context is needed. Prefer search_depth advanced with chunks_per_source 3 for stronger evidence per source.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query to execute. | |
| topic | No | Search category. The server applies `general` when this is absent. `news` automatically enables `include_published_date`. | |
| country | No | Boost results from a country using Tavily's lowercase English country name (for example `united states`). Available only when `topic` is `general`. | |
| endDate | No | Return results before this date (`YYYY-MM-DD`). | |
| language | No | Boost or filter results by language — an ISO 639-1 code (for example `en`, `fr`, `zh-cn`) or English language name (for example `english`, `french`). | |
| startDate | No | Return results after this date (`YYYY-MM-DD`). | |
| timeRange | No | Filter by publish or last-updated date window. | |
| exactMatch | No | Return only results containing the exact quoted phrase(s) in the query. | |
| maxResults | No | Maximum search results to return. The server applies 10 when this is absent. | |
| safeSearch | No | Filter adult or unsafe content. Not supported when `search_depth` is `fast` or `ultra-fast`. | |
| searchDepth | No | Latency/relevance tradeoff. The server applies `basic` when this is absent. `advanced` costs 2 credits; `basic`, `fast` and `ultra-fast` cost 1 credit. | |
| includeUsage | No | Include credit usage in the response. | |
| includeAnswer | No | Include an LLM-generated answer. `true` or `basic` returns a quick answer; `advanced` returns a detailed answer. The server applies `false` when this is absent. | |
| includeImages | No | Include query-related images and per-result `images`. | |
| autoParameters | No | Let Tavily configure parameters from the query. Explicit values override auto-selected ones. `include_answer`, `include_raw_content` and `max_results` must always be set manually when using this. | |
| excludeDomains | No | Domains to exclude (max 150). | |
| includeDomains | No | Domains to include (max 300). | |
| includeFavicon | No | Include a favicon URL per result. | |
| chunksPerSource | No | Maximum relevant chunks per source in each result's `content`. The server applies 3 when this is absent. Available only when `search_depth` is `advanced`, `basic` or `fast`. Each chunk is at most 500 characters and joined with `[...]`. | |
| filterByLanguage | No | Strictly filter out non-matching languages. Requires `language`. | |
| includeRawContent | No | Include cleaned page content per result. `true` or `markdown` returns markdown; `text` returns plain text and may increase latency. The server applies `false` when this is absent. | |
| includeDomainsMode | No | How `include_domains` is applied. Requires `include_domains` to be set. | |
| includePublishedDate | No | Include `published_date` on each result. Beta feature. Automatically enabled when `topic` is `news`. | |
| filterByPublishedDate | No | Remove results outside the date window or with no detectable date. Also enables `include_published_date`. | |
| includeImageDescriptions | No | Add descriptive text per image when `include_images` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, which is an unusual pairing for a read-only search and is left unexplained. The description adds a config recommendation (search_depth advanced, chunks_per_source 3) but doesn't clarify the credit costs or why a read-only operation isn't marked read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it does, when to use it, and an actionable configuration tip. Front-loaded and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a large but fully documented 25-parameter schema with no output schema, the description covers the essential purpose, usage trigger, and a key tuning recommendation. It's adequate, though the odd readOnlyHint=false on a search tool could have been addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 25 parameters is already documented in the schema. The description names two parameters and their preferred values without adding semantics beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Execute a real-time web search') and scopes it ('optimized for AI agents'). It distinguishes itself from tavily_research_create by being a real-time search vs. a research task, though it never explicitly names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage trigger: 'Use when sources are unknown or current web context is needed.' This tells the agent when to reach for it over knowledge-only answers, though it doesn't name alternatives like firecrawl_scrape or tavily_research_create.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
29 tool updates
v0.1.0- First observed
connect - First observed
connection_status - First observed
firecrawl_crawl - First observed
firecrawl_crawl_status - First observed
firecrawl_extract - First observed
firecrawl_map - First observed
firecrawl_monitor_create - First observed
firecrawl_scrape - First observed
gcalendar_events_insert - First observed
gdocs_documents_create - First observed
gforms_forms_create - First observed
gforms_forms_responses_list - First observed
github_releases_list - First observed
github_repos_list_commits - First observed
gmail_drafts_create - First observed
gmail_messages_send - First observed
granola_notes_list - First observed
granola_notes_transcript_get - First observed
gsheets_spreadsheets_values_append - First observed
gsheets_spreadsheets_values_update - First observed
linear_issue_create - First observed
linear_issues_list - First observed
notion_data_sources_query - First observed
notion_pages_create - First observed
slack_chat_post_message - First observed
slack_chat_schedule_message - First observed
stripe_checkout_sessions_list - First observed
stripe_customers_list - First observed
tavily_search
TDQS
Scored across 29 tools
Tools are namespaced by app (linear_, github_, firecrawl_, etc.) with a distinct resource+action per tool, so cross-app confusion is minimal. The only mild overlap is within firecrawl (scrape's 'optional structured extraction' vs extract) and the meta pair connect/connection_status, but descriptions differentiate them clearly.
The dominant pattern is consistent app_resource_action with the verb last (linear_issues_list, gmail_messages_send, stripe_customers_list). Deviations are minor: inconsistent pluralization (linear_issues_list vs linear_issue_create) and the two meta tools (connect, connection_status) that abandon the pattern entirely.
At 29 tools this is on the heavy side of the rubric, though the breadth is explained by spanning 13+ distinct apps, each contributing only 1-3 tools. It feels borderline: justified by multi-app scope but still a large surface an agent must navigate.
Coverage is a thin slice per app and several descriptions point to tools that do not exist in this set: linear_issue_create references teams_list/users_list/workflow_states_list, gdocs_documents_create references documents_batch_update, gforms_forms_create references forms_batch_update, and granola_notes_list references notes_get. These dangling references create dead ends that will cause agent failures.
Maintenance
Related MCP Connectors
Plan and run paced startup, SaaS and AI directory launches from your agent, in your own Chrome.
Marketing MCP: your AI agent runs your organic growth loop — you approve before it ships.
AI social media team for founders: make and publish posts, carousels and short video to LinkedIn, X, Instagram, TikTok, YouTube, Facebook, Threads, Reddit and Bluesky, from any agent. OAuth sign-in, no API key; every draft waits for approval.
Go-to-market tools for AI agents: publish and schedule posts, find customers, SEO, outreach.
Related MCP Servers
- AlicenseBqualityBmaintenanceRuns 20 slash-command workflows across Google Calendar, Gmail, Linear, Slack, Granola, Google Docs, GitHub, Stripe, Notion, Google Forms, Sheets and Drive to produce morning briefs, meeting prep, action items assigned to owners and weekly updates. Reads proceed without asking, while anything that creates, sends, changes or deletes is shown for approval first.47Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables freelancers and agencies to run eight back-office workflows across Stripe, Google Drive, Linear, Google Calendar, Gmail, GitHub, Google Docs, Granola, Google Sheets and Firecrawl, covering client onboarding, invoices from calendar and commits, status reports, scope-creep detection, site audits and overdue invoice chasers. Reads run freely, while anything that creates, sends, changes or deletes is shown for approval first, with credentials kept in your own OS keychain and no proxying through any third-party server.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceAn MCP server that turns CI failures, pull requests, security alerts, and Slack threads into Linear issues, Slack posts, changelogs, docs, and reports with evidence attached, through 25 agent-run workflows over 59 tools spanning GitHub, Linear, Slack, Notion, Google Docs, Google Calendar, Google Sheets, Firecrawl, and Tavily. It also handles credential setup in the OS keychain and pauses for approval before any create, send, change, or delete call.Apache 2.0
- AlicenseBqualityBmaintenanceEnables an agent to run 20 product-management workflows that convert call notes, specs, customer feedback and competitor changes into deduped Linear issues, PRD drafts, roadmap sheets, digests and weekly project updates. Orchestrates only the tools each workflow needs across Granola, Linear, Slack, Notion, GitHub, Stripe, Firecrawl, Tavily and Google Workspace, using credentials kept in the user's own OS keychain.51Apache 2.0