Skip to main content
Glama
r28ai

stripe-billing-ops-mcp

by r28ai

Stripe Billing Ops

Invoices from real work, failed-payment follow-ups, revenue sheets and dispute evidence.

An MCP server with 14 workflows across Google Calendar, GitHub, Stripe, Gmail, Slack, Google Sheets, Shopify, Google Drive, Linear, Google Docs and Notion. Each workflow is a prompt your agent runs as a slash command, over the 37 tools it needs and no others.

uv tool install https://github.com/r28ai/stripe-billing-ops-mcp/releases/download/v0.1.0/stripe_billing_ops_mcp-0.1.0-py3-none-any.whl
claude mcp add billing -- stripe-billing-ops-mcp

It installs with uv from this repository's release, with no git and nothing to build; nothing but Charter and the libraries it uses comes from PyPI. To update, run the install line from the latest release. If a desktop app cannot find stripe-billing-ops-mcp, give it the full path from which stripe-billing-ops-mcp (where stripe-billing-ops-mcp on Windows).

Then ask your agent to connect your apps, or run /mcp__billing__setup.

Connect your apps

Ask the agent to connect one ("connect Linear"). It tells you where to get that app's key and the command that stores it, and the next call works, with no restart. The agent never asks for a key in the chat.

Or connect everything this server uses from a terminal:

stripe-billing-ops-mcp login            # each app in turn
stripe-billing-ops-mcp login github     # just one
stripe-billing-ops-mcp status           # what is connected

Tokens and keys go to your operating system's keychain (macOS Keychain, Windows Credential Manager, the Secret Service on Linux), and are checked with one read-only call to the app's own API before they are kept. Every key, token and OAuth client is yours: we register no app with any of these services, and nothing passes through a server of ours, because there isn't one.

App

How it connects

Or set

Google

Browser sign-in, over your own OAuth client (make one).

GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET

GitHub

Your own key (get one), entered once.

GITHUB_TOKEN

Stripe

Your own key (get one), entered once.

STRIPE_API_KEY

Slack

Your own key (get one), entered once. A bot token from your own Slack app, which the guide sets up in about three minutes.

SLACK_BOT_TOKEN

Shopify

Your own key (get one), entered once.

SHOPIFY_SHOP, SHOPIFY_CLIENT_ID, SHOPIFY_CLIENT_SECRET

Linear

Your own key (get one), entered once.

LINEAR_API_KEY

Notion

Your own key (get one), entered once. Then share the pages it should see with the integration.

NOTION_API_KEY

A variable set in your client's config always wins over the keychain.

Related MCP server: Subotiz MCP

Workflows

Workflow

What you get

Apps

Freelancer invoice from the week's work freelancer_invoice_from_the_week_s_work

Client meetings and commits become line items and a sent invoice.

Google Calendar, GitHub, Stripe

Failed payment follow-up failed_payment_follow_up

Past-due invoices get a human email and the account owner is told.

Stripe, Gmail, Slack

Revenue sheet revenue_sheet

MRR, new, churned and fees appended daily to the sheet finance already uses.

Stripe, Google Sheets

Payout reconciliation payout_reconciliation

Each payout broken into the charges, refunds and fees inside it.

Stripe, Google Sheets

Dispute evidence pack dispute_evidence_pack

Correspondence and fulfillment proof assembled and submitted before the deadline.

Stripe, Gmail, Shopify

Receipts → Drive and the expense sheet receipts_to_drive_and_the_expense_sheet

Every receipt filed and logged without forwarding emails to anyone.

Gmail, Google Drive, Google Sheets

Vendor invoice inbox → approval vendor_invoice_inbox_to_approval

Bills land in a sheet with the PDF and an approval request to the budget owner.

Gmail, Google Drive, Google Sheets, Slack

Plan change by email plan_change_by_email

'Please move us to annual' handled from the email, with confirmation drafted.

Gmail, Stripe

Investor update draft investor_update_draft

Numbers from billing and progress from the tracker, in your template.

Stripe, Linear, Google Docs, Gmail

Price change rollout price_change_rollout

New price created, affected customers listed and notice drafted, with an FAQ page.

Stripe, Gmail, Notion

Platform fee report platform_fee_report

Fees per connected account per month, for marketplaces on Connect.

Stripe, Google Sheets

SaaS spend audit saas_spend_audit

Every recurring vendor charge in the inbox, with an owner asked to justify it.

Gmail, Google Sheets, Slack

Bulk refunds from a sheet bulk_refunds_from_a_sheet

A list of affected orders after an incident, refunded and marked row by row.

Google Sheets, Stripe

Price list from a sheet price_list_from_a_sheet

The pricing sheet the team agreed on becomes the products Stripe sells.

Google Sheets, Stripe

Every prompt takes one optional argument, details: the repo, team, channel, customer or date range you mean, so the agent does not have to ask. In Claude Code, put it in quotes, or only its first word arrives:

/mcp__billing__freelancer_invoice_from_the_week_s_work "client Acme, repo acme/site, last week"

Reads run without asking. Before anything that creates, sends, changes or deletes, the prompt tells the agent to show you the call and wait.

0 of the 14 workflows need no Google or Granola credential.

Other clients

Claude Desktop: install uv if you have not, since Claude Desktop starts the server with it, then open the .mcpb from the latest release. Claude asks for any keys in its own settings and keeps them in your keychain. The first start takes a few seconds longer, while uv installs it.

VS Code (.vscode/mcp.json): VS Code asks for each key the first time the server starts and stores it securely. Leave out any you stored with login.

{
  "inputs": [
    {
      "type": "promptString",
      "id": "google-client-secret",
      "description": "Google: OAuth client secret",
      "password": true
    },
    {
      "type": "promptString",
      "id": "github-token",
      "description": "GitHub: Personal access token",
      "password": true
    },
    {
      "type": "promptString",
      "id": "stripe-api-key",
      "description": "Stripe: Secret or restricted key",
      "password": true
    },
    {
      "type": "promptString",
      "id": "slack-bot-token",
      "description": "Slack: Bot token (xoxb-\u2026)",
      "password": true
    },
    {
      "type": "promptString",
      "id": "shopify-client-secret",
      "description": "Shopify: App client secret",
      "password": true
    },
    {
      "type": "promptString",
      "id": "linear-api-key",
      "description": "Linear: Personal API key",
      "password": true
    },
    {
      "type": "promptString",
      "id": "notion-api-key",
      "description": "Notion: Integration secret (ntn_\u2026)",
      "password": true
    }
  ],
  "servers": {
    "billing": {
      "type": "stdio",
      "command": "stripe-billing-ops-mcp",
      "env": {
        "GOOGLE_CLIENT_SECRET": "${input:google-client-secret}",
        "GITHUB_TOKEN": "${input:github-token}",
        "STRIPE_API_KEY": "${input:stripe-api-key}",
        "SLACK_BOT_TOKEN": "${input:slack-bot-token}",
        "SHOPIFY_CLIENT_SECRET": "${input:shopify-client-secret}",
        "LINEAR_API_KEY": "${input:linear-api-key}",
        "NOTION_API_KEY": "${input:notion-api-key}",
        "GOOGLE_CLIENT_ID": "",
        "SHOPIFY_SHOP": "",
        "SHOPIFY_CLIENT_ID": ""
      }
    }
  }
}

Cursor (.cursor/mcp.json) starts it the same way:

{
  "mcpServers": {
    "billing": {
      "command": "stripe-billing-ops-mcp"
    }
  }
}

Codex (~/.codex/config.toml) starts a turn without waiting for a server unless it is required, and then the agent has none of its tools. required = true makes the session wait for it, and startup_readiness = "catalog" waits for its tool list rather than just its connection:

[mcp_servers.billing]
command = "stripe-billing-ops-mcp"
required = true
startup_readiness = "catalog"
startup_timeout_sec = 30

Name the server billing. A host builds each tool's name from that key, and a longer one can push a tool past the 64 characters a function name allows.

Built with Charter

Every tool here is a Charter declaration: a Pydantic schema saying where each field goes on the wire. Charter's runtime builds the request, attaches and refreshes the credential, and trims the response before the model reads it. It runs in your process, with no proxy and no telemetry.

The 37 tool schemas come to 32,273 tokens.

The same tools work in your own agent, without MCP:

from charter.adapters.openai import to_openai_tools
from charter_packs_mcp import FAMILIES

tools = FAMILIES["finance"].tools()
definitions = to_openai_tools(tools)   # or charter.adapters.langchain

Need an API that isn't here? Write a pack: your coding agent writes the declarations, and Charter's conformance suite checks them.

  • Google Calendar: gcalendar_events_list

  • GitHub: github_search_commits

  • Stripe: stripe_invoice_items_create, stripe_invoices_create, stripe_invoices_send, stripe_invoices_list, stripe_customers_retrieve, stripe_subscriptions_list, stripe_balance_transactions_list, stripe_payouts_list, stripe_disputes_list, stripe_disputes_retrieve, stripe_disputes_update, stripe_customers_list, stripe_subscriptions_update, stripe_balance_retrieve, stripe_prices_create, stripe_accounts_list, stripe_application_fees_list, stripe_charges_list, stripe_refunds_create, stripe_products_create

  • Gmail: gmail_drafts_create, gmail_threads_list, gmail_messages_list, gmail_messages_attachments_get, gmail_threads_get

  • Slack: slack_chat_post_message, slack_users_list

  • Google Sheets: gsheets_spreadsheets_values_append, gsheets_spreadsheets_values_update, gsheets_spreadsheets_values_get

  • Shopify: shopify_order_get

  • Google Drive: gdrive_files_create

  • Linear: linear_projects_list

  • Google Docs: gdocs_documents_create

  • Notion: notion_pages_create

License

Apache 2.0.

Available Tools

39 tools
connectA

Connect one app this server uses. For an app that issues keys, says where to get one and the terminal command that stores it. For Google, once the user's own OAuth client is set, starts the browser sign-in and returns at once: the user approves in the browser and the next call works. To see which apps are connected, call connection_status. Never ask the user for a key in the chat.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesThe app to connect.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=false; the description adds real value by disclosing that Google sign-in is asynchronous ('returns at once: the user approves in the browser and the next call works') and that key-based apps require a stored terminal command. It doesn't cover failure modes or what happens on partial/again-called connections.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then per-app caveats, then the routing hint and guardrail. Four sentences each carry distinct information, though the key-storage mechanics sentence is denser than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one enum parameter fully documented and no output schema, the description carries the behavioral load well, including the async sign-in semantics and the follow-up call. Lacking any mention of errors or retry behavior for a credential-storing mutation keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and only says 'The app to connect', so baseline is 3; the description goes beyond it by explaining that behavior differs per app (key-issuing apps vs the special Google OAuth flow), which meaningfully informs how the enum value changes the interaction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (connect) plus resource (one app this server uses) and clarifies the consequence, that credentials get stored. It explicitly routes away from the sibling by naming `connection_status` for the read case, so an agent can distinguish it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit guidance for the read alternative ('To see which apps are connected, call `connection_status`') and a guardrail ('Never ask the user for a key in the chat'). It does not state conditions under which connecting should be avoided or deferred, so it stops short of full when/when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connection_statusA
Read-only

See which apps this server is connected to, and how to connect each one that is not. Changes nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered by structured data. 'Changes nothing' restates the readOnly hint rather than adding new behavior; the only incremental value is noting that connect instructions are returned for unconnected apps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the primary purpose and appends the secondary benefit with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, read-only status tool with no output schema, the description covers both what is inspected and the shape of the useful payload (connect guidance). Return format details are absent but minimal given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly implies no input is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('See') and resource ('which apps this server is connected to'), and adds the secondary payload of connect instructions for missing apps. This distinguishes it from the sibling 'connect' tool, which performs the connection rather than reporting status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'how to connect each one that is not' implies this tool is the discovery step before using 'connect', but the sibling is never named and there is no explicit when-to-use/when-not statement. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gcalendar_events_listC
Read-onlyIdempotent

List events matching a given search filter.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoFree text search terms to find events that match these terms in the following fields: summary, description, location, attendee's displayName, attendee's email, organizer's displayName, organizer's email, workingLocationProperties.officeLocation.buildingId, workingLocationProperties.officeLocation.deskId, workingLocationProperties.officeLocation.label, workingLocationProperties.customLocation.label. These search terms also match predefined keywords against all display title translations of working location, out-of-office, and focus-time events. Optional.
iCalUIDNoSpecifies an event ID in the iCalendar format to be provided in the response. Optional. Use this if you want to search for an event by its iCalendar ID.
orderByNoThe order of the events returned in the result. Optional. The default is an unspecified, stable order. Acceptable values are: "startTime" (order by the start date/time, ascending; this is only available when querying single events, i.e. the parameter singleEvents is True), "updated" (order by last modification time, ascending).
timeMaxNoUpper bound (exclusive) for an event's start time to filter by. Optional. The default is not to filter by start time. Must be an RFC3339 timestamp with mandatory time zone offset, for example, 2011-06-03T10:00:00-07:00, 2011-06-03T10:00:00Z. Milliseconds may be provided but are ignored. If timeMin is set, timeMax must be greater than timeMin.
timeMinNoLower bound (exclusive) for an event's end time to filter by. Optional. The default is not to filter by end time. Must be an RFC3339 timestamp with mandatory time zone offset, for example, 2011-06-03T10:00:00-07:00, 2011-06-03T10:00:00Z. Milliseconds may be provided but are ignored. If timeMax is set, timeMin must be smaller than timeMax.
timeZoneNoTime zone used in the response. Optional. The default is the time zone of the calendar.
pageTokenNoToken specifying which result page to return. Optional.
syncTokenNoToken obtained from the nextSyncToken field returned on the last page of results from the previous list request. It makes the result of this list request contain only entries that have changed since then. All events deleted since the previous list request will always be in the result set and it is not allowed to set showDeleted to False. There are several query parameters that cannot be specified together with nextSyncToken to ensure consistency of the client state. These are: iCalUID, orderBy, privateExtendedProperty, q, sharedExtendedProperty, timeMin, timeMax, updatedMin. All other query parameters should be the same as for the initial synchronization to avoid undefined behavior. If the syncToken expires, the server will respond with a 410 GONE response code and the client should clear its storage and perform a full synchronization without any syncToken. Optional. The default is to return all entries.
calendarIdYesCalendar identifier. To retrieve calendar IDs call the calendarList.list method. If you want to access the primary calendar of the currently logged in user, use the "primary" keyword.
eventTypesNoEvent types to return. Optional. This parameter can be repeated multiple times to return events of different types. If unset, returns all event types. Acceptable values are: "birthday" (special all-day events with an annual recurrence), "default" (regular events), "focusTime" (focus time events), "fromGmail" (events from Gmail), "outOfOffice" (out of office events), "workingLocation" (working location events).
maxResultsNoMaximum number of events returned on one result page. The number of events in the resulting page may be less than this value, or none at all, even if there are more events matching the query. Incomplete pages can be detected by a non-empty nextPageToken field in the response. By default the value is 250 events. The page size can never be larger than 2500 events. Optional.
updatedMinNoLower bound for an event's last modification time (as a RFC3339 timestamp) to filter by. When specified, entries deleted since this time will always be included regardless of showDeleted. Optional. The default is not to filter by last modification time.
showDeletedNoWhether to include deleted events (with status equals "cancelled") in the result. Cancelled instances of recurring events (but not the underlying recurring event) will still be included if showDeleted and singleEvents are both False. If showDeleted and singleEvents are both True, only single instances of deleted events (but not the underlying recurring events) are returned. Optional. The default is False.
maxAttendeesNoThe maximum number of attendees to include in the response. If there are more than the specified number of attendees, only the participant is returned. Optional.
singleEventsNoWhether to expand recurring events into instances and only return single one-off events and instances of recurring events, but not the underlying recurring events themselves. Optional. The default is False.
showHiddenInvitationsNoWhether to include hidden invitations in the result. Optional. The default is False.
sharedExtendedPropertyNoExtended properties constraint specified as propertyName=value. Matches only shared properties. This parameter might be repeated multiple times to return events that match all given constraints.
privateExtendedPropertyNoExtended properties constraint specified as propertyName=value. Matches only private properties. This parameter might be repeated multiple times to return events that match all given constraints.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that — no pagination behavior, no sync-token semantics, no note that results are scoped to a single calendar. It contributes no behavioral context of its own.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with zero waste, and the action is front-loaded. But the brevity is under-specification rather than disciplined conciseness — it omits any usable context an agent would want for a 18-parameter list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with 18 parameters, no output schema, and paging/sync behavior, a single sentence is insufficient. The rich schema mitigates some of this, but the description supplies no operational framing (single-calendar scope, pagination, recurring-event expansion) that an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 18 parameters, so the schema carries full semantic burden (time bounds, orderBy constraints, syncToken exclusions, etc.). The description adds no parameter meaning whatsoever, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (events) plus a scope qualifier ('matching a given search filter'). However, 'search filter' is vague against an 18-parameter schema and the description offers no differentiation from the sibling gcalendar_events_insert or any hint of what the tool actually filters on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives, no mention of prerequisites (calendarId required, calendarList.list to discover IDs), and no exclusions. The single sentence gives only the bare action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gdocs_documents_createA

Create a blank document with a title. Only the title is honoured — the document is created empty. To add content, call this and then documents_batch_update with the returned documentId.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoThe document to create. Only the title is honoured.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries the behavioral load and does well: it discloses the critical constraint that only the title is honoured and the document is created empty. It also notes the documentId is returned, which matters with no output schema. It stops short of noting auth/permission or quota behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, zero padding, and the most important constraint (empty document, title-only) is front-loaded before the follow-up instructions. Every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by mentioning the returned documentId and the batch_update follow-up, which is what an agent needs to chain calls. For a single-param mutation tool with annotations covering the safety profile, this is nearly complete; only auth/error behavior is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the title field already states it is the only honoured field, so the description largely restates structured data. Baseline 3 applies; the emphasis on the empty-document effect is useful but not new information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create) and resource (document), and immediately disambiguates the scope: it produces a blank/empty document, not a content-bearing one. This is a distinct action an agent can tell apart from content-writing operations. The mention of documents_batch_update further fixes its place in the workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear sequencing guidance: to add content, create first and then call documents_batch_update with the returned documentId. It does not state explicit exclusions or alternatives (e.g., when to use a copy/template flow instead), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gdrive_files_createA

Create a file's metadata: a folder (mimeType application/vnd.google-apps.folder), a Google Doc / Sheet / Slide, or an empty blob. Media upload is not expressed here. parents takes at most one folder ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesThe file resource to create. For a folder, set `mimeType` to `application/vnd.google-apps.folder`. For a Google Doc, Sheet or Slide, use the corresponding `application/vnd.google-apps.*` MIME type. `parents` takes at most one folder ID.
ocrLanguageNoA language hint for OCR processing during image import (ISO 639-1 code).
includeLabelsNoA comma-separated list of IDs of labels to include in the `labelInfo` part of the response.
supportsAllDrivesNoWhether the requesting application supports both My Drives and shared drives.
keepRevisionForeverNoWhether to set the `keepForever` field in the new head revision. This is only applicable to files with binary content in Google Drive. Only 200 revisions for the file can be kept forever. If the limit is reached, try deleting pinned revisions.
ignoreDefaultVisibilityNoWhether to ignore the domain's default visibility settings for the created file. Domain administrators can choose to make all uploaded files visible to the domain by default; this parameter bypasses that behavior for the request. Permissions are still inherited from parent folders.
includePermissionsForViewNoSpecifies which additional view's permissions to include in the response. Only `published` is supported.
useContentAsIndexableTextNoWhether to use the uploaded content as indexable text.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false (a write) and openWorldHint=true, so the safety profile is partly covered. The description adds a meaningful scope constraint — metadata only, not media upload — but says nothing about permissions, reversibility, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no filler. The core action and the two most consequential constraints (no media upload, single parent) appear immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given full schema coverage and existing annotations, the description is adequate for calling the tool. However, with no output schema, it omits any statement about what is returned (the created file resource), and gives no hint about required permissions or behavior when the parent is invalid.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters in detail. The mimeType values and the 'parents takes at most one folder ID' note in the description merely restate what the nested `file` parameter description already provides, adding no new syntax or format information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a file's metadata') and enumerates the three things it can create — a folder, a Google Doc/Sheet/Slide, or an empty blob. It does not, however, explicitly distinguish itself from a possible native-doc creator, so the differentiation is only partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Media upload is not expressed here' gives one negative scoping cue (don't use this to upload content), which is useful. But it names no alternative tool and gives no explicit when-to-use conditions, leaving usage largely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

github_search_commitsA
Read-onlyIdempotent

Search commits by message, author or date across GitHub — 'repo:owner/name fix flaky test'. The way to find where a change was introduced. Rate limited to 30 requests per minute.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesThe search query, with commit qualifiers: `repo:owner/name`, `author:LOGIN`, `committer-date:>2026-01-01`, `merge:false`. A bare term searches commit messages.
pageNoThe page number of the results to fetch. Defaults to 1.
sortNoSorts the results by author or committer date. Absent, results come back by best match.
orderNoDetermines whether the first search result returned is the highest number of matches (`desc`) or lowest (`asc`). Ignored unless `sort` is provided. GitHub uses `desc` when this is absent.
perPageNoThe number of results per page (max 100). Defaults to 30.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds a genuine behavioral constraint not present in structured data: a rate limit of 30 requests per minute, which an agent must respect when issuing repeated searches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scoped qualifiers, then a purpose statement and a rate-limit caveat. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with a fully documented schema, the description covers purpose, the query shape and a rate limit. It says nothing about result shape or pagination behavior, though the page/perPage params imply paging; with no output schema, a brief note on what a result contains would have closed the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters including the qualifier syntax for `q` are already documented. The example query 'repo:owner/name fix flaky test' illustrates composition of qualifiers, but it repeats qualifiers the schema already enumerates, so it adds only marginal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Search) and resource (commits) plus the axes of search (message, author, date) and the target system (GitHub). The second sentence clarifies the higher-order purpose — locating where a change was introduced — which no sibling tool covers, so it is trivially distinguishable from the rest of the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear use case: 'The way to find where a change was introduced,' which tells the agent when this tool is the right pick. It does not name an alternative or state exclusions, but no sibling tool overlaps with commit search, so the omission costs little.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_drafts_createC

Save an email draft to Gmail.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesThe draft to create.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/non-read nature is covered structurally. The description adds nothing beyond that: it does not say the draft is not sent, whether creation is idempotent, or what auth is required, so the behavioral burden is largely unmet.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is well structured. It is arguably under-specified rather than bloated, but as a size/structure judgment it is tight and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with only minimal annotations and no output schema, the description should at least clarify that it saves rather than sends and hint at the returned draft. The rich input schema compensates for parameters, but the core behavioral distinction is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the nested Message/Draft/EmailContent fields are richly documented (threadId rules, bodyHtml multipart behavior, in_reply_to threading). The description adds no parameter meaning at all, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Save an email draft') and destination (Gmail), which is enough to know it creates a draft rather than sending. However, it does not distinguish itself from the adjacent gmail_messages_send sibling, so an agent gets no explicit routing cue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the alternative tool (gmail_messages_send) for actually delivering mail. The agent must infer that this only persists a draft.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_messages_attachments_getA
Read-onlyIdempotent

Read one attachment, by the attachmentId messages_get or threads_get lists. A text file comes back as text; a binary one is reported, not decoded.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe ID of the attachment.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me
messageIdYesThe ID of the message containing the attachment.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld, so safety is covered. The description adds real behavioral value beyond them: text attachments return as text and binary ones are 'reported, not decoded', which sets return expectations for a tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the core action and the id provenance are front-loaded, and the return-format caveat follows.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no output schema, the description covers the important unknown (how content is returned). What 'reported, not decoded' actually looks like for a binary attachment remains slightly ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both required params and userId are documented in the schema; baseline is 3. The description clarifies where the attachmentId originates, but adds no format or constraint detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Read one attachment') and it explicitly names the sibling tools (messages_get, threads_get) that produce the attachmentId, so the agent can place it precisely among the gmail_* family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the context of use by tying the call to the id returned by messages_get/threads_get, which is the key prerequisite. It does not, however, state any exclusion or alternative for fetching attachment content another way.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_messages_listC
Read-onlyIdempotent

List messages in the user's mailbox.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoOnly return messages matching the specified query. Supports the same query format as the Gmail search box. For example, "from:someuser@example.com rfc822msgid:<somemsgid@example.com> is:unread". Parameter cannot be used when accessing the api using the gmail.metadata scope.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me
labelIdsNoOnly return messages with labels that match all of the specified label IDs. Messages in a thread might have labels that other messages in the same thread don't have.
pageTokenNoPage token to retrieve a specific page of results in the list.
maxResultsNoMaximum number of messages to return. This field defaults to 100. The maximum allowed value for this field is 500.
includeSpamTrashNoInclude messages from SPAM and TRASH in the results.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered structurally. The description adds nothing beyond that: it does not mention pagination, the default/max result caps, or how SPAM/TRASH are excluded by default, all of which materially affect behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence is front-loaded and waste-free, but its brevity here reflects under-specification rather than disciplined conciseness for a six-parameter list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full schema documentation and annotations covering safety, the essentials are present. It is incomplete in that it omits pagination/result-cap behavior and return shape, but with no output schema and a fully documented input schema, the omission is a moderate rather than severe gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every one of the six parameters has a detailed description in the schema, including query syntax, page tokens, and maxResults bounds. The description adds no parameter meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List messages') and the scope ('user's mailbox'), which is clear enough for an agent to know it retrieves Gmail messages. However, it offers no differentiation from the sibling gmail_threads_list, which is also a Gmail listing operation, so the agent must infer the distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description contains no when-to-use, when-not-to-use, or alternative-tool guidance. Nothing tells the agent to prefer this over gmail_threads_list or when a query-based search (q) is the right approach versus fetching everything.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_threads_getC
Read-onlyIdempotent

Read a Gmail thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe unique ID of the Gmail thread to retrieve.
formatNoThe format to return the messages in.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me
metadataHeadersNoWhen format is 'METADATA', only include these headers in the response.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds nothing beyond that—no note on auth requirements, what a 'thread' contains, or how format affects the response—so it does not enrich the behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is appropriately sized, though the brevity reflects under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool whose annotations cover safety and whose schema documents all four parameters, the description is minimally adequate. It would be stronger if it hinted at the format options or return structure, but nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with every parameter (id, format, userId, metadataHeaders) documented in the schema itself. The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('a Gmail thread'), which cleanly separates it from the list-oriented sibling gmail_threads_list. It does not, however, explicitly contrast itself with gmail_threads_list or gmail_messages_list, so sibling differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus gmail_threads_list or gmail_messages_list, and no prerequisites or context are given. The agent must infer usage purely from the name and the required 'id' parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_threads_listC
Read-onlyIdempotent

List Gmail threads.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoOnly return threads matching this Gmail search query string.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me
labelIdsNoReturn only threads with all of these label IDs.
pageTokenNoPage token to retrieve a specific page of results in the list.
maxResultsNoMaximum number of threads to return (default 100, max 500).
includeSpamTrashNoInclude threads from SPAM and TRASH in the results.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered by structured data. The description adds nothing beyond that — no note on pagination behavior, result ordering, or the fact that a full mailbox scan may be needed without filters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the key information front-loaded. It is efficient, though the extreme brevity shades into under-specification rather than pure conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter read tool with no output schema, the description is minimally viable: the schema covers all inputs, but the description omits pagination semantics and what a thread result contains. Nothing is misleading, but an agent gets no help beyond the structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (q, userId, labelIds, pageToken, maxResults, includeSpamTrash) is already documented in the schema. The description contributes no additional meaning, which is the baseline 3 case when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('Gmail threads'), which is unambiguous and distinguishable from write-oriented siblings like gmail_messages_send and gmail_drafts_create. It stops short of scope details (e.g. mailbox-wide vs. label-filtered), so it is clear but not maximally informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no mention of prerequisites or the conditions under which a caller should prefer it. The sibling set contains other Gmail operations but the description offers no routing signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_appendC

Appends values to a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation of a range to search for a logical table of data. Values are appended after the last row of the table.
valueRangeYesThe request body contains an instance of ValueRange.
spreadsheetIdYesThe ID of the spreadsheet to update.
insertDataOptionNoHow the input data should be inserted.
valueInputOptionYesHow the input data should be interpreted.
includeValuesInResponseNoDetermines if the update response should include the values of the cells that were appended. By default, responses do not include the updated values.
responseValueRenderOptionNoDetermines how values in the response should be rendered. The default render option is FORMATTED_VALUE.
responseDateTimeRenderOptionNoDetermines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false and openWorldHint=true, indicating a mutating, external operation. The description adds nothing beyond this – it doesn't disclose how appended data interacts with existing tables, whether headers are auto-detected, or rate-limit/permission requirements. For a mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single, front-loaded sentence that wastes no words. However, the extreme brevity contributes to gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, nested value types, mutation behavior), the description is critically underspecified. It omits key behavioral details like how the append range is determined, interaction with existing data, and available options (insertDataOption, valueInputOption). No output schema exists to compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds no parameter-level detail beyond what's already provided. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (appends) and resource (values to a spreadsheet), which is clear but does not differentiate from the sibling gsheets_spreadsheets_values_update or clarify the 'logical table' append behavior. It's clear but lacks sibling differentiation within the gsheets tool family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like gsheets_spreadsheets_values_update, and no mention of required preconditions or idempotency considerations. The description is silent on usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_getC
Read-onlyIdempotent

Returns a range of values from a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation or R1C1 notation of the range to retrieve values from.
spreadsheetIdYesThe ID of the spreadsheet to retrieve data from.
majorDimensionNoThe major dimension that results should use. For example, if the spreadsheet data in Sheet1 is: A1=1,B1=2,A2=3,B2=4, then requesting range=Sheet1!A1:B2?majorDimension=ROWS returns [[1,2],[3,4]], whereas requesting range=Sheet1!A1:B2?majorDimension=COLUMNS returns [[1,3],[2,4]].
valueRenderOptionNoHow values should be represented in the output. The default render option is FORMATTED_VALUE.
dateTimeRenderOptionNoHow dates, times, and durations should be represented in the output. This is ignored if valueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no mention of behavior for missing/empty ranges, error cases, or whether the range must already exist — so it contributes essentially no behavioral context of its own.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero padding. It is efficient, though arguably terse given the tool takes five parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover the safety profile and the schema fully documents inputs, so the description's burden is lighter. Still, with no output schema, it says only that 'values' are returned without hinting at the row/column shape that majorDimension controls, which is the main thing an agent must reason about.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameters (range, spreadsheetId, majorDimension, valueRenderOption, dateTimeRenderOption) are already richly documented in the schema, including an example for majorDimension. The description adds no parameter meaning at all, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a range of values from a spreadsheet'), which is clearly readable. However, it does not differentiate this from siblings such as gsheets_spreadsheets_values_update or gdocs_documents_get, leaving the agent to infer the read-vs-write distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of when NOT to use it, and no reference to the sibling update tool. The agent gets an implied read-only purpose from the verb but nothing explicit about context or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_updateC

Sets values in a range of a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation of the values to update.
valueRangeYesThe request body contains an instance of ValueRange.
spreadsheetIdYesThe ID of the spreadsheet to update.
valueInputOptionYesHow the input data should be interpreted.
includeValuesInResponseNoDetermines if the update response should include the values of the cells that were updated. By default, responses do not include the updated values. If the range to write was larger than the range actually written, the response includes all values in the requested range (excluding trailing empty rows and columns).
responseValueRenderOptionNoDetermines how values in the response should be rendered. The default render option is FORMATTED_VALUE.
responseDateTimeRenderOptionNoDetermines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the write nature is known. But the description adds nothing beyond the name – it does not disclose that existing cell values are overwritten, how valueInputOption affects interpretation, or what the update returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is tight. But for a 7-parameter mutation tool it is under-specified rather than appropriately sized; conciseness here is closer to omission.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with 7 params and no output schema, the description omits overwrite semantics, auth requirements, and response behavior. An agent could call it, but not safely without reading the schema closely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with 7 well-documented parameters, so the schema carries the meaning. The description adds no parameter detail beyond 'in a range', which is baseline 3 for fully documented schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (sets values in a range of a spreadsheet), which an agent can distinguish from gsheets_spreadsheets_values_get by direction of data flow. However it offers no explicit sibling differentiation and largely restates the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the sibling gsheets_spreadsheets_values_get or when reading vs writing applies. The agent must infer context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

linear_projects_listB
Read-onlyIdempotent

List projects in the workspace, with their status and progress.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesNoPaging and filtering.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, so safety profile is explicit. The description adds the workspace scoping and return fields (status, progress), which is useful context beyond annotations. But it omits pagination behavior despite a `first`/`after` paging contract, and doesn't note default result size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single short sentence, front-loaded verb and resource. No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with a rich filter/paging schema and no output schema, the description is too thin: it doesn't mention paging, default page size, filtering capability, or archived handling, all of which an agent needs to call it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents paging, filtering, ordering, and archived-inclusion. The description adds nothing about variables, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: 'List projects in the workspace' with the added scope of status and progress fields. Distinguishable from siblings (none are project-listing), though it doesn't explicitly differentiate itself from other list tools in the sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives, no prerequisites, no mention of paging defaults. The description just states what it does.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notion_pages_createA

Create a page — as a subpage of another page, or as a row of a database by giving its data_source_id as the parent. Content comes as a markdown string Notion parses into blocks, or from a template: one or the other, never both.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesThe page to create.
filterPropertiesNoProperty IDs to return on the page that comes back, instead of all of them. A page that does not have a listed property omits it.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, establishing this as an external write. The description adds the meaningful markdown/template mutual exclusion. It does not disclose auth requirements, rate limits, or the allowAsync async-202 behavior, but with annotations carrying the safety profile a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the verb and the two modes, with the exclusivity constraint phrased crisply. No wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex creation tool with a fully documented schema, the description covers parent selection and content sourcing adequately. It omits the async task path (allowAsync → 202) and return shape, but with no output schema and 100% schema coverage these are minor gaps, and annotations cover the write semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value: it clarifies that `parent` takes `data_source_id` to make a database row and that template vs. markdown are mutually exclusive — beyond the schema's raw field docs. It doesn't add detail on `properties` or `filterProperties`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a page') and immediately delineates the two creation modes — subpage vs. database row — which is exactly what separates it from siblings like notion_pages_update. An agent can identify the tool's scope without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance: use a page parent for a subpage, `data_source_id` for a database row, and content comes from `markdown` OR a template, 'never both.' The mutual-exclusion rule is explicit. However, it names no sibling alternatives (e.g., update vs. create routing) and states no prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_order_getA
Read-onlyIdempotent

Get one order with its line items, shipping address and totals. Takes a global ID (gid://shopify/Order/...), which orders_list returns.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesWhich order to fetch.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint and idempotentHint, so the safety profile is covered externally. The description adds useful shape-of-result information (line items, address, totals) and constrains the ID format, but says nothing about auth requirements, rate limits, or error behavior. With annotations carrying the safety burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the purpose is front-loaded and the ID-format detail is deferred to the second sentence where it belongs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with full annotation coverage, the description does the work of standing in for an absent output schema by naming what gets returned. Nothing essential is missing, though error/not-found handling is unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'id' parameter is already fully documented in the schema with the same gid:// example. The description's mention of the global ID format duplicates rather than extends the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Get') plus a precisely scoped resource ('one order') and it enumerates the payload contents (line items, shipping address, totals). An agent can tell this is a single-record fetch, distinct from any list/collection operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies concrete workflow context by noting the required global ID is the one 'orders_list' returns, which tells the agent when this tool fits (after a listing). It stops short of stating exclusions or what to do if the ID is unknown, so it is not a full when/when-not treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

slack_chat_post_messageA

Send a message to a Slack channel, private group, or DM. Provide text for a plain message; set thread_ts to reply inside an existing thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoThe main body text of the message. Required unless blocks or attachments are provided. Used as the fallback string for notifications when blocks are provided, so it is worth setting even then.
parseNoChange how messages are treated. Accepts 'none' or 'full'.
blocksNoA JSON-based array of structured Block Kit blocks.
mrkdwnNoDisable Slack markup parsing by setting to false. Defaults to true.
channelYesAn encoded ID or channel name that represents a channel, private group, or IM channel to send the message to. Prefer the encoded ID (e.g. 'C123ABC456').
iconUrlNoURL to an image to use as the icon for this message. Requires the chat:write.customize scope.
metadataNoApplication-specific metadata to attach to the message.
threadTsNoProvide another message's 'ts' value to make this message a reply in that thread. Avoid using a reply's ts value; use the parent's.
usernameNoSet the bot's user name. Requires the chat:write.customize scope.
iconEmojiNoEmoji to use as the icon for this message, e.g. ':chart_with_upwards_trend:'. Requires the chat:write.customize scope.
linkNamesNoFind and link user groups.
attachmentsNoA JSON-based array of structured attachments.
unfurlLinksNoPass true to enable unfurling of primarily text-based content.
unfurlMediaNoPass false to disable unfurling of media content.
markdownTextNoAccepts message text formatted in markdown. Limit this field to 12,000 characters. Cannot be used together with blocks or text.
replyBroadcastNoUsed in conjunction with thread_ts and indicates whether the reply should be made visible to everyone in the channel. Defaults to false.
unfurlAppLinksNoPass true to enable unfurling of links to installed apps.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external nature is covered. The description adds the threading behavior, but does not disclose required scopes, rate limits, message-size limits, or what a successful send returns — and most scope info already lives in the schema parameter descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and then the two most important parameter behaviors. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with no output schema, the description is thin but the schema carries full parameter documentation, so an agent can call it correctly. Missing behavioral context (rate limits, required scopes, response shape) keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 17 parameters are documented in the schema itself. The description's notes on `text` and `thread_ts` largely restate what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Send') and resource ('a message to a Slack channel, private group, or DM'), making the action and destination unambiguous. An agent can distinguish this from siblings like slack_conversations_create without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives implied usage for `text` and `thread_ts`, but there is no explicit when-to-use vs. when-not, no mention of prerequisites (e.g. chat:write scope), and no reference to alternative messaging tools such as gmail_messages_send. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

slack_users_listA
Read-onlyIdempotent

List members of the workspace. Use this to resolve a person's name to the user ID that Slack mentions and filters expect.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoThe maximum number of items to return. Fewer than the requested number of items may be returned, even if the end of the users list has not been reached. Providing no limit value will result in Slack attempting to deliver you the entire result set.
cursorNoPaginate through collections of data by setting this to the next_cursor attribute returned by a previous request's response_metadata.
teamIdNoEncoded team id to list users in. Required if the token belongs to an org-wide app.
includeLocaleNoSet this to true to receive the locale for users. Defaults to false.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, and the description adds nothing behavioral on top of them. It says nothing about pagination behavior, the teamId requirement for org-wide tokens, or rate limits, so the structured fields carry the whole load.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose first and the resolution use case second, with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, zero-required list tool with no output schema, the description covers what it returns only implicitly ('user ID'). It omits pagination guidance and the org-wide token/teamId caveat, leaving the agent to discover those from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, cursor, teamId and includeLocale are already fully documented in the schema. The description adds no parameter-level meaning beyond that, which makes the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List members of the workspace') and adds the resolution intent, so an agent can distinguish it from slack_chat_post_message. It does not name a sibling alternative, but the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use case: resolving a person's name to the user ID that mentions and filters expect. There is no explicit when-not-to-use or named alternative, but the triggering context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_accounts_listA
Read-onlyIdempotent

List the connected accounts on this platform. Empty if you are not a platform.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
createdNoOnly return records created in this window.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds the notable behavior that the result is empty for non-platform accounts, which is useful context beyond the annotations, but says nothing about pagination or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no waste, with the core action stated first and the non-platform caveat immediately after. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with fully documented parameters and safety annotations, the description is essentially complete; the only minor gap is pagination/return behavior, which the schema largely covers via the cursor parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with limit, created window, and both pagination cursors documented directly in the schema. The description adds no parameter-level detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('connected accounts on this platform'), making it distinguishable from siblings like stripe_customers_list or stripe_payouts_list. It is clear but does not explicitly differentiate itself from those other Stripe list tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'Empty if you are not a platform' gives useful conditional context about when this tool is relevant/applicable, which is a genuine usage signal. However, there is no explicit guidance on when to prefer this over sibling list tools or how to page through results.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_application_fees_listC
Read-onlyIdempotent

List the platform fees collected from connected accounts.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
chargeNoOnly return application fees for this charge ID.
createdNoOnly return records created in this window.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint and idempotentHint, so the safety profile is covered structurally. The description adds nothing beyond that: no pagination behavior, no note on which filters combine, no indication of the response shape. It contributes essentially zero behavioral context on top of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is appropriately sized, though it does so little work that brevity here reflects minimalism rather than disciplined structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list with fully documented parameters and no output schema, the definition is just barely sufficient to call the tool correctly. It omits pagination expectations and scope notes about connected accounts, which would have been cheap and useful additions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (limit, charge, created window, endingBefore, startingAfter) is already fully documented in the schema. The description adds no parameter meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('platform fees collected from connected accounts'), which is distinct from the many other Stripe list tools among the siblings. However, it offers no explicit differentiation from tools like stripe_balance_transactions_list or stripe_invoices_list, leaving the agent to infer the distinction from the resource name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no prerequisites, and no mention of the sibling list tools it could be confused with. The description gives a purpose but no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_balance_retrieveB
Read-onlyIdempotent

Retrieve the current account balance.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so safety behavior is covered. The description adds no behavioral context beyond what annotations provide, such as authentication needs, rate limits, or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is appropriately sized for a zero-parameter read-only tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool with rich annotations, the description is sufficient to select and invoke it. However, with no output schema, it does not explain what the balance response contains (e.g., available vs. pending), leaving a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema is empty and the baseline for parameter semantics is 4. There are no parameter details to add.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Retrieve') and resource ('current account balance'), making it distinct from siblings like stripe_balance_transactions_list and stripe_payouts_list. It does not explicitly name or contrast a sibling, so it stops short of the highest clarity tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. Usage is only implied by the tool name and resource.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_balance_transactions_listA
Read-onlyIdempotent

List every movement across the Stripe balance: charges, refunds, fees and payouts. For accounting, reporting_category on each result groups them better than type does.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoOnly return transactions of this type.
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
payoutNoOnly return transactions paid out in this payout. Automatic payouts only.
sourceNoOnly return transactions for this object ID.
createdNoOnly return records created in this window.
currencyNoThree-letter lowercase ISO currency code.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds the useful detail that each result carries a reporting_category field, but says nothing about pagination behavior, volume limits, or return shape beyond that. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the scope statement is front-loaded ahead of the accounting tip. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, all-optional list endpoint with no output schema, the description gives the resource scope and a key field-level hint. It stops short of mentioning pagination or the shape of a result, which an agent would have to infer from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema carries the parameter burden and the baseline is 3. The description earns a point above baseline by explaining the relationship between the `type` filter and the `reporting_category` output field, telling the agent which is more useful for accounting grouping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (balance transactions) and enumerates what those movements are: charges, refunds, fees and payouts. This scope is precise enough to separate it from stripe_balance_retrieve (totals) and from the narrower list endpoints like stripe_payouts_list or stripe_charges_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The accounting framing ('for accounting, reporting_category groups them better than type') implies the reconciliation use case, but there is no explicit when-to-use or when-not-to-use guidance and no routing to alternative list endpoints such as stripe_payouts_list or stripe_charges_list. Usage is inferable, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_charges_listB
Read-onlyIdempotent

List charges, most recently created first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
customerNoOnly return charges for the customer with this ID.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
paymentIntentNoOnly return charges for this PaymentIntent.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds one genuinely new fact beyond structured data: default sort order is most-recently-created first. It says nothing about pagination behavior, page size limits, or how cursors interact, though those live in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero filler, and the sort-order fact is front-loaded rather than buried. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list endpoint with fully documented params, no required fields, annotations covering the safety profile, and no output schema, the description covers the essentials. A brief note on pagination or result shape would close the remaining gap, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (limit, customer, paymentIntent, startingAfter, endingBefore) are already documented in the schema, including the mutual exclusivity of the cursors. The description adds no parameter-level meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (charges), and the resource name cleanly separates it from the other stripe_*_list siblings like invoices, payouts, and disputes. It doesn't explicitly contrast with any alternative, but the resource noun does the disambiguation work.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of filtering alternatives, and no indication of when this is preferable to a more targeted retrieval. Usage is only implied by the tool name itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_customers_listA
Read-onlyIdempotent

List customers, most recently created first. Filter by email to find one.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNoA case-sensitive filter on the list based on the customer's email field. The value must be a string.
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered without description help. The description adds the non-obvious ordering guarantee (most recently created first), which is genuinely useful, but says nothing about rate limits, page size defaults, or result shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the ordering behavior is front-loaded before the filtering hint. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no output schema, the description plus the rich parameter schema and annotations give an agent enough to call it correctly. The main omission is any note about pagination workflow or default result count, though the schema covers the cursor mechanics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (email, limit, endingBefore, startingAfter) are already fully documented in the schema. The description echoes the email filter without adding syntax, case-sensitivity, or pagination semantics beyond what the schema states, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (customers) plus a default sort order, so the agent knows exactly what it retrieves. It does not explicitly distinguish itself from the sibling stripe_checkout_sessions_list, but the resource noun is unambiguous enough to route correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Filter by email to find one' implies the lookup use case, giving some usage context. However, there is no guidance on when to use this versus other Stripe list endpoints, and no mention of pagination workflow for iterating beyond the default limit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_customers_retrieveB
Read-onlyIdempotent

Retrieve a single customer by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
customerYesThe identifier of the customer, e.g. 'cus_NffrFeUfNV2Hib'.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, covering the safety profile. The description adds no behavioral context beyond what the annotations and the verb 'Retrieve' already imply, such as authentication needs, rate limits, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for a simple retrieval tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple read operation, full schema coverage, and annotations that cover safety traits, the description is almost complete. It does not describe the returned customer object, which is a minor gap in the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is fully documented with an example. The description adds only 'by ID', which is already conveyed by the schema's parameter name and description, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Retrieve'), resource ('customer'), and scope ('single ... by ID'). It clearly distinguishes retrieval from creation, but does not explicitly name the sibling tool (stripe_customers_create) or otherwise differentiate itself from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description implies a straightforward lookup, but gives no context about when it is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_disputes_listA
Read-onlyIdempotent

List disputes, most recently created first. Stripe has no status filter here, so check status on the results.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
chargeNoOnly return disputes for this charge ID.
createdNoOnly return disputes created in this window.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
paymentIntentNoOnly return disputes for this PaymentIntent ID.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so safety is covered. The description adds genuine behavioral context beyond that: the default ordering and the important limitation that status filtering is unavailable server-side.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler; the ordering fact and the status-filter caveat are both front-loaded and each earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full schema coverage and no output schema, the description covers what the schema cannot: ordering and the missing status filter. Pagination behavior is only in the schema, but that is acceptable given the rich schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are documented in the schema itself (including the seconds-not-milliseconds timestamps and pagination cursors). The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List disputes') plus the ordering ('most recently created first'), which distinguishes it from stripe_disputes_retrieve and stripe_disputes_update. It does not name a sibling explicitly, but the list semantics are clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It warns that there is no status filter and that the caller must check `status` on results, which is a useful usage note. However, it gives no guidance on when to use this versus stripe_disputes_retrieve or when to apply the charge/paymentIntent/created filters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_disputes_retrieveA
Read-onlyIdempotent

Retrieve a dispute by ID. evidence_details.due_by is the deadline; evidence submitted after it is wasted.

ParametersJSON Schema
NameRequiredDescriptionDefault
disputeYesThe ID of the dispute.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint), so the bar is lower. The description adds genuinely useful domain context beyond the annotations: that evidence_details.due_by is the submission deadline and late evidence is wasted — a behavioral/operational fact not present in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, the core operation front-loaded, and the second sentence delivers a high-value domain warning with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-param read tool with full annotation coverage and no output schema, the description is essentially sufficient, and the due_by warning adds helpful context. It stops short of 5 only because it says nothing about the returned dispute structure or error behavior for an invalid ID.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema fully documents the dispute ID. The description adds nothing about the parameter itself, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (dispute) with the scoping qualifier 'by ID', which implicitly separates it from stripe_disputes_list. It does not explicitly name or contrast with the sibling retrieve/update tools, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by ID' implies the precondition (you must already have a dispute ID), which is the natural selector versus stripe_disputes_list. However, there is no explicit when-to-use statement or mention of alternatives such as the list tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_disputes_updateA

Add evidence to a dispute. Writing any evidence field submits every evidence field to the bank unless submit is false, and submission cannot be undone. Set submit to false to stage it for review first.

ParametersJSON Schema
NameRequiredDescriptionDefault
submitNoWhether to send the evidence to the bank now. False stages it on the dispute instead, where it can be read and changed. Stripe submits if this is not set, and submitting cannot be undone.
disputeYesThe ID of the dispute.
evidenceNoEvidence supporting the charge.
metadataNoSet of key-value pairs attached to the dispute.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses critical behavior beyond the annotations: writing any evidence field submits every evidence field to the bank, submission is irreversible, and submit=false stages for review. These are exactly the side-effect and reversibility details an agent needs for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and immediately followed by the critical side-effect warning. There is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with irreversible side effects and no output schema, the description provides the essential safety context: default submission, staging via submit=false, and the all-fields submission behavior. The rich schema covers the evidence fields themselves.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents parameters well. The description still adds important semantics: evidence fields are submitted collectively and submit controls staging versus immediate irreversible submission, which is not fully captured by the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: adding evidence to a dispute. It is clear enough for an agent to distinguish from stripe_disputes_retrieve/list by intent, though it does not explicitly name those siblings or explain that this tool updates rather than reads disputes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditional guidance: set submit to false to stage evidence for review, otherwise submission happens automatically. It lacks explicit alternatives among sibling tools, such as using stripe_disputes_retrieve to review evidence before calling this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_invoice_items_createA

Add a line to an invoice. Name an existing price as pricing.price, or give an amount directly. Without invoice the line waits for the customer's next subscription invoice.

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNoThe amount in the smallest currency unit. Negative reduces what the invoice is due. Cents, not dollars: $15.00 is 1500, and 15 puts fifteen cents on the invoice. Multiply a decimal amount by 100.
periodNoThe service period this line covers.
invoiceNoThe draft invoice to add this line to, at most 250 lines. Absent, the line waits for the customer's next subscription invoice and a standalone invoice will not collect it on its own.
pricingNoAn existing price to bill, named as `pricing.price`.
taxCodeNoThe tax code for what is being billed.
currencyNoThree-letter lowercase ISO currency code.
customerYesThe ID of the customer to bill.
metadataNoSet of key-value pairs attached to the line.
quantityNoHow many units this line bills.
discountsNoCoupons or promotion codes applying to this line alone.
priceDataNoA price created inline for this line.
descriptionNoWhat this line says on the invoice.
taxBehaviorNoWhether the amount includes tax. Once set to inclusive or exclusive it cannot be changed.
discountableNoWhether invoice-level discounts apply to this line.
subscriptionNoBill this line on that subscription's invoices only.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations disclose mutation (readOnlyHint=false) and open-world reach, so the safety profile is covered. The description usefully adds lifecycle behavior — that a line without `invoice` defers to the customer's next subscription invoice — which is non-obvious and not captured by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero filler, with the core action front-loaded and the deferred-invoice caveat last. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter mutation tool with full schema coverage, annotations for the mutation profile, and no output schema, the description covers purpose, the two pricing paths, and the missing-invoice behavior. The required `customer` parameter is left to the schema, a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents all 15 parameters in depth (units, formats, limits). The description's `pricing.price` vs `amount` note restates what the schema says, adding little beyond it, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Add a line to an invoice" gives a specific verb and resource that clearly maps to creating a Stripe invoice item, and the pricing/amount guidance sharpens it. It does not explicitly name or rule out sibling tools such as stripe_invoices_create, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states two concrete ways to supply the line's value (name an existing price via `pricing.price`, or pass `amount` directly) and explains the conditional behavior when `invoice` is omitted. It gives clear operating context but no explicit when-to-use-this-vs-an-alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_invoices_createA

Create a draft invoice for a customer. It bills nothing until finalised, so this is the safe half of invoicing.

ParametersJSON Schema
NameRequiredDescriptionDefault
footerNoFooter text displayed on the invoice.
dueDateNoUnix timestamp when payment is due. Valid only when `collection_method` is 'send_invoice'. Seconds, not milliseconds: 1700000000, not 1700000000000.
currencyNoThree-letter lowercase ISO currency code. Absent, the customer's currency.
customerYesThe ID of the customer to bill.
metadataNoSet of key-value pairs attached to the invoice.
discountsNoCoupons and promotion codes to apply. Absent, the invoice inherits the customer's discount.
onBehalfOfNoThe connected account the funds are intended for. Its branding and support information appear on the invoice.
autoAdvanceNoWhether Stripe collects the invoice automatically. False leaves the invoice where it is until you act on it.
descriptionNoAn arbitrary string attached to the invoice. Shown as the memo in the Dashboard.
daysUntilDueNoDays until the invoice is due. Valid only when `collection_method` is 'send_invoice'.
subscriptionNoBill this subscription. The invoice then includes only that subscription's pending items, and its billing cycle is untouched.
collectionMethodNoHow to collect payment. Absent, Stripe charges the customer's default payment method.
statementDescriptorNoWhat the customer sees on their card statement. Must contain at least one letter.
defaultPaymentMethodNoThe payment method to charge. It must belong to the invoice's customer.
pendingInvoiceItemsBehaviorNoWhether to pull the customer's pending invoice items onto this invoice. Absent, Stripe excludes them and the draft is empty.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag a mutation (readOnlyHint false) and open-world scope; the description adds genuinely new behavioral context: the invoice does not bill until finalised, so creation itself is side-effect-free with respect to payment. It does not cover auth/permission needs or that a draft is empty without line items.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the key behavioral reassurance. Nothing is wasted and nothing important is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter mutation with no output schema, the description is thin: it omits the workflow detail that the draft is empty unless invoice items are added (via stripe_invoice_items_create) and says nothing about finalization as the follow-up step. The rich schema compensates on parameters, but the end-to-end workflow is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 15 parameters are already documented in the schema. The description adds no additional parameter meaning (e.g., defaults, interactions), which matches the baseline of 3 when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create a draft invoice for a customer.' The phrase 'the safe half of invoicing' nicely frames it as distinct from a finalization step, though it never names the sibling (stripe_invoices_finalize) explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when this is appropriate ('it bills nothing until finalised') which nudges the agent toward creating a draft before finalizing, but it never states prerequisites, when-not-to-use, or the alternative tool by name. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_invoices_listA
Read-onlyIdempotent

List invoices, most recently created first. Filter by customer, subscription, status or collection method.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
statusNoOnly return invoices with this status.
createdNoOnly return invoices created in this window.
customerNoOnly return invoices for this customer ID.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
subscriptionNoOnly return invoices for this subscription ID.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.
collectionMethodNoOnly return invoices collected this way.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint/idempotentHint/openWorldHint annotations already declare that this is a safe, non-mutating, idempotent read against external data, so the behavioral bar is lowered. The description adds one real behavioral fact not covered by annotations, the default sort order, but says nothing about default page size or pagination behavior beyond what the schema already carries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler, and the ordering behavior is front-loaded immediately after the core verb-resource statement. Every clause carries information an agent would otherwise have to infer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an eight-parameter, zero-required list tool with 100% schema coverage and annotations that cover the safety profile, the description supplies the two things the schema cannot: ordering and the set of filterable dimensions. Pagination semantics are left to the schema, which is acceptable given the rich cursor descriptions, though a brief note would have closed the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all eight parameters including timestamps, cursors, and enums, making 3 the baseline. The description names four of the filter dimensions but adds no semantics beyond the schema and omits the created-window and pagination parameters entirely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (invoices) plus the sort order (most recently created first), which lets an agent distinguish it from the mutation siblings stripe_invoices_create and stripe_invoices_send without opening a schema. It does not explicitly name those alternatives, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence enumerates the available filter dimensions (customer, subscription, status, collection method), which implies when to use the tool for narrowing a result set. It offers no explicit when-not guidance, no mention of pagination as a usage consideration, and no routing to sibling list tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_invoices_sendA

Email an invoice to the customer outside the normal schedule. Test mode sends nothing but still emits the event.

ParametersJSON Schema
NameRequiredDescriptionDefault
invoiceYesThe ID of the invoice.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries weight by disclosing a real behavioral trait: 'Test mode sends nothing but still emits the event.' That tells the agent about a side effect that survives even when no email goes out, which is not derivable from the annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action front-loaded ahead of the test-mode caveat. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation tool with no output schema, the description covers the action and a key behavioral caveat. It leaves out error conditions (e.g., sending an already-paid or draft invoice) and required permissions, but the annotations cover the safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one parameter and the schema already documents it at 100% coverage ('The ID of the invoice.'). The description adds no format, ID prefix, or constraint detail beyond the schema, so the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Email an invoice to the customer.' An agent can distinguish this from stripe_invoices_create and stripe_invoices_list without opening any schema. It stops short of naming a sibling alternative, so it lands at a clear-but-undifferentiated 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Outside the normal schedule' implies when this tool is preferred over the automatic sending behavior, which is useful implied context. However, there is no explicit when-not guidance, no mention of prerequisites like invoice finalization, and no named alternative among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_payouts_listB
Read-onlyIdempotent

List payouts to your own bank account or card, most recently created first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
statusNoOnly return payouts with this status. A payout Stripe has sent to the bank but not settled reports 'in_transit' on the object, which this filter does not document.
createdNoOnly return records created in this window.
arrivalDateNoOnly return payouts expected to arrive in this window.
destinationNoOnly return payouts sent to this external account: a bank account or card, not a connected account.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds only the default sort order ('most recently created first'); it says nothing about pagination mechanics or result size beyond what the schema provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. The resource is named immediately and the scoping/ordering detail follows compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full schema coverage and annotations covering safety, the definition is nearly sufficient. The main gap is the absence of guidance on pagination workflow and how this list differs from related payout/balance/transaction lists, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters (limit, status, created, arrivalDate, destination, cursors) are already documented in the schema with good detail. The description adds no parameter-level meaning beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (payouts), plus a distinguishing scope: 'to your own bank account or card' (as opposed to connected accounts). It also communicates default ordering, but does not name a sibling tool to contrast against, leaving differentiation to be inferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use guidance and never mentions alternatives such as stripe_balance_transactions_list or stripe_disputes_list. Usage is only implied by the resource name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_prices_createB

Create a price for a product. Include recurring for a subscription price, omit it for a one-off.

ParametersJSON Schema
NameRequiredDescriptionDefault
activeNoWhether the price can be used. Absent, it can.
productYesThe ID of the product this price is for.
currencyYesThree-letter lowercase ISO currency code.
metadataNoSet of key-value pairs attached to the price.
nicknameNoAn internal label. Customers never see it.
recurringNoBilling interval. Absent, this is a one-off price.
unitAmountNoThe amount in the smallest currency unit. Mutually exclusive with unit_amount_decimal. Cents, not dollars: $15.00 is 1500, and 15 prices this at fifteen cents. Multiply a decimal amount by 100.
taxBehaviorNoWhether the amount includes tax. Once set to inclusive or exclusive it cannot be changed.
unitAmountDecimalNoSame as `unit_amount`, but accepts a decimal value in the smallest currency unit with at most 12 decimal places. Cents, not dollars: $15.00 is "1500". The decimal places are fractions of a cent, so "1500.5" is fifteen dollars and half a cent.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the agent knows this is a write to an external system. Beyond that the description discloses nothing behavioral: no note that price objects are immutable, no idempotency or retry guidance, no indication of what side effects or errors to expect. For a financial mutation tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero padding, and the core action is front-loaded before the conditional parameter advice. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The 9-parameter schema is fully documented and the description covers the key create-vs-subscription distinction, but there is no output schema and the description says nothing about the returned price object or the immutability of amounts after creation. Adequate for the mechanics, thin on consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including unitAmount, unitAmountDecimal, taxBehavior and the nested recurring object is already documented in detail. The description's recurring/one-off clarification is essentially a restatement of the schema's own 'Absent, this is a one-off price' note, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a price for a product'), which is enough to distinguish it from unrelated siblings like stripe_invoices_create. It does not explicitly distinguish itself from stripe_products_create, which an agent could plausibly confuse it with when setting up a new product.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives actionable conditional logic for the recurring parameter (include for subscription, omit for one-off), which is real usage guidance. However it says nothing about prerequisites (the product must already exist) or when to reach for stripe_products_create first, so guidance is implied rather than complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_products_createA

Create a product. Prices attach to it separately.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoA public webpage for the product.
nameYesThe product's name, shown to customers.
activeNoWhether the product can be bought. Absent, it can.
imagesNoUp to eight image URLs.
metadataNoSet of key-value pairs attached to the product.
descriptionNoCustomer-facing description of the product.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external-effect profile is covered. The description adds the relational fact that prices are attached by a separate call, but says nothing about idempotency, what identifier is returned for chaining, or whether the product is immediately active — modest added context against a lowered bar.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action, and the second sentence earns its place by preventing a mistaken assumption that pricing happens here. Zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description is thin: it doesn't say what the call returns or how to obtain the product id needed for the follow-up price creation it references. The required-parameter constraint is covered by the schema, so the gap is mainly the create-then-price workflow linkage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented in the schema (name, url, active, images, metadata, description). The description adds no parameter-level meaning beyond that, which is the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create a product" is a specific verb+resource, and the second sentence distinguishes it from the sibling stripe_prices_create by noting prices attach separately. It stops short of describing what kind of product (Stripe billing object) or its scope, but an agent can identify the operation unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note that prices attach separately implicitly routes the agent to stripe_prices_create for pricing, which is genuine usage guidance. However, it never states when to reach for this tool versus other create tools, nor any prerequisites (e.g., that a product must exist before a price can reference it) beyond that hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_refunds_createA

Refund a charge. Provide either charge or payment_intent. Omit amount to refund the full sum.

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNoA positive integer in the smallest currency unit representing how much to refund. Defaults to the entire charge. Cents, not dollars: $15.00 is 1500, and 15 refunds fifteen cents. Multiply a decimal amount by 100.
chargeNoThe identifier of the charge to refund.
reasonNoThe reason for the refund. If set to 'fraudulent', the associated payment is marked as fraudulent.
metadataNoSet of key-value pairs attached to the refund.
paymentIntentNoThe identifier of the PaymentIntent to refund.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries most of the behavioral burden. It usefully discloses the default-to-full-amount behavior and the mutual exclusivity of charge/payment_intent, but says nothing about irreversibility, idempotency, or how partial vs full refunds affect the underlying charge.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the two operative constraints. Zero filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with thin annotations (no destructiveHint) and no output schema, the description covers the key parameter constraints but omits meaningful behavioral context: that refunds are irreversible state changes, whether they can be repeated safely, and what happens to the parent charge. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds a real constraint the schema does not express: charge and payment_intent are alternatives (provide either one), which prevents an agent from populating both. The amount default is restated from the schema, so the gain is modest.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Refund a charge') that no sibling tool duplicates, so an agent can identify it immediately. It doesn't explicitly contrast with a sibling, but none of the listed siblings is a competing refund operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives invocation guidance ('Provide either charge or payment_intent', 'Omit amount to refund the full sum') rather than tool-selection guidance. It never says when this tool is preferred over alternatives or what prerequisites apply (e.g. charge must be uncaptured/refundable).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_subscriptions_listB
Read-onlyIdempotent

List subscriptions. Filter by customer or status.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
priceNoFilter for subscriptions that contain this recurring price ID.
statusNoThe status of the subscriptions to retrieve. Pass 'all' to return subscriptions of all statuses.
customerNoThe ID of the customer whose subscriptions will be retrieved.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds no behavioral context beyond that — nothing about the default status set returned, pagination cursor semantics, or rate/limit behavior — so it earns little credit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action front-loaded before the filtering hint. Nothing needs trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter list tool with no output schema and no required params, the description is minimal but workable since the schema carries full parameter documentation. It omits any note on pagination, default page size, or what a subscription object contains, which an agent would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including limit, price, and both cursors is already documented in the schema. The description only echoes customer and status, adding no syntax or format detail beyond the structured fields; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('subscriptions'), which cleanly separates it from siblings like stripe_prices_list and stripe_customers_retrieve. It does not, however, explicitly contrast itself with any sibling or state the scope of what is listed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence implies the two main filtering paths (customer, status), giving implied usage context, but there is no explicit when-to-use guidance, no mention of alternatives for narrower queries, and no note on default result behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_subscriptions_updateA

Update a subscription. Parameters not provided are left unchanged. To change an item send its id; omitting the id adds a new item instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsNoThe items to change, up to 20. Absent, the items are left alone.
cancelAtNoWhen to cancel: a Unix timestamp, or one of 'max_billed_until', 'max_period_end', 'min_period_end'. A timestamp before the current period ends prorates if prorations are enabled. Pass an empty string to clear a cancellation that was scheduled earlier; `cancel_at_period_end` does not clear a timestamp. A timestamp is seconds, not milliseconds: 1700000000, not 1700000000000.
metadataNoSet of key-value pairs attached to the subscription.
trialEndNoWhen the trial ends: the string 'now' to end it immediately, or a Unix timestamp at most two years out. Overrides any trial the plan defines. A timestamp is seconds, not milliseconds: 1700000000, not 1700000000000.
discountsNoCoupons and promotion codes. A populated array replaces the subscription's existing discounts. Absent, they are left alone.
descriptionNoAn arbitrary string attached to the subscription.
daysUntilDueNoDays until an invoice is due. Valid only when collection_method is 'send_invoice'.
subscriptionYesThe ID of the subscription to update.
prorationDateNoA Unix timestamp to calculate prorations against, as though the update happened then. Absent, Stripe uses now. A timestamp is seconds, not milliseconds: 1700000000, not 1700000000000.
paymentBehaviorNoWhat to do if payment for the update fails. Absent, Stripe allows the subscription to go past_due.
collectionMethodNoHow to collect payment. Absent, Stripe charges automatically.
cancelAtPeriodEndNoWhether the subscription cancels at the end of the current period. Send it only to change the setting.
prorationBehaviorNoHow to handle prorations when the billing cycle or an item quantity changes. Absent, Stripe creates them.
cancellationDetailsNoWhy the subscription is being cancelled.
defaultPaymentMethodNoThe payment method to charge. It must already belong to this subscription's customer.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag this as a mutation (readOnlyHint=false) with open-world effects. The description adds real value beyond that: the PATCH-style semantics that unprovided parameters are left unchanged, and the add-vs-update branching for items. It does not mention auth or rate limits, but the partial-update contract is the key behavioral trait an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: purpose first, then the partial-update contract, then the one non-obvious branching rule. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter mutation with full schema coverage and annotations declaring the write semantics, the description surfaces the two traps an agent could hit (silent partial update, item add-vs-update). No output schema exists, so no return-value explanation is needed. Coverage is strong, though it omits any pointer to related tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 15 parameters thoroughly, including the items.id add/update semantics. The description restates the items id rule rather than adding meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Update a subscription'), which is unambiguous and separable from read-oriented siblings like stripe_subscriptions_list. It stops short of naming any sibling or scope constraint, so it is clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives useful conditional behavior for the items parameter (send `id` to modify, omit it to add), which is genuine usage guidance. However it never states when to reach for this tool versus stripe_subscriptions_list or a create path, so context is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 39 tool updatesv0.1.0
    • First observedconnect
    • First observedconnection_status
    • First observedgcalendar_events_list
    • First observedgdocs_documents_create
    • First observedgdrive_files_create
    • First observedgithub_search_commits
    • First observedgmail_drafts_create
    • First observedgmail_messages_attachments_get
    • First observedgmail_messages_list
    • First observedgmail_threads_get
    • First observedgmail_threads_list
    • First observedgsheets_spreadsheets_values_append
    • First observedgsheets_spreadsheets_values_get
    • First observedgsheets_spreadsheets_values_update
    • First observedlinear_projects_list
    • First observednotion_pages_create
    • First observedshopify_order_get
    • First observedslack_chat_post_message
    • First observedslack_users_list
    • First observedstripe_accounts_list
    • First observedstripe_application_fees_list
    • First observedstripe_balance_retrieve
    • First observedstripe_balance_transactions_list
    • First observedstripe_charges_list
    • First observedstripe_customers_list
    • First observedstripe_customers_retrieve
    • First observedstripe_disputes_list
    • First observedstripe_disputes_retrieve
    • First observedstripe_disputes_update
    • First observedstripe_invoice_items_create
    • First observedstripe_invoices_create
    • First observedstripe_invoices_list
    • First observedstripe_invoices_send
    • First observedstripe_payouts_list
    • First observedstripe_prices_create
    • First observedstripe_products_create
    • First observedstripe_refunds_create
    • First observedstripe_subscriptions_list
    • First observedstripe_subscriptions_update

TDQS

B3.1/5.0

Scored across 39 tools

Disambiguation4/5

Tools are mostly distinguishable by service, resource, and action, with clear separation between Stripe billing operations and the other app integrations. Minor ambiguity exists within similar Google Sheets value operations and between Stripe balance retrieval vs. balance transaction listing, but descriptions help resolve these.

Naming Consistency4/5

Most tools follow a consistent service_resource_action snake_case pattern, e.g. stripe_invoices_send, gmail_threads_list, gsheets_spreadsheets_values_append. The meta tools connect and connection_status break the pattern, and a few singular/plural choices vary, but the convention is still highly readable.

Tool Count2/5

39 tools is heavy for a server named stripe-billing-ops-mcp, especially since many are one-off tools across unrelated apps like GitHub, Slack, Notion, and Shopify. The set sprawls beyond a focused billing-ops surface and many tools do not feel like they earn their place.

Completeness2/5

The surface has significant gaps: shopify_order_get references an orders_list tool that is absent, gmail_messages_attachments_get references messages_get, and gdocs_documents_create references documents_batch_update. Several integrations expose only list or create operations, and Stripe itself lacks core lifecycle operations such as customer create/update, subscription create/cancel, and payment intent handling.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents and human operators to manage Stripe payment operations, SaaS billing, subscriptions, refunds with safety caps, and Austrian/EU tax compliance including VAT and BAO record retention.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to manage subscriptions, usage-based billing, payments, refunds, tax compliance, and invoicing to drive revenue growth.
    1
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Runs 20 slash-command workflows across Google Calendar, Gmail, Linear, Slack, Granola, Google Docs, GitHub, Stripe, Notion, Google Forms, Sheets and Drive to produce morning briefs, meeting prep, action items assigned to owners and weekly updates. Reads proceed without asking, while anything that creates, sends, changes or deletes is shown for approval first.
    47
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables freelancers and agencies to run eight back-office workflows across Stripe, Google Drive, Linear, Google Calendar, Gmail, GitHub, Google Docs, Granola, Google Sheets and Firecrawl, covering client onboarding, invoices from calendar and commits, status reports, scope-creep detection, site audits and overdue invoice chasers. Reads run freely, while anything that creates, sends, changes or deletes is shown for approval first, with credentials kept in your own OS keychain and no proxying through any third-party server.
    Apache 2.0