Skip to main content
Glama

Server Details

Stripe billing for indie apps: plans, coupons, subscriptions, customers and revenue reporting.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
38.8% over 21 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.8/5.0

Scored across 25 tools

Disambiguation4/5

Most tools have clearly distinct resource+action targets (list_coupons vs invite_to_coupon vs set_coupon_listed; update_plan_pricing vs update_plan_limits vs update_plan_meta). There is some overlap between single/bulk coupon variants and between setup_pricing_page and the standalone plan tools, but descriptions make the boundaries explicit (batch vs single, wizard vs standalone).

Naming Consistency4/5

Strong, predictable verb_noun snake_case throughout: create_coupon, delete_coupon, list_apps, update_plan_pricing, revoke_coupon_invite. Minor deviations are the bulk_ prefix and the noun-first outliers plan_app_kit, revenue_by_app, and revenue_summary, but overall it is consistent and readable.

Tool Count3/5

At 25 tools this sits at the heavy end for a single server, spanning coupons, plans, subscriptions, revenue, referrals, and onboarding. Reasonable given the broad domain, but the single/bulk coupon pairs and the wizard-vs-standalone duplication add surface area that could be consolidated.

Completeness4/5

Covers the core lifecycle well: apps, plan catalog (create/archive/reprice/limits/metadata), full coupon CRUD plus invites/listing and bulk ops, customers, subscriptions, payments, cancellations, referrals, and revenue rollups. Minor gaps like referral status mutation and explicit refund operations are the only notable omissions.

Available Tools

25 tools
archive_planAInspect

Remove a tier from an app's plan catalog. WRITE, human-gated: a real write needs an elicitation-capable client and the human must approve; dry_run previews anywhere. Every other plan tool edits tiers that already exist, so a catalog could grow but never shrink, and collapsing a multi-tier page down to a single offer was a dashboard-only job. Stripe products and prices are deliberately left intact, exactly as update_plan_pricing leaves the old price behind: removing the row stops the plan being offered and never reaches into anyone's existing subscription. Refuses while the tier still has non-cancelled subscribers (matched on its Stripe price ids, since subscriptions carry a price rather than a tier slug), and refuses to remove an app's last remaining tier. Run bun run plans:sync in the app repo afterwards to regenerate plans.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
tierYes
app_idYes
dry_runNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it declares the write is human-gated and requires an elicitation-capable client, that Stripe products/prices are deliberately left intact, that it refuses when non-cancelled subscribers exist (matching on Stripe price ids) and when it would remove the last tier, and that `bun run plans:sync` must follow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purposethen layers constraints efficiently in four dense sentences. There is minor narrative padding ('was a dashboard-only job') that could be trimmed, but every sentence still carries actionable constraint information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no annotations, no output schema, and 0% schema coverage, the description covers safety profile, refusal conditions, side effects on Stripe, and the required follow-up step. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the tier's semantics (matched to live subscriptions via its Stripe price ids) and dry_run (previews anywhere), but app_id is left as self-evident and no format/example is given for tier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Remove a tier from an app's plan catalog.' It also differentiates from siblings by noting that every other plan tool only edits existing tiers, so an agent can place this apart from update_plan_meta/update_plan_pricing without reading their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the exact conditions: dry_run previews anywhere, a real write needs an elicitation-capable client and human approval, and it names update_plan_pricing as the tool that leaves old prices behind in the analogous case. When-to-use and when-not are both explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bulk_create_couponAInspect

Create the SAME discount code across many apps in one call: the portfolio-promo shortcut (e.g. a summer sale on every owned app). WRITE, human-gated: a real create requires an elicitation-capable client; the human enters the coupon terms ONCE in a form (overriding the arguments here) and that one approval covers the whole batch. Clients without elicitation are refused (dry_run still works anywhere). Runs each app through the same fully-guarded create as create_coupon. Unlike a single call it does NOT fail the batch when an app already has the code: that app is reported skipped and the rest are still created. Returns a per-app outcome matrix plus created/skipped/failed counts. Use dry_run:true first to see which apps would create vs skip. Pass explicit app_ids (from list_apps). There is deliberately no 'all apps' shortcut, so a discount can't hit an app you didn't name.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
listedNo
app_idsYes
dry_runNo
max_usesNo
expires_atNo
descriptionNo
applies_to_tiersNo
discount_percentYes
confirm_high_discountNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it declares WRITE, human-gated approval via elicitation that overrides arguments, refusal for clients lacking elicitation (except dry_run), and a key contrast with the single call — existing codes are reported `skipped` rather than failing the batch. It also discloses the return shape (per-app outcome matrix plus created/skipped/failed counts).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and dense with useful detail; nearly every sentence adds a distinct operational fact. Slightly repetitive around the elicitation/human-gating point, which costs it a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Batch behavior, approval gating, skip semantics, and return counts are well covered for a complex tool, and no output schema exists to offload that. However, for a 10-parameter tool with 0% schema coverage, six parameters remain undefined, notably confirm_high_discount and max_uses/expires_at, leaving the agent under-informed on how to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 parameters, so the description must compensate but only meaningfully explains app_ids (pass explicit ids from list_apps) and dry_run (use first to preview). Code and discount are only implied, and listed, max_uses, expires_at, description, applies_to_tiers, and confirm_high_discount are left entirely undocumented in both description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: 'Create the SAME discount code across many apps in one call.' It explicitly frames itself as the bulk counterpart to create_coupon and adds the portfolio-promo use case, so an agent can distinguish it from the single-create sibling without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance ('summer sale on every owned app'), a prerequisite (requires an elicitation-capable client; non-elicitation clients are refused), a recommended first step ('Use dry_run:true first'), and the alternative (create_coupon for single calls). It also warns there is deliberately no 'all apps' shortcut and to pass explicit app_ids from list_apps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bulk_delete_couponAInspect

Delete the SAME coupon code across many apps in one call: the inverse of bulk_create_coupon, for retiring a portfolio-wide promo. WRITE, human-gated: a real delete requires an elicitation-capable client; the human sees every app the code would be removed from (with per-app redemption counts) and must click approve, and confirm_used is only consulted for dry_run previews. Clients without elicitation are refused (dry_run still works anywhere). An app that never had the code is reported skipped rather than failing the batch, since the desired end state already holds there. Returns a per-app outcome matrix plus deleted/skipped/failed counts. Use dry_run:true first to see exactly which apps would lose the code. Pass explicit app_ids (from list_apps) so a delete can't hit an app you didn't name.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
app_idsYes
dry_runNo
confirm_usedNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it declares this is a WRITE, is human-gated via elicitation, is refused on non-elicitation clients, reports never-had-the-code apps as `skipped` rather than failing, and returns a per-app outcome matrix with deleted/skipped/failed counts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The verb and scope are front-loaded, and nearly every clause adds a distinct behavioral fact. It runs long and dense, but the length is justified by the elicitation gate, skip semantics, and return-shape details rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating batch tool with no annotations and no output schema, the description supplies the missing safety profile, failure semantics, preview path, and return-value shape. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate; it explains dry_run as a preview toggle, that confirm_used is only consulted for dry_run previews, that app_ids should be explicit (sourced from list_apps), and that code targets the same coupon across apps. Coverage is strong though it does not spell out types/defaults, which the schema already holds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Delete the SAME coupon code across many apps in one call.' It explicitly differentiates itself from siblings by naming bulk_create_coupon as its inverse and positioning itself as a multi-app operation distinct from single-app delete_coupon.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit operational guidance: use dry_run:true first to preview affected apps, and pass explicit app_ids from list_apps so a delete cannot hit unnamed apps. It also states the client prerequisite (elicitation-capable) and when confirm_used applies, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_app_kitAInspect

Step 2 of onboarding a new app onto SparkPay. Applies answers.plan_catalog by minting real Stripe prices (same helper /api/apps/sync-plans uses) and registers/updates the app, then returns a downloadable integration kit (lib/spark-pay client + plans.json + sync script + INTEGRATION.md, modeled on still-applying's golden pattern) plus exact post_steps for env vars and secrets. Refuses to overwrite an existing app's plans unless answers.confirm_overwrite is true. Real secrets (a push-mode webhook secret) are only ever returned in post_steps, never written into a kit file.

ParametersJSON Schema
NameRequiredDescriptionDefault
answersYes
app_nameYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses real side effects (minting actual Stripe prices via the sync-plans helper, registering/updating the app), a destructive-action guard (overwrite refused without confirm_overwrite), and a security property (secrets only returned in post_steps, never written to a kit file). These are exactly the traits an agent needs before invoking a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The workflow step is front-loaded and the sentences are dense but each carries distinct information (side effects, guard, output contents, secret handling). Slightly heavy for a single paragraph, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, nested, no-annotation, no-output-schema tool, the description covers side effects, outputs (kit contents, post_steps), and mutation guards. Remaining gap is that several nested input fields and the exact return shape are unspecified, but the essential calling context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with nested objects, so the description should compensate. It adds real meaning to answers.plan_catalog (minted into real Stripe prices), answers.confirm_overwrite (overwrite gate), and push_mode (push-mode webhook secret), but leaves success_url, cancel_url, trial_days, and the monthly/yearly/payment_type distinctions undocumented. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Applies answers.plan_catalog by minting real Stripe prices ... registers/updates the app') and frames itself precisely as 'Step 2 of onboarding a new app onto SparkPay'. The unique outputs (integration kit + post_steps) distinguish it from siblings like plan_app_kit and setup_pricing_page without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Step 2 of onboarding' framing implies a preceding step and a clear point in the workflow, and it states a conditional guard ('Refuses to overwrite ... unless answers.confirm_overwrite is true'). However, it never names the sibling alternative (e.g. plan_app_kit as the preview/first step), so the routing is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_couponAInspect

Create a percentage discount coupon for an app. WRITE, human-gated: a real create requires an elicitation-capable client (e.g. Claude Desktop). The human enters the final code, percent, expiry, visibility, and tier scope in a form; those values OVERRIDE the arguments here, which only pre-fill the form message as a proposal. Clients without elicitation are refused (dry_run still works anywhere). Inserts a coupons row and best-effort-mints the matching Stripe coupon (checkout backfills it on first use if Stripe is unreachable). The code is uppercased and must be unique per app. By default the coupon applies to every paid tier. Set dry_run:true first to preview the exact row without writing. Cannot discount a free-only app. Three house rules are enforced and cannot be argued out of: every coupon expires within 6 months (omitting expires_at means the 6-month term, not a code that never expires); a coupon at 90%+ is created invite_only, so it is redeemable only by addresses added with invite_to_coupon; and a 90%+ coupon is pinned to the second-highest paid recurring tier and may never cover the top tier or a one-time (lifetime) plan. Use list_coupons afterwards to confirm, and list_plans to see an app's tier slugs.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
app_idYes
listedNo
dry_runNo
max_usesNo
expires_atNo
descriptionNo
applies_to_tiersNo
discount_percentYes
confirm_high_discountNo

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the human-gated write mechanics, elicitation requirement, Stripe best-effort minting with checkout backfill, the 6-month expiry cap, 90%+ invite_only pinning, tier restrictions, code uppercasing/uniqueness, and the free-only-app prohibition — excellent transparency. However, it does not state permissions/auth requirements or error conditions beyond refusal, and the return/response shape is left unspecified; against zero annotation coverage these remain meaningful gaps, warranting a 3.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the WRITE/human-gated flag are front-loaded, and nearly every sentence adds a distinct rule or precondition rather than filler. It is nonetheless a dense wall of clauses, which slightly hurts scanability for a 10-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 10-parameter mutation with no annotations and no output schema, the description covers behavior, failure modes, the three enforced house rules, dry-run semantics, and next-step tooling. An agent has everything needed to call it correctly or avoid it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description must compensate, and it does for most inputs: code (uppercased, unique per app), discount_percent (90%+ rules), expires_at (omitting means the 6-month default), dry_run, listed/tier scope mentioned as form fields, and applies_to_tiers defaulting to every paid tier. It omits max_uses, description, and confirm_high_discount, so it is strong but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'Create a percentage discount coupon for an app.' It is immediately distinguishable from bulk_create_coupon, update-style siblings, and the read-only list_coupons without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when the tool works (elicitation-capable client), when it is refused (non-elicitation clients), that dry_run works anywhere, to run dry_run first to preview, that the form values override arguments, and names follow-up tools (list_coupons, list_plans, invite_to_coupon). Routing and preconditions are fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_couponAInspect

Delete one coupon from an app. WRITE, human-gated: a real delete requires an elicitation-capable client; the human sees the exact coupon (code, percent, redemption count) and must click approve. The model cannot confirm on its own, and confirm_used is only consulted for dry_run previews. Clients without elicitation are refused (dry_run still works anywhere). Removes the Stripe coupon object and then the coupons row, exactly like the dashboard's delete, so no orphaned Stripe object is left behind. Deleting is permanent and the code becomes reusable afterwards. Set dry_run:true to see exactly which coupon would go (and its redemption count) without touching anything. Use list_coupons first to confirm the code.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
app_idYes
dry_runNo
confirm_usedNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: human-gated elicitation requirement, model cannot self-confirm, permanence of the delete, the code becoming reusable afterwards, ordered side effects (Stripe object then DB row), and the client-capability refusal behavior. This is behavior far beyond the bare name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The safety-relevant info (human-gated, permanent) is front-loaded and every clause carries weight, but the middle section is dense and run-on, stacking mechanism, capability gating, and side effects into long sentences. Slightly more structure would help without losing content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a permanent, human-gated mutation with no annotations, no output schema, and 4 parameters, the description covers safety, side effects, reversibility, dry-run semantics, and prerequisite discovery. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate: it explains dry_run (preview without touching anything), confirm_used (only for dry_run previews), and code (surfaced to the human with percent and redemption count). app_id is left to inference, which is the only gap in an otherwise strong compensation for zero schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete one coupon from an app') and pins the singular scope, which distinguishes it from the sibling bulk_delete_coupon. The mechanism is spelled out ('removes the Stripe coupon object and then the `coupons` row... no orphaned Stripe object'), so an agent knows exactly what is destroyed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use list_coupons first to confirm the code,' use dry_run:true to preview, and confirms that non-elicitation clients are refused while dry_run still works anywhere. It also distinguishes confirm_used's narrow role ('only consulted for dry_run previews'), which prevents misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customerAInspect

Look up a customer by email. Returns every subscription/registration for that email across apps (or within one app if app_id is given), with full detail: verification, plan, billing period, usage, refunds. Empty when the email is unknown.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYes
app_idNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the burden. It usefully discloses scope behavior (all apps vs one) and the terminal case ('Empty when the email is unknown'), which tells the agent not to expect an error. It does not state read-only nature, auth requirements, or whether results are paginated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action, and every clause earns its place by adding scope or return-shape information. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 0% parameter coverage, the description does the necessary work by summarizing the returned fields and the empty-result case. It is complete enough to invoke correctly, though auth/rate-limit context is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it does: email is the lookup key, app_id scopes to one app. The enumeration of returned detail (verification, plan, billing period, usage, refunds) further characterizes what the key resolves to. Format/syntax of the email value is left unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Look up a customer by email'. This is clearly distinguishable from sibling list-style tools (list_subscriptions, list_coupons) since it pivots on a single customer identity. No sibling is named for routing, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one conditional: pass app_id to narrow the lookup to a single app, omit it to span all apps. That is implied usage guidance but there is no explicit when-to-use-this-vs-siblings or when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

invite_to_couponAInspect

Add email addresses to a coupon's invite list and, by default, switch the coupon to invite-only so the list is actually enforced. WRITE, human-gated: a real write requires an elicitation-capable client and the human sees every address plus whether the coupon is about to be closed to everyone else; clients without elicitation are refused (dry_run works anywhere). Addresses are lowercased and de-duplicated, malformed ones are reported rather than dropped, and an address already on the list is left exactly as it is -- so re-inviting someone who already redeemed does NOT hand them a second free account, and re-inviting someone who was revoked does NOT un-block them. Set enforce_invite_only:false only to stage a list before gating the code; the result then carries a warning that the coupon is still open. Use dry_run:true first to see exactly who would be added.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
app_idYes
emailsYes
sourceNo
dry_runNo
enforce_invite_onlyNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it is declared a WRITE that is human-gated, requires an elicitation-capable client (clients without it are refused, though dry_run works anywhere), lowercases/de-duplicates addresses, reports rather than drops malformed ones, leaves existing entries untouched (no second free account, no re-blocking of revoked users), and warns in the result when enforce_invite_only:false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph, but it is front-loaded with purpose and the enforcement behavior before drilling into edge cases. Nearly every clause earns its place; it leans long, though the density is justified by the behavioral detail being conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter write tool with no annotations and no output schema, this is largely complete: it covers safety gating, idempotency, error handling, and a result warning. The only gap is the unexplained 'source' parameter and no explicit statement of the returned payload shape beyond the warning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains dry_run, enforce_invite_only, and emails semantics (lowercasing, de-dup, malformed handling) thoroughly, but the 'source' parameter is never mentioned and code/app_id are only implied, leaving one of six params undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (add addresses) plus resource (coupon's invite list) and discloses the key side effect (by default switches coupon to invite-only). This distinguishes it cleanly from siblings like revoke_coupon_invite, list_coupon_invites, and create_coupon.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance: set enforce_invite_only:false only to stage a list before gating, and use dry_run:true first. It covers when-not for one parameter but never explicitly routes to a sibling alternative (e.g., revoke_coupon_invite for removal), so it stops short of full alternatives coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_appsAInspect

List the registered apps (app_id + name). Use this first to discover valid app_id values, then pass one to the other tools to scope results to a single product.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It indicates a read-only enumeration and reveals the return fields, but says nothing about pagination, ordering, volume limits, or whether apps are user- or workspace-scoped. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what it returns and immediately followed by the workflow instruction. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description sensibly summarises the returned fields (app_id + name), and with no parameters and no annotations there is little else an agent needs. Minor gaps around result size or ordering keep it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to disambiguate and the baseline is 4. The description correctly implies no input is required to discover app ids.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (list the registered apps) and even names the payload fields (app_id + name), which is unusually concrete. It does not explicitly contrast itself with the sibling list_* tools, but its role as an app-id discovery tool is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit sequencing guidance: call this first to discover valid app_id values, then pass one to other tools to scope results. That is a clear when-to-use directive with a stated purpose. It stops short of naming alternatives or conditions under which not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_coupon_invitesAInspect

List who a coupon was offered to and what they did with it: per-address status (invited / redeemed / revoked), where the invite came from, and the timestamps. Also returns the coupon's own invite_only flag -- read that FIRST: when it is false the invite rows are recorded but NOT enforced, and anyone holding the code can still redeem it. Read-only. Use list_coupons for the coupon's terms, invite_to_coupon to add addresses, revoke_coupon_invite to withdraw one.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
app_idYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares 'Read-only' and discloses a non-obvious behavioral nuance — that invite rows are recorded but NOT enforced when invite_only is false. It does not mention auth/permission requirements or result-size limits, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the return shape front-loaded, then the critical flag caveat, then routing guidance. Every clause earns its place, though the parenthetical status list and the routing sentence make it slightly denser than the minimum needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description properly takes on the job of describing the return shape (statuses, source, timestamps, invite_only) and safety profile. Only the unexplained app_id parameter and absent auth/limit notes leave a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is described in the schema, so the description must compensate. It implies that 'code' identifies the coupon being inspected, but never names or explains 'app_id', leaving one of two required parameters unclarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource ('List who a coupon was offered to') and enumerates the returned fields (per-address status, invite source, timestamps, invite_only flag). It also explicitly distinguishes itself from list_coupons, invite_to_coupon, and revoke_coupon_invite, so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The final sentence names three sibling tools and states what each is for instead, giving explicit when-to-use and when-not-to-use guidance. It also instructs the agent to read the invite_only flag FIRST and explains why that ordering matters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_couponsAInspect

List discount coupons (code, percent off, expiry, max uses, use count, which paid tiers they apply to). Optionally scope by app_id and/or restrict to currently-active ones (not expired, not maxed out). Read-only: cannot create, edit, or delete coupons.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
app_idNo
active_onlyNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the read-only safety profile ('cannot create, edit, or delete coupons'), which is the key behavioral trait. It omits pagination behavior and what the default 50-item limit means in practice, leaving meaningful gaps for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences, front-loaded with the resource and its return fields, followed by scope options and the read-only guarantee. No filler and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the enumerated return fields are a valuable substitute, and scope/safety are covered. The only shortfall is the undocumented limit parameter and pagination behavior, which keeps it just under fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds genuine meaning for app_id (scoping) and active_only ('not expired, not maxed out'), which the bare schema does not convey, but the limit parameter and its 50 default are never explained, so compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (discount coupons), then enumerates the returned fields (code, percent off, expiry, max uses, use count, tier applicability), which pins the purpose down precisely. It does not explicitly distinguish itself from the similar sibling list_coupon_invites, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real usage context — optional scoping by app_id and optional restriction to active coupons — but never states when to prefer this over list_coupon_invites or when filtering is required. Usage is implied rather than prescribed, with no alternatives or exclusions named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_plansAInspect

Download an app's live plan catalog: tiers, names, prices, trial, and the machine-readable limits that gate features inside the app. This is the source of truth each app repo's bun run plans:sync regenerates lib/spark-pay/plans.json from, so diff against this rather than trusting the committed file. Stripe price/product IDs are omitted unless include_stripe_ids is set.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idYes
include_stripe_idsNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work: it characterizes the data as live rather than committed, and discloses the default suppression of Stripe price/product IDs. It does not mention auth requirements, rate limits, or pagination, which are the remaining gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place, with the return contents front-loaded and the guidance and parameter caveat following. No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly describes the return shape, and with no annotations on a read-style catalog tool the remaining burden is light. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does for include_stripe_ids by explaining exactly what that flag gates. app_id is left undefined, though its meaning is inferable from the sibling list_apps tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Download') and resource ('an app's live plan catalog') and enumerates the payload contents (tiers, names, prices, trial, machine-readable limits). This is clearly distinguishable from siblings like list_apps, list_coupons, and the update_plan_* mutators.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage context: this is the source of truth that `bun run plans:sync` regenerates plans.json from, so the agent should diff against it rather than the committed file. It does not name alternative tools or state when not to use it, but the intended context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recent_cancellationsAInspect

Recently canceled paid subscriptions (excludes free rows), most recent cancellation first, with what each customer was paying. Use for churn review; list_subscriptions orders by signup date instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
app_idNo
include_samplesNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the free-row exclusion, the sort order, and that payment amounts are returned, but says nothing about read-only semantics, pagination/limit behavior, or how app_id scoping changes results. Solid but incomplete for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste: the scope/filtering is front-loaded and the sibling routing comes after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema it covers the return content, ordering, and exclusions well. However, with no annotations and undocumented parameters, an agent still lacks any signal on limit semantics, app_id scoping, or what include_samples does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, and the description mentions none of them. The opaque include_samples flag and the app_id scoping parameter are left entirely unexplained, so the description does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (recently canceled paid subscriptions), defines scope precisely (excludes free rows), and specifies ordering (most recent cancellation first). It also names a sibling, list_subscriptions, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use for churn review' and names the alternative with the condition that selects it ('list_subscriptions orders by signup date instead'). This is exactly the when-to-use/when-to-use-something-else guidance the dimension asks for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recent_paymentsAInspect

List recent paid transactions (excludes free-tier rows), newest first, ordered by payment date. Optionally scope by app_id. Use to review recent sales activity. Amounts are in the transaction currency (cents).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
app_idNo
include_samplesNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden; it usefully discloses the free-tier exclusion, the sort order, and the currency/cents unit. However, it omits pagination behavior, whether results are truncated by the default limit of 25, and any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action and scope. 'Newest first, ordered by payment date' is mildly redundant but the description contains little waste overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no annotations and no output schema, the description covers scope, ordering, and units, which is a reasonable baseline. It is still incomplete on include_samples semantics and limit/result-window behavior, which an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains the meaning of app_id scoping. The limit parameter (default 25) and the include_samples flag are never mentioned, leaving two of three parameters semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List recent paid transactions') plus the defining scope ('excludes free-tier rows') and ordering ('newest first, ordered by payment date'). This clearly separates it from siblings like list_recent_cancellations and revenue_summary without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use to review recent sales activity' implies a use case but names no alternatives or when-not conditions. Given siblings such as revenue_by_app and revenue_summary, an agent gets no explicit routing guidance between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_referralsBInspect

List referral records (who referred whom, reward coupon code, pending/rewarded status), newest first. Optionally scope by app_id and/or status.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
app_idNo
statusNo
include_samplesNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose useful traits: the default ordering (newest first) and the shape of what is returned. It says nothing about pagination, the default limit, what include_samples toggles, or any auth/permission requirements, which leaves meaningful behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no wasted words; the resource and ordering are front-loaded and the optional scoping trails. Efficient, though the parenthetical field list is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and 4 parameters with 0% schema description coverage. Two inputs (limit, include_samples) are unexplained anywhere, and return-volume/pagination behavior is unaddressed, so the definition is not complete enough for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must document 4 parameters and only explains two (app_id, status) — and the status values it repeats are already an enum in the schema. limit (default 25) and include_samples are entirely undocumented, leaving half the inputs opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (referral records) and even enumerates the returned fields (who referred whom, reward coupon code, status) plus the sort order (newest first). No sibling tool covers referrals, so no differentiation is needed, but the description stops short of anything beyond a clear purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Optionally scope by app_id and/or status' implies when the filters apply, but there is no explicit when-to-use guidance, no prerequisites, and no alternatives to route between. Adequate but leaves the agent to infer the purpose of a listing call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_subscriptionsAInspect

List subscriptions/registrations, newest first. Optionally filter by app_id, status, or payment_type. Returns metadata (email, status, plan, amounts). Use get_customer for a single user, or revenue_summary for totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
app_idNo
statusNo
payment_typeNo
include_samplesNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses sort order ('newest first'), that filters are optional, and roughly what is returned (email, status, plan, amounts), but says nothing about pagination/limit behavior, permissions, or result caps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and ordering, then filters, then return content, then alternatives. No filler and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations, so the description does well to name return fields and alternatives, but the two undocumented parameters (limit, include_samples) leave a gap for a list tool where pagination matters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It adds real meaning for three of five params (app_id, status, payment_type as optional filters), but 'limit' and 'include_samples' are left entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List subscriptions/registrations'), a sort order ('newest first'), and the filterable dimensions. It explicitly names the sibling tools it is not (get_customer, revenue_summary), so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear routing conditions: get_customer for a single user, revenue_summary for totals. It does not address when to prefer this over other list-style siblings (list_recent_payments, list_recent_cancellations, list_plans), so the guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_app_kitAInspect

Step 1 of onboarding a new app onto SparkPay. Recons app_name against existing apps and the rest of the portfolio's plan catalogs, then returns automated findings plus the minimum set of questions that genuinely need a human answer (plan prices, billing model, redirect URLs, trial length, push-webhook opt-in). Pass the answers to create_app_kit next.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_nameYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does useful work by disclosing the return shape (automated findings plus a minimum set of human questions) despite there being no output schema, but it never states whether the call is side-effect-free, what permissions are needed, or what happens if the app_name already exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero waste, front-loaded with the step position and workflow. The parenthetical enumerating the question categories is dense but each item is load-bearing information the agent would otherwise have to guess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter planning tool with no annotations and no output schema, the description supplies the missing return-value context and the downstream step. Only the absence of any statement about safety/side effects keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 0% schema description coverage, so the description must compensate. It does: app_name is the string reconciled against existing apps and portfolio plan catalogs, which is more meaning than the bare schema string type conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (recon app_name against existing apps and plan catalogs) on a specific resource, and explicitly positions itself as "Step 1 of onboarding a new app onto SparkPay." That sequencing distinguishes it cleanly from create_app_kit and the plan-mutation siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it (onboarding a new app) and routes the agent forward: "Pass the answers to create_app_kit next." It lacks a when-not clause (e.g., don't run for an app that already exists), so it falls just short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revenue_by_appAInspect

Per-app rollup in USD cents (latest_payments/refunded/MRR + paid customer count), largest refund-adjusted latest payments first. latest_payments sums the most recent payment per customer, so it is a run-rate and NOT lifetime turnover. No lifetime figure is available per app (accrual is recorded per owner). Use to compare products at a glance; use revenue_summary for one app in depth.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_samplesNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers important nuance: the sort order (largest refund-adjusted latest payments first) and the crucial warning that latest_payments is a run-rate, NOT lifetime turnover, plus why no lifetime figure exists. It still omits return shape/pagination and the effect of include_samples, so it falls short of fully self-sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and metric, then the caveats, then the sibling routing. Every sentence carries information, though the run-rate explanation could be trimmed slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the metric semantics are well covered, but the include_samples parameter is undocumented and there is no hint of return structure or pagination. Adequate but with a clear gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter include_samples has 0% schema coverage and is never mentioned in the description, so the agent gets no signal about what it does or when to toggle it. The description compensates well for metric semantics but ignores the one actual input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and output scope: a per-app rollup with named metrics (latest_payments/refunded/MRR + paid customer count) in USD cents. It explicitly distinguishes itself from the sibling revenue_summary, so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when to use this tool ('compare products at a glance') versus the alternative ('use revenue_summary for one app in depth'). The alternative and its selecting condition are both stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revenue_summaryAInspect

Revenue and subscriber summary in USD cents (normalized from each transaction currency). Returns latest_payments (the MOST RECENT payment summed per customer: a run-rate, NOT lifetime turnover, because amount_paid is overwritten on renewal), refunds, approximate MRR, active/trialing/paid counts, and lifetime_net (the only cumulative total held, accrued per OWNER across all their apps and net of refunds; null when no accrual record exists). lifetime_net counts only charges whose Stripe webhook reached this app, so a missing renewal event under-counts it silently. A 0 or a low figure means "not recorded here", never "no sales". For turnover, accounts or tax, read Stripe directly: it is authoritative, and this database cannot reconstruct lifetime gross. Optionally scope by app_id; omit for platform-wide totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idNo
include_samplesNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so exceptionally: it explains that latest_payments is a run-rate not lifetime turnover because amount_paid is overwritten on renewal, that lifetime_net is accrued per OWNER across apps and net of refunds, that missing webhooks under-count silently, and that 0 means 'not recorded here' rather than 'no sales'. These are exactly the caveats an agent needs to avoid misreporting.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense and front-loaded with the return payload before the caveats. Nearly every clause earns its place by preventing misinterpretation; the only mild cost is the run-on phrasing of the lifetime_net sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe returns, and it covers the payload and its pitfalls thoroughly, including null/under-count semantics. The one gap is the undocumented include_samples parameter, which a caller cannot infer from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with two parameters, so the description must compensate. It does explain app_id well ('Optionally scope by app_id; omit for platform-wide totals'), but include_samples is never mentioned anywhere, leaving one of two parameters entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names the specific resource (revenue and subscriber summary) and its unit (USD cents, normalized), then enumerates the exact payload returned (latest_payments, refunds, MRR, active/trialing/paid counts, lifetime_net). This lets an agent distinguish it from siblings like revenue_by_app or list_recent_payments without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the agent away for adjacent needs — 'For turnover, accounts or tax, read Stripe directly: it is authoritative' — and explains the app_id scoping choice ('omit for platform-wide totals'). It stops short of naming the obvious sibling revenue_by_app as the per-app alternative, so the exclusion set is clear but the sibling routing is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_coupon_inviteAInspect

Withdraw one address's right to spend a coupon: the targeted block. WRITE, human-gated: a real write requires an elicitation-capable client and the human sees the address and its current status; clients without elicitation are refused (dry_run works anywhere). Works on an address that was never invited too, writing a revoked tombstone, so blocking someone does not depend on whether they were on the list and survives a later bulk re-invite. This only blocks FUTURE redemptions: if the status was already redeemed the discount has been granted, and undoing that means cancelling or refunding the subscription in the dashboard. Prefer this over delete_coupon when one recipient misbehaves -- deleting the code punishes every other invitee. Use list_coupon_invites first to see the current status.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
emailYes
app_idYes
reasonNo
dry_runNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it declares this is a WRITE that is human-gated, requires an elicitation-capable client, refuses non-elicitation clients, and that dry_run works anywhere. It also discloses tombstone semantics (works on never-invited addresses, survives re-invite) and that it only blocks future redemptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb and the WRITE/gating caveat, then structured in earned sentences. It is dense and slightly overstuffed, but nearly every clause adds decision-relevant behavior rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, yet the description covers the write model, gating, edge cases, and sibling relationships thoroughly. The remaining gap is the undocumented app_id and reason parameters, which are neither in the schema nor the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for 5 params, so the description must compensate. It clarifies the address/email target and dry_run behavior well, but app_id, code, and reason are never explained. Partial compensation only justifies a middling score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Withdraw one address's right to spend a coupon') and immediately scopes it ('the targeted block'). It distinguishes itself from siblings by naming delete_coupon and list_coupon_invites, so an agent can route correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Prefer this over delete_coupon when one recipient misbehaves'), when-not ('if the status was already redeemed the discount has been granted'), and a required prerequisite ('Use list_coupon_invites first to see the current status').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_coupon_listedAInspect

Show, hide or SCHEDULE an EXISTING coupon on the public pricing pages by setting its listed flag and optional listed_from start date. WRITE, human-gated: a real write requires an elicitation-capable client and the human sees the code, the discount, which way it is moving and any start date; clients without elicitation are refused (dry_run works anywhere). These are deliberately the ONLY fields an agent can change on a live coupon: the percent, expiry and tier scope are terms somebody may already have been quoted, while listed/listed_from are purely a display decision. Together they decide whether the public coupons API returns the code today, which is what every app's promo bar draws from; the discount itself does not move and no Stripe object is touched. listed_from in the future STAGES a seasonal coupon so it appears on its own date with no cron: it never blocks anyone holding the code from redeeming it, it only holds the code back from the promo bar. Listing with no listed_from clears any existing schedule (lists now). Listing an INVITE-ONLY coupon is refused outright, since that would advertise a near-free code to people who cannot redeem it. A listed_from with listed:false is refused as a contradiction. Listing an expired/maxed coupon, or scheduling a start at or after expiry, is allowed but warns. A coupon already in the requested state reports noop. Use list_coupons first to see codes and their current visibility.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
app_idYes
listedYes
dry_runNo
listed_fromNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it discloses that writes are human-gated via elicitation, that non-elicitation clients are refused, that `dry_run` is safe anywhere, that no Stripe object is touched, that future `listed_from` stages without a cron, and that listing invite-only coupons or contradictory schedules are refused. It also explains warnings and `noop` behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded, starting with the operation and then layering write gating, field semantics, and edge cases. It is longer than average, but most sentences add distinct and useful constraints for a complex mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% schema description coverage, the description is remarkably complete. It covers authentication/elicitation requirements, dry-run behavior, refusal conditions, warning cases, idempotent `noop` responses, and the fact that redemption is not blocked by future scheduling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It richly explains the semantics of `listed`, `listed_from`, and `dry_run`, including scheduling, clearing, and contradiction rules, but does not explicitly define `app_id` or `code` beyond context, leaving minor explanatory work to the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource+scope: setting the `listed` flag and optional `listed_from` on an EXISTING coupon to show, hide, or schedule it on public pricing pages. It distinguishes this from creation, deletion, and listing siblings such as `create_coupon` and `list_coupons`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to use `list_coupons` first to see codes and visibility, explains the human-gated write flow, notes that `dry_run` works without elicitation, and lists when writes are refused or warned. This gives an agent clear when-to-use and when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

setup_pricing_pageAInspect

Guided, stage-chained wizard that configures an app's hosted pricing page end to end: plan billing model + price (with clickable recommendations), plan name and badge (elite/recommended), optional free tier, a seasonal coupon (SUMMER26-style code suggested from the current season), favicon-derived theme color + card style, and FAQ content. WRITE, human-gated per stage: each call runs AT MOST ONE elicitation form and commits only that stage, then returns next_call for the model to continue (stages: overview -> plans -> coupon -> branding -> content -> finish). The values the HUMAN types in each form OVERRIDE the arguments here - args only pre-fill form defaults (pass price_suggestions from market research and faq_drafts to seed them). Testimonials are NOT configurable here: they are collected from real customers on Sellular and rendered from there. Form stages are refused on clients without elicitation; overview/finish are read-only recon and work anywhere, as does dry_run (echoes the form + would-be writes, writes nothing). A declined stage leaves earlier committed stages applied - each stage is individually approved, and each effect is reversible via the standalone tools. Stripe prices are minted only by the plans stage with human-typed amounts (existing paid tiers only; free tiers never touch Stripe). After plan changes, run bun run plans:sync in consumer app repos. Use list_apps first for app_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
tierNo
stageNooverview
app_idYes
domainNo
dry_runNo
faq_draftsNo
coupon_codeNo
discount_percentNo
price_suggestionsNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: WRITE, human-gated per stage, one elicitation form per call, human-typed values override args, declined stages leave earlier commits intact, each effect reversible via standalone tools, Stripe prices minted only by plans stage. Return contract (returns next_call) is disclosed. Only minor gap is it doesn't state authorization/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core purpose and stage-chain model appear before the finer rules. It is long (roughly 200 words), yet most sentences encode distinct constraints (overrides, refusal behavior, Stripe minting, sync command) rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 9-param, 6-stage, write-plus-elicitation tool with no output schema and no annotations, the description covers behavior, sequencing, override semantics, and reversibility thoroughly. The remaining omission is the meaning/format of the tier and domain params.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

At 0% schema coverage with 9 params, the description compensates substantially: it explains stage ordering, price_suggestions ('pass from market research'), faq_drafts ('to seed them'), dry_run ('echoes the form + would-be writes, writes nothing'), and coupon_code shape (SUMMER26-style, seasonal). It leaves tier, domain, and discount_percent unexplained, but coverage is strong overall.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('configures an app's hosted pricing page end to end') and enumerates the exact configuration surface (plans, coupon, branding, content). It is clearly distinguishable from the sibling standalone tools like update_plan_pricing or create_coupon, which it positions itself as orchestrating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit on when and how: 'Use list_apps first for app_id', the stage chain sequence, 'Testimonials are NOT configurable here', and the note that form stages are refused on clients without elicitation while overview/finish/dry_run work anywhere. It also tells the agent to re-run `bun run plans:sync` after plan changes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_plan_limitsAInspect

Change the usage limits on an app's existing tiers (e.g. set ai_generations_daily to 200 on pro). WRITE, human-gated: a real write requires an elicitation-capable client; the human sees the per-tier before -> after diff and must click approve. Clients without elicitation are refused (dry_run still works anywhere). Tier-keyed, so one call can set a new limit key across every tier in a single atomic write. Use -1 for unlimited, matching the catalog convention. This writes ONLY the limits object: it never touches Stripe, prices, or plan names, so it cannot mint a price by accident. Changing prices or adding a tier is a different job: use create_app_kit or the dashboard. Set dry_run to preview the before/after without writing. After this succeeds, run bun run plans:sync in the app repo and commit the regenerated plans.json so the code and the catalog agree.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNomerge
app_idYes
dry_runNo
updatesYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden and does so: it declares this is a WRITE, human-gated via elicitation with a before->after diff and approval click, that non-elicitation clients are refused, that dry_run works anywhere, that the write is atomic across tiers, and that Stripe/prices/plan names are untouched. This is exactly the kind of safety and scope context an agent needs before invoking a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long but front-loaded with the core action, and nearly every clause adds operational value (gating, scope, dry_run, follow-up sync). A few clauses could be tightened, but there is little pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param, nested, mutation tool with no annotations and no output schema, the description covers authorization, gating, refusal behavior, preview mode, scope limits, and post-write synchronization. An agent has everything needed to call it correctly or avoid it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage and a nested updates object, the description compensates well: it explains the tier-keyed shape of updates, the -1 unlimited convention, and what dry_run does. It does not explain the mode enum (merge vs replace) or the default, which is the one remaining semantic gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Change the usage limits on an app's existing tiers', with a concrete example (set ai_generations_daily to 200 on pro). It explicitly demarcates itself from siblings by saying it writes ONLY the limits object and that changing prices or adding a tier is a different job handled by create_app_kit or the dashboard.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (existing tiers, limit keys, -1 for unlimited), when-not ('Changing prices or adding a tier is a different job: use create_app_kit or the dashboard'), and a preview path via dry_run. It also names the prerequisite (elicitation-capable client) and the required follow-up step after success.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_plan_metaAInspect

Set an existing tier's display metadata: its name and the two badge flags the public pricing page styles cards from: is_elite (gold/carbon BEST VALUE card with the rainbow title and app icon) and recommended (MOST POPULAR). A tier called "Elite" is NOT styled as elite unless is_elite is actually true on the stored plan; the page reads the flag, never the tier slug, so a catalog whose top tier lacks the flag renders as a plain grey card. Use this to repair that. WRITE, human-gated: the human sees the before -> after diff and approves (clients without elicitation pass through, since this touches no money: it never reaches Stripe, never changes an amount, limits or billing model, and is reversible by calling it again). Setting a flag false REMOVES the key rather than storing false, keeping exported plans.json minimal. Convention is at most ONE is_elite and ONE recommended tier per app: this tool does not clear the flag from its siblings, so unset the old holder in a second call. Use list_plans first for tier slugs and current flags, and run bun run plans:sync in the app repo afterwards to regenerate plans.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
tierYes
app_idYes
dry_runNo
is_eliteNo
recommendedNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses that this is a WRITE, human-gated operation, that clients without elicitation pass through, that it never reaches Stripe or changes amounts/limits/billing, that it is reversible by calling again, and that setting a flag false removes the key rather than storing false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loads the core action and then layers behavioral and workflow details in a logical order. Most sentences earn their place, though the density and multiple parenthetical asides make it slightly heavy rather than maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no annotations and no output schema, the description covers purpose, safety, side effects, formatting convention, and required follow-up sync steps very well. The main gap is that it never accounts for the app_id or dry_run parameters, so an agent lacks full parameter context despite the otherwise comprehensive operational detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all six parameters. It gives rich meaning for name, is_elite, and recommended, and implies tier via "list_plans first for tier slugs," but it never explains app_id or dry_run, leaving two parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: set an existing tier's display metadata, naming the exact fields (name, is_elite, recommended) and their visual effects on the pricing page. It clearly distinguishes this tool from siblings like update_plan_limits and update_plan_pricing by focusing on display metadata rather than money or limits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use context: repair a catalog whose top tier lacks the is_elite flag, use list_plans first for slugs and current flags, and run bun run plans:sync afterwards. It also states the one-is_elite/one-recommended convention and that the tool does not clear siblings, so a second call is required to unset the old holder.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_plan_pricingAInspect

Change what an app's EXISTING tiers charge: the amount, the billing model (one_time vs subscription), or both. WRITE, human-gated at the highest tier: a real write requires an elicitation-capable client, and the human types the tier, billing model and price into a form. Those typed values OVERRIDE the arguments here, which only pre-fill the proposal, so the model cannot move a price on its own. Clients without elicitation are refused (dry_run still works anywhere). Stripe prices are immutable, so this mints a NEW price and repoints the plan; the old price object is left intact and existing subscribers keep the price they signed up on, so this changes what NEW buyers pay. Switching a tier to one_time also clears its monthly/yearly price ids, so checkout cannot fall back to a recurring price the pricing page no longer advertises. Free, contact and payg tiers are refused: converting those is a catalog change, not a reprice. Use list_plans first for tier slugs and current amounts, and run bun run plans:sync in the app repo afterwards to regenerate plans.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idYes
dry_runNo
updatesYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: human-gated elicitation at the highest tier, typed form values overriding the arguments, Stripe price immutability, minting a new price while leaving the old object intact, existing subscribers retaining their price, and one_time switching clearing monthly/yearly price ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but front-loaded with the core purpose, then constraints, then edge cases and follow-up. Nearly every sentence adds operational value, though the density is near the upper limit for a three-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema and 0% schema description coverage, the description is unusually complete: it covers authorization, override semantics, immutability outcomes, subscriber impact, refusal cases and a required follow-up sync step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: it maps tier, amount_cents, payment_type (one_time vs subscription) and monthly_amount_cents to their meaning and notes that arguments only pre-fill the proposal. It does not explain app_id or clarify amount units (cents) beyond the schema's naming.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Change what an app's EXISTING tiers charge') and enumerates the mutable dimensions (amount, billing model). It is clearly distinguishable from siblings like update_plan_limits and update_plan_meta, which touch different aspects of a plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names prerequisites and routing ('Use list_plans first for tier slugs and current amounts'), the post-step (run plans:sync), and the exclusion set (free, contact and payg tiers are refused because that is a catalog change, not a reprice). It also states when the tool will not work at all (non-elicitation clients, unless dry_run).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 25 tool updates
    • First observedarchive_plan
    • First observedbulk_create_coupon
    • First observedbulk_delete_coupon
    • First observedcreate_app_kit
    • First observedcreate_coupon
    • First observeddelete_coupon
    • First observedget_customer
    • First observedinvite_to_coupon
    • First observedlist_apps
    • First observedlist_coupon_invites
    • First observedlist_coupons
    • First observedlist_plans
    • First observedlist_recent_cancellations
    • First observedlist_recent_payments
    • First observedlist_referrals
    • First observedlist_subscriptions
    • First observedplan_app_kit
    • First observedrevenue_by_app
    • First observedrevenue_summary
    • First observedrevoke_coupon_invite
    • First observedset_coupon_listed
    • First observedsetup_pricing_page
    • First observedupdate_plan_limits
    • First observedupdate_plan_meta
    • First observedupdate_plan_pricing

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    F
    maintenance
    Provides real-time Stripe subscription analytics including MRR, churn, failed payments, and expiring trials. Enables AI assistants to answer business health questions like 'How's my business doing?'
    8
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Universal billing gateway for MCP servers. Add Stripe subscription billing to any MCP server with one line of code. Supports tiered plans, usage tracking, and automatic access control.
    1
    2
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to manage subscriptions, usage-based billing, payments, refunds, tax compliance, and invoicing to drive revenue growth.
    1
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources