Skip to main content
Glama

Scaffold the cost-governor-kit spend-safety pattern

scaffold_cost_governor

Generate a runtime cost-governor kit for paid AI API calls: pre-call spend estimates, cache-aware pricing, advisory usage counting, install steps, starter code, and worked examples.

Instructions

cost-governor-kit is a RUNTIME LIBRARY your app installs and calls at the moment it's about to make an AI API call -- it is not something this MCP server can 'check' on demand the way it checks a document's citations. Its advisory check-then-commit piece (withReserveConfirm) needs a live UsageLedger backed by YOUR database and an async callback making the real call, neither of which exist as content to hand this tool. Call this to get all three pieces explained (estimated pre-call spend check, cache-aware pricing math, advisory successful-call usage counting), an install step, a copy-pasteable starter snippet, and (optionally) a live worked example: real checkPreCallCeiling calls (one allowed, one blocked), a real withReserveConfirm run against an in-memory demo ledger (one call under the limit, one over, one whose commit fails after a successful call), and a real estimateCostUsd comparison of the default 0.1x cache-read rate against a caller-supplied cacheReadPerMillion override. Use this when you're building or reviewing anything that calls a paid AI API and want a pre-call spend estimate plus usage counting that skips a call that threw -- NOT a strict concurrent limit and not proof a timed-out call was never billed by the provider (see concurrency_and_recovery in the output, and use withCapacityReservation instead if you need a real reservation).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
includeWorkedExampleNoAlso run live examples against the real checkPreCallCeiling/withReserveConfirm, not just show a snippet.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it explains that the withReserveConfirm piece needs a live UsageLedger and async callback that cannot be handed to this tool, and that the worked example runs against an in-memory demo ledger (so no external data dependency). It also scopes the tool's limits (advisory only, not a hard concurrency limit, not proof of provider billing). It stops short of stating side effects of running the live example (latency/cost/credentials), keeping it just under a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is substantive and mostly earns its place, but it is delivered as one dense, meandering paragraph with heavy parenthetical nesting for a tool that takes one optional boolean. It is not front-loaded on the single most important fact, and readability suffers; it would be far stronger as a short lead sentence plus structured bullets.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must carry the full load, and it does enumerate what is returned (three explained pieces, install step, snippet, optional live examples with specific cases) and names the concurrency_and_recovery section of the output. It is nearly complete, only leaving unclear whether the live examples incur latency or require credentials.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage for the single includeWorkedExample parameter, so the baseline is 3. The description restates the same idea ('optionally... a live worked example') and enumerates what the examples contain, which adds content detail but no syntax or format meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (explain the three cost-governor-kit pieces, provide an install step, a starter snippet, and optionally live worked examples) for a specific resource. It also distinguishes itself from the sibling check_* tools by stressing it is a runtime library, not an on-demand checker. The purpose is clear, though it is buried in a long first sentence rather than stated crisply up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when (building or reviewing anything that calls a paid AI API and wanting a pre-call spend estimate plus usage counting that skips a call that threw). It names exclusions (NOT a strict concurrent limit, not proof a timed-out call was never billed) and routes to an alternative explicitly: use withCapacityReservation instead if you need a real reservation. This is model usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.