compute_estimate
Estimate the USD cost of an LLM call by model, input tokens, and output tokens, returning base provider price, routing fee, and billed amount for budgeting.
Instructions
Nominal USD cost for any oracle-tracked model given input/output token counts — index members and catalog-only models on identical terms. Cache reads and cache writes belong inside input_tokens and are charged at the full input rate here; no cache discount is applied. Returns base_usd_cost (provider list price), routing_fee_usd and billed_usd_cost (what compute.finance charges), plus the routing_fee_rate they derive from. Compare models on base_usd_cost; budget on billed_usd_cost. A provenance mark says how far a number has been checked: 'verified' — an operator recorded a vendor source; 'inferred' — derived from a sibling number or a vendor default, with no source recorded; 'promotional' — a discounted list price expected to end. Every value bills exactly as shown: the mark rates trust in the number, not the amount charged, and holds as of the operator's last pass rather than a live vendor check. base_price_provenance marks the input and output prices this cost is built from, populated for every tracked model. The cost is quoted at the rung input_tokens selects on the model's ladder, returned whole as applied_context_tier so the rate behind the number is visible; for every rung read context_tiers on data_get_price. A long-context price ladder ascends by from_input_tokens and always has at least one rung: the first starts at 0 and restates the model's flat rate, so a model priced the same at every input size has exactly one rung and needs no special case. The rung is chosen by the whole input side of a request — prompt plus cache reads plus cache writes — over half-open ranges, so an input landing exactly on from_input_tokens takes that rung. Rung rates carry base_*/billed_* like every other price; only the flat rate enters the SCU index. Each rung carries the same {input, output} provenance pair as elsewhere: the first repeats the base price's mark; a higher rung is a catalogue number and takes on both directions the single mark the vendor quotes it under. max_input_tokens is the largest input the model accepts. It is null when the model declares no window of its own — not unbounded: the request-body ceiling still applies, there is just no per-model limit. Above a declared window the request is refused before it reaches the provider. A catalogue that cannot be read, or that does not list the model, errors this tool instead of quoting, because a cost silently computed at the flat rate would understate the long context the ladder exists to price. exceeds_max_input_tokens is true when input_tokens is above max_input_tokens: the cost is still quoted, because a refused request is worth pricing before you reshape it, but the request as supplied would be rejected. price_source ('oracle-basket' | 'oracle-catalog') names the serving endpoint only and does not change the pricing basis, so two models with the same provider price return the same cost. Reasoning tokens are billed inside output_tokens, so this estimate adds no separate reasoning leg; read the model's reasoning price from data_get_price. Errors with 'Model not tracked by oracle' for unknown keys. Source: Oracle API (/v1/oracle/resolve + /v1/oracle/catalog). For a cost with the cache discount applied, use analyze_session on a real transcript. Models are identified by their canonical vendor-prefixed id ('anthropic/claude-sonnet-4.6', 'openai/gpt-5.5'); the bare name ('gpt-5.5') resolves to the same model, and the response echoes the canonical id. The vendor slug is not always provider.key (alibaba → qwen, xai → x-ai, moonshot → moonshotai), so pass an id the API returned rather than assembling one.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Canonical vendor-prefixed model id (e.g. 'anthropic/claude-sonnet-4.6'); the bare name also resolves. | |
| input_tokens | Yes | Whole input side of the request — prompt plus cache reads plus cache writes. Cache tokens are charged at the full input rate here, but they also count toward the size that picks the rung, so leaving them out quotes a rung too low. | |
| output_tokens | Yes |