InferIndex
Server Details
LLM API prices across 70+ providers: cheapest offer, comparisons, history and cost estimates.
- Status
- Healthy
- Uptime
- 100.0% over 21 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- InferIndex/inferindex-docs
- GitHub Stars
- 0
TDQS
Scored across 8 tools
cheapest, compare_providers, and estimate_cost all answer 'what does model X cost across providers' and all return offers sorted cheapest-first with the winner in detail. estimate_cost is separable by its workload inputs, but the boundary between cheapest and compare_providers (which differ mainly in filters and a detail flag) is blurry and invites misselection.
A verb_noun core exists (compare_providers, estimate_cost, list_gpus, search_models), but it is mixed with bare noun phrases (price_history, gpu_rentals), an adjective (cheapest), and a phrase-like name (self_host_or_api). Readable but no single predictable pattern.
Eight tools is well-scoped for a price-tracking service: search, per-provider comparison, cost estimation, GPU listing/rentals, history, and a self-host decision tool. Each maps to a distinct piece of the domain without filler.
The lifecycle is largely covered: find a model, compare offers, estimate workload cost, inspect history, price GPUs, and decide self-host vs API. Minor gaps remain, e.g. no tool to browse or enumerate the full model/provider catalog without a name query, but agents can work around this via search_models.
Available Tools
8 toolscheapestCheapest offers for a modelARead-onlyIdempotentInspect
Cheapest current API offers for one model across direct providers and aggregators, in USD per 1M tokens (input, output, blended 3:1). Stale prices, and flex/batch tiers, are excluded by default. Optional filters (context, tools, JSON, vision, region, no training on prompts, open sign-up) and usage (tokens per request, requests per day) to get an estimated cost per request and per month. Returns the winner in detail and one short line per following offer. A condition that is absent was not published by the provider, it never means "no"; a flag not listed in signals is false.
| Name | Required | Description | Default |
|---|---|---|---|
| json | No | Only offers that support JSON output | |
| limit | No | Number of offers to return, winner included (default 5, max 25) | |
| model | Yes | Model id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure. | |
| tools | No | Only offers that support tool calling | |
| detail | No | compact (default): the winner in detail and one short line per other offer (provider, prices, signals, training on prompts when published). full: every field of every offer (all published conditions, reliability), much longer | |
| region | No | Only providers that process data in this region: eu, us, … | |
| strict | No | Exclude offers whose provider does not publish the filtered information (by default they are kept and flagged) | |
| vision | No | Only offers that accept image input | |
| min_context | No | Minimum context window in tokens | |
| no_training | No | Only providers whose published terms say they do not train on your prompts | |
| no_waitlist | No | Only providers with open sign-up (no waitlist, invitation or country restriction) | |
| cached_ratio | No | Share of input tokens served from the provider's prompt cache (0 to 1) | |
| include_tiers | No | Also include lower-priority service tiers, comma-separated: flex, batch (hidden by default) | |
| output_tokens | No | Output tokens per request, for the estimated cost | |
| prompt_tokens | No | Input tokens per request, for the estimated cost | |
| requests_per_day | No | Requests per day, to also get an estimated monthly cost |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnly/idempotent/non-destructive, the description adds real semantic context: default exclusion of stale prices and flex/batch tiers, and the important interpretation rule that an absent condition means 'not published', never 'no', and that unlisted flags are false. It does not discuss rate limits or fresh-vs-cached data timeliness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and price unit, then filters, then return shape. Sentences are dense but every clause carries information; the interpretation rule about absent conditions is worth its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool with no output schema and no destructive semantics, the description covers purpose, default behavior, filters, return shape, and interpretation of missing data. A brief pointer to which sibling to use when would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description adds meaning beyond the schema by explaining the default behavior of detail (compact winner+short lines), default limit (5), and the strict/absent-condition interaction. This helps the agent interpret filters correctly rather than just pass them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: current cheapest API offers for one model across direct providers and aggregators, with the price unit (USD per 1M tokens, input/output/blended 3:1). This clearly separates it from siblings like compare_providers and price_history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: it explains default exclusions (stale prices, flex/batch tiers) and that optional filters can narrow results, but never says when to prefer this over compare_providers or estimate_cost. The one routing hint (use search_models when unsure) lives in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_providersCompare providers for a modelARead-onlyIdempotentInspect
Current offers for one model, one line per provider and source (direct or via an aggregator), cheapest first (10 by default), with price, quantization, signals and training on prompts when published; the cheapest offer comes in detail. Pass detail: "full" for context, every published condition (data regions, sign-up) and reliability from official status pages on each offer.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | Sort order (default blended); estimated_cost needs prompt_tokens or output_tokens | |
| limit | No | Number of offers to return, cheapest included (default 10, max 50) | |
| model | Yes | Model id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure. | |
| detail | No | compact (default): the winner in detail and one short line per other offer (provider, prices, signals, training on prompts when published). full: every field of every offer (all published conditions, reliability), much longer | |
| region | No | Only providers that process data in this region: eu, us, … | |
| strict | No | Exclude offers whose provider does not publish the filtered information (by default they are kept and flagged) | |
| no_training | No | Only providers whose published terms say they do not train on your prompts | |
| no_waitlist | No | Only providers with open sign-up (no waitlist, invitation or country restriction) | |
| cached_ratio | No | Share of input tokens served from the provider's prompt cache (0 to 1) | |
| include_tiers | No | Also include lower-priority service tiers, comma-separated: flex, batch (hidden by default) | |
| output_tokens | No | Output tokens per request, for the estimated cost | |
| prompt_tokens | No | Input tokens per request, for the estimated cost | |
| requests_per_day | No | Requests per day, to also get an estimated monthly cost |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/non-destructive profile, and the description adds real behavioral value on top: default result count, cheapest-first ordering, the default-vs-full verbosity tradeoff, and that providers lacking published info are kept and flagged unless `strict` is set. It does not describe pagination or how missing data appears in output, but the annotation bar is met.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core behavior (what a row is, ordering, default size, winner-in-detail). Both sentences carry information, though the second is long and slightly run-on with its list of caveats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter comparison tool with no output schema, the description covers the result shape and the compact/full distinction well enough to call it correctly. Remaining gaps (pagination, precise flagging format) are minor given the annotations and high schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 13 parameters, including the two enums. The description restates the `detail` semantics (context, published conditions, status-page reliability) but adds little that the schema does not already carry, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource precisely (current offers for one model, one line per provider and source) and the ordering (cheapest first, 10 by default), which lets an agent distinguish it from `cheapest` and `estimate_cost`. It stops short of explicitly contrasting itself with those siblings, so a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-to-use for the `detail: "full"` option, but never says when to reach for this tool over `cheapest`, `estimate_cost` or `price_history`, nor does it state prerequisites for the comparison. Usage is implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_costEstimate the cost of a workloadARead-onlyIdempotentInspect
Estimated cost of a workload on one model at each provider: cost per request, and per month if requests_per_day is given, taking the provider's tiered pricing and prompt-cache price into account. Offers sorted by estimated cost, cheapest first: the cheapest in detail, then one short line per offer with its estimated cost.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of offers to return, cheapest included (default 5, max 25) | |
| model | Yes | Model id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure. | |
| detail | No | compact (default): the winner in detail and one short line per other offer (provider, prices, signals, training on prompts when published). full: every field of every offer (all published conditions, reliability), much longer | |
| cached_ratio | No | Share of input tokens served from the provider's prompt cache (0 to 1) | |
| output_tokens | No | Output tokens per request, for the estimated cost | |
| prompt_tokens | No | Input tokens per request, for the estimated cost | |
| requests_per_day | No | Requests per day, to also get an estimated monthly cost |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds real behavioral context the annotations don't: tiered pricing and prompt-cache pricing are factored in, offers are sorted cheapest-first, and the return layout (one detailed winner plus short lines) is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and then the output structure, with no filler. It is dense but every clause carries information; slightly heavy sentence construction keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by describing the returned offer list and its ordering. For a 7-parameter tool it leaves the alternative-selection question unanswered, but nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema by tying requests_per_day to a monthly estimate and explaining that the prompt-cache price feeds the cached_ratio calculation. It does not, however, supply defaults or formats for prompt_tokens/output_tokens.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('estimated cost of a workload on one model at each provider') and even characterizes the output shape, so the agent knows exactly what it computes. It does not, however, contrast itself with the sibling tools (cheapest, compare_providers), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the conditional 'per month if requests_per_day is given' hints at when the monthly figure appears, but there is no explicit statement of when to pick this over compare_providers or cheapest, nor any prerequisite guidance. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gpu_rentalsRental offers for one GPU, by tierARead-onlyIdempotentInspect
Rental offers for one GPU type: provider, price in USD per GPU per hour, number of GPUs in the published configuration, billing and region when published. One block per tier (guaranteed, community, spot), each with its own cheapest offer, also given per configuration (cheapest_by_gpu_count: the price per GPU of a single GPU and of a multi-GPU node are not comparable); tiers are never compared with each other. guaranteed: capacity the provider presents as not interrupted; community: third-party hosts; spot: interruptible capacity. Prices are as published by the provider, and dated. Get the GPU id from list_gpus.
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | Yes | GPU id from list_gpus, for example h100-sxm-80gb | |
| tier | No | Only this tier (default: one block per tier) | |
| detail | No | compact (default): the 3 cheapest offers of each tier. full: up to 50 offers per tier | |
| region | No | Only this region, as published by the provider | |
| gpu_count | No | Only this published configuration: 1, 2, 4, 8 or 16 GPUs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent and non-destructive, so safety is covered. The description goes beyond them by defining each tier's meaning (guaranteed vs community vs spot), noting default volume (3 cheapest per tier, up to 50 for full), and stating prices are provider-published and dated with non-comparable normalizations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what is returned before the tier and comparability caveats, and no sentence is pure filler. The telegraphic style is dense but compact enough for the amount of semantics it conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does so: one block per tier, each with its own cheapest offer and per-configuration breakdown. Combined with tier definitions and the price caveat, an agent has enough to call and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real meaning beyond the schema by explaining tier semantics and the cheapest_by_gpu_count normalization rule (single-GPU vs multi-GPU node prices are not comparable), which affects how the agent should interpret results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope: rental offers for one GPU type, broken out by tier, with the fields returned named (provider, price, GPU count, billing, region). It is clear what the tool returns, but it does not explicitly differentiate itself from the sibling 'cheapest', which sounds like it also surfaces low-price offers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to list_gpus for the required GPU id and warns that tiers and per-configuration prices are never comparable, which is useful usage context. However, it never says when to prefer this tool over 'cheapest', 'compare_providers' or 'estimate_cost', so alternative selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_gpusGPU types tracked for rental pricesARead-onlyIdempotentInspect
GPU types whose rental price InferIndex tracks (H100, A100, L40S…), with the lowest price per GPU per hour in USD for each tier (guaranteed, community, spot) and the number of GPUs of that configuration. Prices are per GPU per hour, as published by each provider, and dated. Tiers are different products and are never compared with each other.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds real semantic context beyond that: prices are per GPU per hour as published by each provider and dated, and the explicit rule that tiers are different products and are never compared with each other — a caveat that prevents incorrect cross-tier aggregation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is tracked and then the per-tier pricing rule. Dense but every clause carries meaning; the parenthetical examples (H100, A100, L40S) and tier list are useful rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must describe the return payload, and it does so precisely: one row per tracked GPU type with lowest price per GPU/hour, tier, and GPU count. It omits ordering and any notion of how many entries come back, but for a small enumerated list this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the schema cannot carry semantic load and the baseline for a zero-parameter tool applies. The description appropriately spends its words on what is returned rather than on nonexistent inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (GPU types tracked by InferIndex) and enumerates the returned fields: lowest price per GPU/hour in USD per tier, tier names, and GPU counts. This is far more specific than a tautology, though it never names a sibling tool to distinguish itself from 'cheapest' or 'gpu_rentals'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named, despite siblings like 'cheapest', 'gpu_rentals', and 'compare_providers' that plausibly overlap. The only implicit cue is that this lists tracked GPU types rather than answering a pricing question, which the agent must infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_historyPrice history of a modelARead-onlyIdempotentInspect
Price history of one model: every offer tracked by InferIndex (daily or weekly min/max/last price in USD, or raw price changes), plus the official price of the model's lab over time. Give either days, or from/to (YYYY-MM-DD), or at (a date) for the prices in effect that day.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | A single date, YYYY-MM-DD: prices in effect that day | |
| to | No | End date, YYYY-MM-DD | |
| days | No | Number of days back from today (default 7) | |
| from | No | Start date, YYYY-MM-DD (with to, instead of days) | |
| limit | No | Maximum number of points (default 100, max 500) | |
| model | Yes | Model id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure. | |
| provider | No | Only this provider | |
| granularity | No | day (default), week, or raw price changes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so safety is covered. The description adds behavioral detail beyond that by specifying the data scope: 'every offer tracked by InferIndex', granularity options (daily/weekly min/max/last price, raw price changes), and the inclusion of the lab's official price over time.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence front-loads the resource ('one model') and the return content; the second gives a compact usage rule for date selection. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description needs to cover return content, which it does by listing USD min/max/last prices, raw changes, and official lab price. Missing details like ordering or pagination are minor and not essential for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds value by clarifying that the date parameters are alternatives ('Give either days, or from/to, or at'), which is not made explicit in the individual parameter descriptions. It does not need to repeat schema details, so 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool returns 'Price history of one model', making the resource and scope clear. It contrasts with sibling tools like cheapest and compare_providers by focusing on historical tracking rather than comparisons or estimates, so an agent can distinguish it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to specify date windows ('Give either days, or from/to, or at'), which is parameter-level guidance. However, it does not explicitly state when to prefer this tool over search_models, cheapest, or compare_providers; the 'one model' phrase provides only weak inference rather than clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_modelsSearch modelsARead-onlyIdempotentInspect
Find the exact id of an LLM tracked by InferIndex from a name or partial name (e.g. 'deepseek', 'qwen3 max', 'claude opus'). Returns matching model ids and names, best match first.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Model name or part of it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive. The description adds valuable context: it returns matching IDs and names, orders by best match first, and supports partial names. This goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the primary purpose front-loaded and no filler. Examples are embedded naturally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description appropriately states what is returned (IDs and names, best match first). It omits edge-case handling (e.g., no results) but is sufficient for a typical lookup scenario.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single query parameter. The description adds examples and clarifies partial matching, which is helpful but only a modest enhancement over the schema's 'Model name or part of it'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Find') and names the resource ('exact id of an LLM tracked by InferIndex'), with concrete examples. It clearly distinguishes this lookup tool from cost-focused siblings like cheapest and estimate_cost.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied for retrieving model IDs from names, but there are no explicit alternatives or when-not-to-use conditions. The tool's purpose makes the context clear, but it does not state exclusions or trade-offs versus sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_host_or_apiSelf-host an open model or use the API?ARead-onlyIdempotentInspect
Is it cheaper to host an open-weights model yourself on rented GPUs, or to use the cheapest API offer? Returns a one-sentence verdict (headline) and the numbers behind it: the break-even volume in million tokens per day, the throughput the whole configuration must deliver for self-hosting to cost less (required_tokens_per_second: compare it with what your setup does), the break-even utilization when a throughput is known, the GPU configuration and its hourly price, and what the verdict rests on (confidence). Without a volume you still get the verdict and the break-even volume. It is an estimate: the throughput is a published measurement (measured), an estimate scaled from one (derived; estimated for a wide range), or a published minimum (lower_bound: a floor, valid for requests of up to 2,048 tokens in total, that can show self-hosting wins, never that the API is cheaper); or you give your own tokens_per_second. When no published figure decides, the headline gives the break-even volume and the required throughput, to compare with yours, rather than a yes or no, and may point out a dearer setup that published figures do decide (settled_alternative: for comparison, not a recommendation). headline_kind and headline_values give the headline as fields. It declines to give a number (status refused, with a reason) when it cannot do so reliably, for example a closed model, weights that do not fit the requested GPU, or a GPU nobody rents. Hardware rental only: the hourly price the provider publishes for the machine. It does not count the engineering to deploy and run the model, monitoring, redundancy, start-up time, or storage and network costs billed separately.
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | No | Impose a GPU: id from list_gpus, for example h100-sxm-80gb | |
| tier | No | GPU rental tier (default guaranteed) | |
| model | Yes | Model id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure. | |
| detail | No | compact (default): verdict, headline and key numbers. full: also the three throughput scenarios, the model and every assumption | |
| gpu_count | No | Impose a number of GPUs: 1, 2, 4, 8 or 16 | |
| utilization | No | Average share of time the GPUs serve requests, in percent (default 50) | |
| quantization | No | Precision of the weights (default fp8); must be published by at least one provider unless fp16 or bf16 | |
| output_tokens | No | Output tokens per request | |
| prompt_tokens | No | Input tokens per request | |
| tokens_per_day | No | Your volume in tokens per day (input and output together). Or give requests_per_day, prompt_tokens and output_tokens | |
| weights_margin | No | Percent of the GPU memory the model weights may fill (default 70; the rest is the engine reserve and the context cache) | |
| requests_per_day | No | Requests per day (with prompt_tokens and output_tokens) | |
| tokens_per_second | No | Your own measured throughput for the whole configuration, input and output tokens together, instead of our estimate |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the read-only/idempotent safety profile, and the description adds substantial context beyond that: measured vs derived vs lower_bound throughput semantics (including the 2,048-token floor for lower_bound), refusal behavior and its causes, the settled_alternative field being for comparison not recommendation, and an explicit list of what costs are excluded (engineering, monitoring, redundancy, storage/network).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core question and the headline verdict, and every clause carries information. It is dense with nested parentheses in a single block, which hurts scanability, but there is little pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 13-parameter analytical tool with no output schema, the description fully specifies the return payload (headline, headline_kind/values, break-even volume, required throughput, confidence), the refusal path, and the modeling limitations, so an agent has what it needs to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, setting the baseline at 3, but the description adds real meaning: it distinguishes tokens_per_day from the requests_per_day + prompt/output token path, explains tokens_per_second as a user-supplied override of the estimate, and clarifies the compact/full detail trade-off.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise decision question (self-host an open-weights model vs cheapest API) and enumerates exactly what is returned (verdict/headline, break-even volume, required throughput, GPU config, confidence). This is clearly distinguishable from siblings like estimate_cost and cheapest, which price rather than compare hosting strategies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operational conditions well: it works with or without a volume, tokens_per_second overrides the estimate, and it declines with 'status refused' for closed models or unfit GPUs. It does not, however, explicitly route the agent away from or toward sibling tools (e.g., when to prefer estimate_cost or gpu_rentals).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- Added
gpu_rentals - Added
list_gpus - Added
self_host_or_api
3 tool updates
- Changed
cheapest2 fields changed- added
Input schema / properties / detailAdded value: +{ + "description": "compact (default): the winner in detail and one short line per other offer (provider, prices, signals, training on prompts when published). full: every field of every offer (all published conditions, reliability), much longer", + "enum": [ + "compact", + "full" + ], + "type": "string" +} - changed
Input schema / properties / limit / descriptionPrevious value: -"Number of offers to return (default 5, max 25)"New value: +"Number of offers to return, winner included (default 5, max 25)"
- Changed
compare_providers2 fields changed- added
Input schema / properties / detailAdded value: +{ + "description": "compact (default): the winner in detail and one short line per other offer (provider, prices, signals, training on prompts when published). full: every field of every offer (all published conditions, reliability), much longer", + "enum": [ + "compact", + "full" + ], + "type": "string" +} - changed
Input schema / properties / limit / descriptionPrevious value: -"Number of offers to return (default 10, max 50)"New value: +"Number of offers to return, cheapest included (default 10, max 50)"
- Changed
estimate_cost2 fields changed- added
Input schema / properties / detailAdded value: +{ + "description": "compact (default): the winner in detail and one short line per other offer (provider, prices, signals, training on prompts when published). full: every field of every offer (all published conditions, reliability), much longer", + "enum": [ + "compact", + "full" + ], + "type": "string" +} - changed
Input schema / properties / limit / descriptionPrevious value: -"Number of offers to return (default 5, max 25)"New value: +"Number of offers to return, cheapest included (default 5, max 25)"
2 tool updates
- Changed
cheapest1 field changed- changed
Input schema / properties / no_waitlist / descriptionPrevious value: -"Only providers with open sign-up (no waitlist or invitation)"New value: +"Only providers with open sign-up (no waitlist, invitation or country restriction)"
- Changed
compare_providers1 field changed- changed
Input schema / properties / no_waitlist / descriptionPrevious value: -"Only providers with open sign-up (no waitlist or invitation)"New value: +"Only providers with open sign-up (no waitlist, invitation or country restriction)"
5 tool updates
- First observed
cheapest - First observed
compare_providers - First observed
estimate_cost - First observed
price_history - First observed
search_models
Related MCP Connectors
Live LLM API pricing: token prices, comparisons, cheapest-model lookups. No key required.
Live LLM API price + status radar across 11 providers, with public per-model price HISTORY.
Find AI model pricing, estimate token costs and compare offers. No API key required.
Compare LLM API prices, search models and providers, and access reviewed benchmark results.
Related MCP Servers
- AlicenseAqualityAmaintenanceLive LLM API pricing: current token prices, model comparisons, cheapest-model lookups, and The LLM Price Index for 150+ models across 20+ providers, re-verified daily. No API key required.41MIT
- AlicenseAqualityAmaintenanceDaily-verified LLM API pricing dataset (44+ models, CN & global) with a hosted MCP server for live price queries and token cost estimation.2CC BY-4.0
- AlicenseNot gradedqualityCmaintenanceToken cost math for LLM API calls: current per-million-token rates for 69 models across 17 providers, with local arithmetic for estimates, comparisons and monthly budgets. Rates are verified and date-stamped.37 npm2MIT
- AlicenseNot gradedqualityDmaintenanceCompare AI inference pricing across 9 providers in real time. Routing recommendations, spend tracking, and budget alerts for AI agents.119 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.