FastGPU
Server Details
Compare live GPU cloud rental prices and match workloads to the cheapest provider.
- Status
- Healthy
- Uptime
- 100.0% over 38 days
- Last Tested
- Transport
- Streamable HTTP ยท MCP 2025-11-25
- URL
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: list_gpu_prices is for looking up prices of named GPUs, while match_workload is for sizing and ranking configurations for a described job. Descriptions explicitly guide when to use each, leaving no overlap.
Both names follow a consistent snake_case verb_noun pattern (list_gpu_prices, match_workload). No deviations or mixed conventions.
With only two tools, the surface feels thin for a price-comparison service; common operations like provider-specific queries or historical data are absent. The two tools are broad but the count is borderline low.
The tools cover the core domain: listing live GPU prices and matching workloads to configurations. Minor gaps exist (e.g., no direct way to list all offers for a specific GPU or provider, only the cheapest per model), but the page_url links to full details.
Available Tools
2 toolslist_gpu_pricesList cheapest GPU pricesARead-onlyIdempotentInspect
One entry per GPU model with its cheapest live on-demand cloud rental price in USD per GPU-hour, the provider offering it, and how many providers rent it, from FastGPU's live price feed across marketplaces, neoclouds and hyperscalers (RunPod, Vast.ai, Lambda, AWS and more). Use it to compare GPU rental prices or to find where a named GPU is cheapest to rent. Each price is the provider's own per-GPU rate: cheapest_min_gpu_count says when it is only sold as a multi-GPU instance, and cheapest_billed_separately names what the provider bills on top of it (CPU, memory, storage), null when the rate includes CPU and memory. page_url links to the GPU's live price page with every offer. Rental prices only: it does not rent, reserve or launch GPUs. No key required.
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | Filter by tier. | |
| vendor | No | Filter by GPU vendor. |
Output Schema
| Name | Required | Description |
|---|---|---|
| gpus | Yes | |
| count | Yes | |
| stale | No | |
| updated_at | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial behavior beyond them: the data is a live feed, prices are per-GPU-hour in USD, it clarifies that cheapest_min_gpu_count signals multi-GPU-only sales and cheapest_billed_separately names extra billed resources. It also states no key is required and reinforces the read-only scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core definition is front-loaded in the first sentence, followed by usage, then field-level clarifications. It is dense but every clause carries information; only the parenthetical provider list feels slightly padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not re-explain return structure, and it still goes further by interpreting the key field names and stating the rental-only limitation. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters enum-constrained, so the schema already carries the filter semantics. The description offers no additional meaning for 'tier' or 'vendor' beyond restating that they are filters, so this is the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('one entry per GPU model with its cheapest live on-demand cloud rental price') plus the exact scope (marketplaces, neoclouds, hyperscalers). An agent immediately knows this is a price-listing tool, not a rental or matching tool, which distinguishes it from match_workload.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit use cases ('compare GPU rental prices or find where a named GPU is cheapest to rent') and a clear exclusion ('it does not rent, reserve or launch GPUs'). It does not name match_workload as the alternative for workload-suitability queries, so the routing guidance stops one step short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_workloadMatch a workload to the cheapest GPUARead-onlyIdempotentInspect
Describe a job (an open-weight model to serve or fine-tune, a model size, or a GPU need) and get up to eight ranked GPU configurations that can run it at the lowest live price, each with the GPU count, the effective USD per hour for the whole configuration, and a one-line reason; the result also states the VRAM the job needs and, when a hyperscaler can run the same job, which one and what percent cheaper the top match is. Use it when the user asks where to run a workload or which GPU they need, rather than the price of a named GPU. It sizes open-weight models only: an API-only model such as GPT-4 returns no matches and a note. Results mirror FastGPU's website and apply a small, disclosed tie-break toward providers that pay FastGPU a referral, only between otherwise equal offers (each match reports partner true/false). Optional hard limits (budget_usd_hr, max_total_usd with duration_hours, max_quote_age_minutes) leave out every configuration that breaks them, and limits reports what was applied and how many were left out. page_url links to the GPU's live price page. It does not rent, reserve or launch GPUs. No key required.
| Name | Required | Description | Default |
|---|---|---|---|
| spot | No | Set true to include interruptible spot capacity for a cheaper rate. | |
| task | No | What the job does. | |
| model | No | Open model name to size against, e.g. "Llama 3 70B", "Qwen 72B", "Mixtral". | |
| query | No | Plain-language job, e.g. "cheapest to serve Llama 3 70B", "2x H100 for fine-tuning" or "a GPU with 80GB of memory". Provide this OR a structured spec below. | |
| region | No | Restrict to a data-residency region. | |
| vram_gb | No | Rough VRAM the job needs, in GB, if you already know it. | |
| params_b | No | Model size in billions of parameters when no exact model is named. | |
| reserved | No | Set true to include reserved / committed-term capacity for a lower rate. | |
| gpu_count | No | Exact positive GPU count. Overrides a count in query text. Returns no matches if no supported configuration fits; omit for automatic sizing. | |
| precision | No | Numeric precision to size the model at. | |
| budget_usd_hr | No | Hard limit on the hourly price of the whole configuration, in USD. Configurations above it are not returned. | |
| max_total_usd | No | Hard limit on estimated_total_usd, the GPU rental cost of the whole job, in USD. Needs duration_hours. Configurations above it are not returned. | |
| duration_hours | No | How long the job will run, in hours (fractions are fine). When sent, each match carries billed_hours and estimated_total_usd. | |
| min_gpu_memory_gb | No | Memory each GPU card must have, in GB (e.g. 80 for 80GB-class cards such as the H100 or A100 80GB). Smaller cards are never returned. On its own it lists every card that size, cheapest first. | |
| max_quote_age_minutes | No | Hard limit on how old a price quote may be, in minutes. Configurations whose price was fetched from the provider longer ago are not returned. |
Output Schema
| Name | Required | Description |
|---|---|---|
| hero | No | |
| note | No | Why there are no matches, when there are none. |
| as_of | No | When this answer was computed (ISO 8601). Quote ages are measured against it. |
| count | Yes | |
| error | No | Present when the request was not valid; says what to send instead. |
| stale | No | |
| limits | No | The hard limits this answer applied (null when one was not sent) and how many configurations each one left out. |
| matches | Yes | |
| workload | No | |
| updated_at | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds real behavioral context beyond them: the disclosed referral tie-break between otherwise equal offers (with partner true/false per match), open-weight-only sizing, API-only models returning no matches plus a note, and that it does not rent/reserve/launch. It stops short of describing pagination or result ordering tie-break beyond price, keeping it at a strong 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and output shape are front-loaded in the first sentence, and each subsequent sentence carries distinct value (usage routing, scope limits, tie-break disclosure, hard limits, non-rental, no key). The opening sentence is very dense and could be split, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 15 optional parameters, an existing output schema, and a described output shape, the description covers purpose, usage routing, scope exclusions, cost/limit semantics, provider-bias disclosure, and auth needs (no key required). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by grouping the hard-limit parameters (budget_usd_hr, max_total_usd with duration_hours, max_quote_age_minutes), explaining that any configuration breaking them is excluded, and that the 'limits' report shows what was applied and how many were dropped.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (describe a job) and resource (ranked GPU configurations) with concrete output detail: up to eight configurations, GPU count, effective USD/hr, and one-line reason. It explicitly distinguishes itself from the sibling by noting it is for 'where to run a workload or which GPU they need, rather than the price of a named GPU', which routes cleanly against list_gpu_prices.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('user asks where to run a workload or which GPU they need') and when-not ('rather than the price of a named GPU'), effectively naming the alternative use case served by the sibling. It also gives a clear exclusion: open-weight models only, with API-only models like GPT-4 returning no matches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
match_workload14 fields changed- changed
Input schema / properties / budget_usd_hr / descriptionPrevious value: -"Only recommend configs at or under this hourly budget."New value: +"Hard limit on the hourly price of the whole configuration, in USD. Configurations above it are not returned." - added
Input schema / properties / duration_hoursAdded value: +{ + "description": "How long the job will run, in hours (fractions are fine). When sent, each match carries billed_hours and estimated_total_usd.", + "type": "number" +} - added
Input schema / properties / max_quote_age_minutesAdded value: +{ + "description": "Hard limit on how old a price quote may be, in minutes. Configurations whose price was fetched from the provider longer ago are not returned.", + "type": "number" +} - added
Input schema / properties / max_total_usdAdded value: +{ + "description": "Hard limit on estimated_total_usd, the GPU rental cost of the whole job, in USD. Needs duration_hours. Configurations above it are not returned.", + "type": "number" +} - added
Output schema / properties / as_ofAdded value: +{ + "description": "When this answer was computed (ISO 8601). Quote ages are measured against it.", + "type": "string" +} - added
Output schema / properties / errorAdded value: +{ + "description": "Present when the request was not valid; says what to send instead.", + "type": "string" +} - added
Output schema / properties / limitsAdded value: +{ + "description": "The hard limits this answer applied (null when one was not sent) and how many configurations each one left out.", + "properties": { + "budget_usd_hr": { + "type": [ + "number", + "null" + ] + }, + "duration_hours": { + "type": [ + "number", + "null" + ] + }, + "excluded": { + "properties": { + "older_than_max_quote_age_minutes": { + "type": "integer" + }, + "over_budget_usd_hr": { + "type": "integer" + }, + "over_max_total_usd": { + "type": "integer" + } + }, + "type": "object" + }, + "max_quote_age_minutes": { + "type": [ + "number", + "null" + ] + }, + "max_total_usd": { + "type": [ + "number", + "null" + ] + }, + "nearest_excluded": { + "description": "When the limits leave no match: the cheapest configuration that can still run the job, and the limits it breaks.", + "properties": { + "billed_hours": { + "description": "Hours the provider bills for duration_hours: rounded up to whole hours on a per-hour meter and to whole minutes on a per-minute meter. Null when duration_hours was not sent.", + "type": [ + "number", + "null" + ] + }, + "billing_increment": { + "description": "How this provider's meter ticks for on-demand rental; varies means FastGPU has no verified figure.", + "enum": [ + "per-second", + "per-minute", + "per-hour", + "varies" + ], + "type": "string" + }, + "breaks": { + "items": { + "type": "string" + }, + "type": "array" + }, + "effective_usd_hr": { + "type": "number" + }, + "estimated_total_usd": { + "description": "effective_usd_hr times billed_hours, rounded up to the cent. GPU rental only: storage, data transfer, CPU and memory billed separately, and provider minimums are not included. Null when duration_hours was not sent.", + "type": [ + "number", + "null" + ] + }, + "fetched_at": { + "description": "When this price was fetched from the provider (ISO 8601).", + "type": [ + "string", + "null" + ] + }, + "fits_single_card": { + "type": "boolean" + }, + "gpu": { + "type": "string" + }, + "gpu_count": { + "type": "integer" + }, + "monthly_usd": { + "type": "number" + }, + "offer_type": { + "type": "string" + }, + "over_budget": { + "type": "boolean" + }, + "page_url": { + "description": "Absolute URL of this GPU's live price page, with every live offer.", + "type": "string" + }, + "partner": { + "type": "boolean" + }, + "provider": { + "type": "string" + }, + "provider_label": { + "type": [ + "string", + "null" + ] + }, + "quote_age_seconds": { + "description": "Seconds between fetched_at and as_of. Null when the fetch time is unknown.", + "type": [ + "integer", + "null" + ] + }, + "reason": { + "type": "string" + }, + "reliability": { + "type": "string" + }, + "score": { + "type": [ + "number", + "null" + ] + }, + "tokens_per_sec": { + "type": [ + "number", + "null" + ] + }, + "url": { + "description": "Site-relative path of this GPU's live price page, e.g. /gpus/h100-sxm.", + "type": "string" + }, + "usd_per_million_tokens": { + "type": [ + "number", + "null" + ] + } + }, + "type": [ + "object", + "null" + ] + } + }, + "type": "object" +} - added
Output schema / properties / matches / items / properties / billed_hoursAdded value: +{ + "description": "Hours the provider bills for duration_hours: rounded up to whole hours on a per-hour meter and to whole minutes on a per-minute meter. Null when duration_hours was not sent.", + "type": [ + "number", + "null" + ] +} - added
Output schema / properties / matches / items / properties / billing_incrementAdded value: +{ + "description": "How this provider's meter ticks for on-demand rental; varies means FastGPU has no verified figure.", + "enum": [ + "per-second", + "per-minute", + "per-hour", + "varies" + ], + "type": "string" +} - added
Output schema / properties / matches / items / properties / estimated_total_usdAdded value: +{ + "description": "effective_usd_hr times billed_hours, rounded up to the cent. GPU rental only: storage, data transfer, CPU and memory billed separately, and provider minimums are not included. Null when duration_hours was not sent.", + "type": [ + "number", + "null" + ] +} - added
Output schema / properties / matches / items / properties / fetched_atAdded value: +{ + "description": "When this price was fetched from the provider (ISO 8601).", + "type": [ + "string", + "null" + ] +} - added
Output schema / properties / matches / items / properties / quote_age_secondsAdded value: +{ + "description": "Seconds between fetched_at and as_of. Null when the fetch time is unknown.", + "type": [ + "integer", + "null" + ] +} - added
Output schema / properties / noteAdded value: +{ + "description": "Why there are no matches, when there are none.", + "type": "string" +} - added
Output schema / properties / workload / properties / gpu_vendorsAdded value: +{ + "items": { + "type": "string" + }, + "type": [ + "array", + "null" + ] +}
1 tool update
- Changed
match_workload3 fields changed- added
Input schema / properties / min_gpu_memory_gbAdded value: +{ + "description": "Memory each GPU card must have, in GB (e.g. 80 for 80GB-class cards such as the H100 or A100 80GB). Smaller cards are never returned. On its own it lists every card that size, cheapest first.", + "type": "integer" +} - changed
Input schema / properties / query / descriptionPrevious value: -"Plain-language job, e.g. \"cheapest to serve Llama 3 70B\" or \"2x H100 for fine-tuning\". Provide this OR a structured spec below."New value: +"Plain-language job, e.g. \"cheapest to serve Llama 3 70B\", \"2x H100 for fine-tuning\" or \"a GPU with 80GB of memory\". Provide this OR a structured spec below." - added
Output schema / properties / workload / properties / min_gpu_memory_gbAdded value: +{ + "type": [ + "integer", + "null" + ] +}
2 tool updates
- Changed
list_gpu_prices3 fields changed- added
Output schema / properties / gpus / items / properties / page_urlAdded value: +{ + "description": "Absolute URL of this GPU's live price page, with every live offer.", + "type": "string" +} - added
Output schema / properties / gpus / items / properties / slugAdded value: +{ + "type": "string" +} - added
Output schema / properties / gpus / items / properties / url / descriptionAdded value: +"Site-relative path of this GPU's live price page, e.g. /gpus/h100-sxm."
- Changed
match_workload2 fields changed- added
Output schema / properties / matches / items / properties / page_urlAdded value: +{ + "description": "Absolute URL of this GPU's live price page, with every live offer.", + "type": "string" +} - added
Output schema / properties / matches / items / properties / url / descriptionAdded value: +"Site-relative path of this GPU's live price page, e.g. /gpus/h100-sxm."
1 tool update
- Changed
list_gpu_prices1 field changed- added
Output schema / properties / gpus / items / properties / cheapest_billed_separatelyAdded value: +{ + "type": [ + "string", + "null" + ] +}
1 tool update
- Changed
list_gpu_prices1 field changed- added
Output schema / properties / gpus / items / properties / cheapest_min_gpu_countAdded value: +{ + "type": [ + "integer", + "null" + ] +}
1 tool update
- Changed
match_workload2 fields changed- changed
Input schema / properties / gpu_count / descriptionPrevious value: -"Force a specific GPU count instead of letting the engine size it."New value: +"Exact positive GPU count. Overrides a count in query text. Returns no matches if no supported configuration fits; omit for automatic sizing." - added
Output schema / properties / workload / properties / gpu_countAdded value: +{ + "type": [ + "integer", + "null" + ] +}
2 tool updates
- First observed
list_gpu_prices - First observed
match_workload
Related MCP Connectors
Hourly cloud GPU prices from 60+ providers: find, compare, cost and track GPUs.
Cloud GPU prices with live stock from 22 providers. Find where to rent an H100 or B300 now.
Live GPU rental market: 2,500+ offers across a dozen provider feeds. History, watches, limit orders.
Live GPU spot market: 1,700+ offers, 10 provider feeds. History, watches, limit orders, no fee
Related MCP Servers
- AlicenseBqualityCmaintenanceEnables users to describe their LLM fine-tuning job once and get the cheapest, fastest, and most balanced GPU options across a dozen cloud providers in seconds.7104MIT
- FlicenseNot gradedqualityDmaintenanceEnables comparing and renting vGPUs from 30+ cloud providers via Shadeform API, with tools to list, filter, rent, and manage GPU instances.-
- AlicenseAqualityBmaintenanceGlobal price benchmarking for AI inference across 2,600+ SKUs from 47 vendors. Query live pricing, market indexes, and model specs via 8 tools. Free tier available.855 npmMIT
- AlicenseNot gradedqualityDmaintenanceCompare AI inference pricing across 9 providers in real time. Routing recommendations, spend tracking, and budget alerts for AI agents.119 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.