Skip to main content
Glama

Match a workload to the cheapest GPU

match_workload
Read-onlyIdempotent

Describe a job (an open-weight model to serve or fine-tune, a model size, or a GPU need) and get up to eight ranked GPU configurations that can run it at the lowest live price, each with the GPU count, the effective USD per hour for the whole configuration, and a one-line reason; the result also states the VRAM the job needs and, when a hyperscaler can run the same job, which one and what percent cheaper the top match is. Use it when the user asks where to run a workload or which GPU they need, rather than the price of a named GPU. It sizes open-weight models only: an API-only model such as GPT-4 returns no matches and a note. Results mirror FastGPU's website and apply a small, disclosed tie-break toward providers that pay FastGPU a referral, only between otherwise equal offers (each match reports partner true/false). Optional hard limits (budget_usd_hr, max_total_usd with duration_hours, max_quote_age_minutes) leave out every configuration that breaks them, and limits reports what was applied and how many were left out. page_url links to the GPU's live price page. It does not rent, reserve or launch GPUs. No key required.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
spotNoSet true to include interruptible spot capacity for a cheaper rate.
taskNoWhat the job does.
modelNoOpen model name to size against, e.g. "Llama 3 70B", "Qwen 72B", "Mixtral".
queryNoPlain-language job, e.g. "cheapest to serve Llama 3 70B", "2x H100 for fine-tuning" or "a GPU with 80GB of memory". Provide this OR a structured spec below.
regionNoRestrict to a data-residency region.
vram_gbNoRough VRAM the job needs, in GB, if you already know it.
params_bNoModel size in billions of parameters when no exact model is named.
reservedNoSet true to include reserved / committed-term capacity for a lower rate.
gpu_countNoExact positive GPU count. Overrides a count in query text. Returns no matches if no supported configuration fits; omit for automatic sizing.
precisionNoNumeric precision to size the model at.
budget_usd_hrNoHard limit on the hourly price of the whole configuration, in USD. Configurations above it are not returned.
max_total_usdNoHard limit on estimated_total_usd, the GPU rental cost of the whole job, in USD. Needs duration_hours. Configurations above it are not returned.
duration_hoursNoHow long the job will run, in hours (fractions are fine). When sent, each match carries billed_hours and estimated_total_usd.
min_gpu_memory_gbNoMemory each GPU card must have, in GB (e.g. 80 for 80GB-class cards such as the H100 or A100 80GB). Smaller cards are never returned. On its own it lists every card that size, cheapest first.
max_quote_age_minutesNoHard limit on how old a price quote may be, in minutes. Configurations whose price was fetched from the provider longer ago are not returned.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
heroNo
noteNoWhy there are no matches, when there are none.
as_ofNoWhen this answer was computed (ISO 8601). Quote ages are measured against it.
countYes
errorNoPresent when the request was not valid; says what to send instead.
staleNo
limitsNoThe hard limits this answer applied (null when one was not sent) and how many configurations each one left out.
matchesYes
workloadNo
updated_atNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed14 schema fields changed
    • changedInput schema / properties / budget_usd_hr / description
      Previous value: -"Only recommend configs at or under this hourly budget."New value: +"Hard limit on the hourly price of the whole configuration, in USD. Configurations above it are not returned."
    • addedInput schema / properties / duration_hours
      Added value: +{
      +  "description": "How long the job will run, in hours (fractions are fine). When sent, each match carries billed_hours and estimated_total_usd.",
      +  "type": "number"
      +}
    • addedInput schema / properties / max_quote_age_minutes
      Added value: +{
      +  "description": "Hard limit on how old a price quote may be, in minutes. Configurations whose price was fetched from the provider longer ago are not returned.",
      +  "type": "number"
      +}
    • addedInput schema / properties / max_total_usd
      Added value: +{
      +  "description": "Hard limit on estimated_total_usd, the GPU rental cost of the whole job, in USD. Needs duration_hours. Configurations above it are not returned.",
      +  "type": "number"
      +}
    • addedOutput schema / properties / as_of
      Added value: +{
      +  "description": "When this answer was computed (ISO 8601). Quote ages are measured against it.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / error
      Added value: +{
      +  "description": "Present when the request was not valid; says what to send instead.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / limits
      Added value: +{
      +  "description": "The hard limits this answer applied (null when one was not sent) and how many configurations each one left out.",
      +  "properties": {
      +    "budget_usd_hr": {
      +      "type": [
      +        "number",
      +        "null"
      +      ]
      +    },
      +    "duration_hours": {
      +      "type": [
      +        "number",
      +        "null"
      +      ]
      +    },
      +    "excluded": {
      +      "properties": {
      +        "older_than_max_quote_age_minutes": {
      +          "type": "integer"
      +        },
      +        "over_budget_usd_hr": {
      +          "type": "integer"
      +        },
      +        "over_max_total_usd": {
      +          "type": "integer"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "max_quote_age_minutes": {
      +      "type": [
      +        "number",
      +        "null"
      +      ]
      +    },
      +    "max_total_usd": {
      +      "type": [
      +        "number",
      +        "null"
      +      ]
      +    },
      +    "nearest_excluded": {
      +      "description": "When the limits leave no match: the cheapest configuration that can still run the job, and the limits it breaks.",
      +      "properties": {
      +        "billed_hours": {
      +          "description": "Hours the provider bills for duration_hours: rounded up to whole hours on a per-hour meter and to whole minutes on a per-minute meter. Null when duration_hours was not sent.",
      +          "type": [
      +            "number",
      +            "null"
      +          ]
      +        },
      +        "billing_increment": {
      +          "description": "How this provider's meter ticks for on-demand rental; varies means FastGPU has no verified figure.",
      +          "enum": [
      +            "per-second",
      +            "per-minute",
      +            "per-hour",
      +            "varies"
      +          ],
      +          "type": "string"
      +        },
      +        "breaks": {
      +          "items": {
      +            "type": "string"
      +          },
      +          "type": "array"
      +        },
      +        "effective_usd_hr": {
      +          "type": "number"
      +        },
      +        "estimated_total_usd": {
      +          "description": "effective_usd_hr times billed_hours, rounded up to the cent. GPU rental only: storage, data transfer, CPU and memory billed separately, and provider minimums are not included. Null when duration_hours was not sent.",
      +          "type": [
      +            "number",
      +            "null"
      +          ]
      +        },
      +        "fetched_at": {
      +          "description": "When this price was fetched from the provider (ISO 8601).",
      +          "type": [
      +            "string",
      +            "null"
      +          ]
      +        },
      +        "fits_single_card": {
      +          "type": "boolean"
      +        },
      +        "gpu": {
      +          "type": "string"
      +        },
      +        "gpu_count": {
      +          "type": "integer"
      +        },
      +        "monthly_usd": {
      +          "type": "number"
      +        },
      +        "offer_type": {
      +          "type": "string"
      +        },
      +        "over_budget": {
      +          "type": "boolean"
      +        },
      +        "page_url": {
      +          "description": "Absolute URL of this GPU's live price page, with every live offer.",
      +          "type": "string"
      +        },
      +        "partner": {
      +          "type": "boolean"
      +        },
      +        "provider": {
      +          "type": "string"
      +        },
      +        "provider_label": {
      +          "type": [
      +            "string",
      +            "null"
      +          ]
      +        },
      +        "quote_age_seconds": {
      +          "description": "Seconds between fetched_at and as_of. Null when the fetch time is unknown.",
      +          "type": [
      +            "integer",
      +            "null"
      +          ]
      +        },
      +        "reason": {
      +          "type": "string"
      +        },
      +        "reliability": {
      +          "type": "string"
      +        },
      +        "score": {
      +          "type": [
      +            "number",
      +            "null"
      +          ]
      +        },
      +        "tokens_per_sec": {
      +          "type": [
      +            "number",
      +            "null"
      +          ]
      +        },
      +        "url": {
      +          "description": "Site-relative path of this GPU's live price page, e.g. /gpus/h100-sxm.",
      +          "type": "string"
      +        },
      +        "usd_per_million_tokens": {
      +          "type": [
      +            "number",
      +            "null"
      +          ]
      +        }
      +      },
      +      "type": [
      +        "object",
      +        "null"
      +      ]
      +    }
      +  },
      +  "type": "object"
      +}
    • addedOutput schema / properties / matches / items / properties / billed_hours
      Added value: +{
      +  "description": "Hours the provider bills for duration_hours: rounded up to whole hours on a per-hour meter and to whole minutes on a per-minute meter. Null when duration_hours was not sent.",
      +  "type": [
      +    "number",
      +    "null"
      +  ]
      +}
    • addedOutput schema / properties / matches / items / properties / billing_increment
      Added value: +{
      +  "description": "How this provider's meter ticks for on-demand rental; varies means FastGPU has no verified figure.",
      +  "enum": [
      +    "per-second",
      +    "per-minute",
      +    "per-hour",
      +    "varies"
      +  ],
      +  "type": "string"
      +}
    • addedOutput schema / properties / matches / items / properties / estimated_total_usd
      Added value: +{
      +  "description": "effective_usd_hr times billed_hours, rounded up to the cent. GPU rental only: storage, data transfer, CPU and memory billed separately, and provider minimums are not included. Null when duration_hours was not sent.",
      +  "type": [
      +    "number",
      +    "null"
      +  ]
      +}
    • addedOutput schema / properties / matches / items / properties / fetched_at
      Added value: +{
      +  "description": "When this price was fetched from the provider (ISO 8601).",
      +  "type": [
      +    "string",
      +    "null"
      +  ]
      +}
    • addedOutput schema / properties / matches / items / properties / quote_age_seconds
      Added value: +{
      +  "description": "Seconds between fetched_at and as_of. Null when the fetch time is unknown.",
      +  "type": [
      +    "integer",
      +    "null"
      +  ]
      +}
    • addedOutput schema / properties / note
      Added value: +{
      +  "description": "Why there are no matches, when there are none.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / workload / properties / gpu_vendors
      Added value: +{
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": [
      +    "array",
      +    "null"
      +  ]
      +}
  2. Changed3 schema fields changed
    • addedInput schema / properties / min_gpu_memory_gb
      Added value: +{
      +  "description": "Memory each GPU card must have, in GB (e.g. 80 for 80GB-class cards such as the H100 or A100 80GB). Smaller cards are never returned. On its own it lists every card that size, cheapest first.",
      +  "type": "integer"
      +}
    • changedInput schema / properties / query / description
      Previous value: -"Plain-language job, e.g. \"cheapest to serve Llama 3 70B\" or \"2x H100 for fine-tuning\". Provide this OR a structured spec below."New value: +"Plain-language job, e.g. \"cheapest to serve Llama 3 70B\", \"2x H100 for fine-tuning\" or \"a GPU with 80GB of memory\". Provide this OR a structured spec below."
    • addedOutput schema / properties / workload / properties / min_gpu_memory_gb
      Added value: +{
      +  "type": [
      +    "integer",
      +    "null"
      +  ]
      +}
  3. Changed2 schema fields changed
    • addedOutput schema / properties / matches / items / properties / page_url
      Added value: +{
      +  "description": "Absolute URL of this GPU's live price page, with every live offer.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / matches / items / properties / url / description
      Added value: +"Site-relative path of this GPU's live price page, e.g. /gpus/h100-sxm."
  4. Changed2 schema fields changed
    • changedInput schema / properties / gpu_count / description
      Previous value: -"Force a specific GPU count instead of letting the engine size it."New value: +"Exact positive GPU count. Overrides a count in query text. Returns no matches if no supported configuration fits; omit for automatic sizing."
    • addedOutput schema / properties / workload / properties / gpu_count
      Added value: +{
      +  "type": [
      +    "integer",
      +    "null"
      +  ]
      +}
  5. First observed

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds real behavioral context beyond them: the disclosed referral tie-break between otherwise equal offers (with partner true/false per match), open-weight-only sizing, API-only models returning no matches plus a note, and that it does not rent/reserve/launch. It stops short of describing pagination or result ordering tie-break beyond price, keeping it at a strong 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and output shape are front-loaded in the first sentence, and each subsequent sentence carries distinct value (usage routing, scope limits, tie-break disclosure, hard limits, non-rental, no key). The opening sentence is very dense and could be split, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 15 optional parameters, an existing output schema, and a described output shape, the description covers purpose, usage routing, scope exclusions, cost/limit semantics, provider-bias disclosure, and auth needs (no key required). Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by grouping the hard-limit parameters (budget_usd_hr, max_total_usd with duration_hours, max_quote_age_minutes), explaining that any configuration breaking them is excluded, and that the 'limits' report shows what was applied and how many were dropped.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (describe a job) and resource (ranked GPU configurations) with concrete output detail: up to eight configurations, GPU count, effective USD/hr, and one-line reason. It explicitly distinguishes itself from the sibling by noting it is for 'where to run a workload or which GPU they need, rather than the price of a named GPU', which routes cleanly against list_gpu_prices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('user asks where to run a workload or which GPU they need') and when-not ('rather than the price of a named GPU'), effectively naming the alternative use case served by the sibling. It also gives a clear exclusion: open-weight models only, with API-only models like GPT-4 returning no matches.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources