Skip to main content
Glama

Self-host an open model or use the API?

self_host_or_api
Read-onlyIdempotent

Is it cheaper to host an open-weights model yourself on rented GPUs, or to use the cheapest API offer? Returns a one-sentence verdict (headline) and the numbers behind it: the break-even volume in million tokens per day, the throughput the whole configuration must deliver for self-hosting to cost less (required_tokens_per_second: compare it with what your setup does), the break-even utilization when a throughput is known, the GPU configuration and its hourly price, and what the verdict rests on (confidence). Without a volume you still get the verdict and the break-even volume. It is an estimate: the throughput is a published measurement (measured), an estimate scaled from one (derived; estimated for a wide range), or a published minimum (lower_bound: a floor, valid for requests of up to 2,048 tokens in total, that can show self-hosting wins, never that the API is cheaper); or you give your own tokens_per_second. When no published figure decides, the headline gives the break-even volume and the required throughput, to compare with yours, rather than a yes or no, and may point out a dearer setup that published figures do decide (settled_alternative: for comparison, not a recommendation). headline_kind and headline_values give the headline as fields. It declines to give a number (status refused, with a reason) when it cannot do so reliably, for example a closed model, weights that do not fit the requested GPU, or a GPU nobody rents. Hardware rental only: the hourly price the provider publishes for the machine. It does not count the engineering to deploy and run the model, monitoring, redundancy, start-up time, or storage and network costs billed separately.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
gpuNoImpose a GPU: id from list_gpus, for example h100-sxm-80gb
tierNoGPU rental tier (default guaranteed)
modelYesModel id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure.
detailNocompact (default): verdict, headline and key numbers. full: also the three throughput scenarios, the model and every assumption
gpu_countNoImpose a number of GPUs: 1, 2, 4, 8 or 16
utilizationNoAverage share of time the GPUs serve requests, in percent (default 50)
quantizationNoPrecision of the weights (default fp8); must be published by at least one provider unless fp16 or bf16
output_tokensNoOutput tokens per request
prompt_tokensNoInput tokens per request
tokens_per_dayNoYour volume in tokens per day (input and output together). Or give requests_per_day, prompt_tokens and output_tokens
weights_marginNoPercent of the GPU memory the model weights may fill (default 70; the rest is the engine reserve and the context cache)
requests_per_dayNoRequests per day (with prompt_tokens and output_tokens)
tokens_per_secondNoYour own measured throughput for the whole configuration, input and output tokens together, instead of our estimate

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the read-only/idempotent safety profile, and the description adds substantial context beyond that: measured vs derived vs lower_bound throughput semantics (including the 2,048-token floor for lower_bound), refusal behavior and its causes, the settled_alternative field being for comparison not recommendation, and an explicit list of what costs are excluded (engineering, monitoring, redundancy, storage/network).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core question and the headline verdict, and every clause carries information. It is dense with nested parentheses in a single block, which hurts scanability, but there is little pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 13-parameter analytical tool with no output schema, the description fully specifies the return payload (headline, headline_kind/values, break-even volume, required throughput, confidence), the refusal path, and the modeling limitations, so an agent has what it needs to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, setting the baseline at 3, but the description adds real meaning: it distinguishes tokens_per_day from the requests_per_day + prompt/output token path, explains tokens_per_second as a user-supplied override of the estimate, and clarifies the compact/full detail trade-off.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise decision question (self-host an open-weights model vs cheapest API) and enumerates exactly what is returned (verdict/headline, break-even volume, required throughput, GPU config, confidence). This is clearly distinguishable from siblings like estimate_cost and cheapest, which price rather than compare hosting strategies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the operational conditions well: it works with or without a volume, tokens_per_second overrides the estimate, and it declines with 'status refused' for closed models or unfit GPUs. It does not, however, explicitly route the agent away from or toward sibling tools (e.g., when to prefer estimate_cost or gpu_rentals).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.