Self-host an open model or use the API?
self_host_or_apiIs it cheaper to host an open-weights model yourself on rented GPUs, or to use the cheapest API offer? Returns a one-sentence verdict (headline) and the numbers behind it: the break-even volume in million tokens per day, the throughput the whole configuration must deliver for self-hosting to cost less (required_tokens_per_second: compare it with what your setup does), the break-even utilization when a throughput is known, the GPU configuration and its hourly price, and what the verdict rests on (confidence). Without a volume you still get the verdict and the break-even volume. It is an estimate: the throughput is a published measurement (measured), an estimate scaled from one (derived; estimated for a wide range), or a published minimum (lower_bound: a floor, valid for requests of up to 2,048 tokens in total, that can show self-hosting wins, never that the API is cheaper); or you give your own tokens_per_second. When no published figure decides, the headline gives the break-even volume and the required throughput, to compare with yours, rather than a yes or no, and may point out a dearer setup that published figures do decide (settled_alternative: for comparison, not a recommendation). headline_kind and headline_values give the headline as fields. It declines to give a number (status refused, with a reason) when it cannot do so reliably, for example a closed model, weights that do not fit the requested GPU, or a GPU nobody rents. Hardware rental only: the hourly price the provider publishes for the machine. It does not count the engineering to deploy and run the model, monitoring, redundancy, start-up time, or storage and network costs billed separately.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | No | Impose a GPU: id from list_gpus, for example h100-sxm-80gb | |
| tier | No | GPU rental tier (default guaranteed) | |
| model | Yes | Model id or name, e.g. 'deepseek-v3.2', 'deepseek/deepseek-v4-pro', 'gpt-5.6-luna'. Use search_models when unsure. | |
| detail | No | compact (default): verdict, headline and key numbers. full: also the three throughput scenarios, the model and every assumption | |
| gpu_count | No | Impose a number of GPUs: 1, 2, 4, 8 or 16 | |
| utilization | No | Average share of time the GPUs serve requests, in percent (default 50) | |
| quantization | No | Precision of the weights (default fp8); must be published by at least one provider unless fp16 or bf16 | |
| output_tokens | No | Output tokens per request | |
| prompt_tokens | No | Input tokens per request | |
| tokens_per_day | No | Your volume in tokens per day (input and output together). Or give requests_per_day, prompt_tokens and output_tokens | |
| weights_margin | No | Percent of the GPU memory the model weights may fill (default 70; the rest is the engine reserve and the context cache) | |
| requests_per_day | No | Requests per day (with prompt_tokens and output_tokens) | |
| tokens_per_second | No | Your own measured throughput for the whole configuration, input and output tokens together, instead of our estimate |