Estimate VRAM
estimate_vramCalculate VRAM needed to run an LLM: weights, KV cache and overhead per quantisation at any context, plus the smallest GPU that fits it.
Instructions
How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | A model from list_models (id or name, e.g. "llama-3.3-70b" or "Llama 3.3 70B"), or any Hugging Face repo id (e.g. "Qwen/Qwen3-8B"), read live from its config.json. | |
| quant | No | Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6. | |
| context | No | Tokens held in context: prompt plus conversation. | |
| kv_cache | No | KV cache precision. fp16 is what most runtimes use; q8 halves the cache. | fp16 |
| architecture | No | A model described by its config.json values instead of a name. | |
| active_params_b | No | Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling. |