Can I run it?
can_i_runCheck if your GPU can run an open-weight LLM: get the memory split, decode-speed ceiling, longest fitting context, and the changes that make it fit.
Instructions
Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | No | A GPU from list_gpus (id or name, e.g. "rtx-4090", "4090" or "M4 Max"). | |
| model | No | A model from list_models (id or name, e.g. "llama-3.3-70b" or "Llama 3.3 70B"), or any Hugging Face repo id (e.g. "Qwen/Qwen3-8B"), read live from its config.json. | |
| quant | No | Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default. | q4 |
| context | No | Tokens held in context: prompt plus conversation. | |
| vram_gb | No | Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true. | |
| kv_cache | No | KV cache precision. fp16 is what most runtimes use; q8 halves the cache. | fp16 |
| architecture | No | A model described by its config.json values instead of a name. | |
| apple_silicon | No | vram_gb is Apple unified memory; the GPU can use about 75% of it by default. | |
| bandwidth_gb_s | No | Memory bandwidth from the maker's spec, GB/s, for a speed ceiling. | |
| active_params_b | No | Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling. |