What fits my GPU?
what_fitsCheck which open-weight LLMs fit your GPU's VRAM at a given quantization and context, returning recommended everyday, largest, Q8, and first out-of-reach options.
Instructions
Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | No | A GPU from list_gpus (id or name, e.g. "rtx-4090", "4090" or "M4 Max"). | |
| quant | No | Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default. | q4 |
| context | No | Tokens held in context: prompt plus conversation. | |
| vram_gb | No | Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true. | |
| apple_silicon | No | vram_gb is Apple unified memory; the GPU can use about 75% of it by default. | |
| bandwidth_gb_s | No | Memory bandwidth from the maker's spec, GB/s, for a speed ceiling. |