Estimate VRAM from a Hugging Face repo
estimate_from_hf_repoReads a Hugging Face repo's config.json to estimate VRAM for weights, KV cache, and overhead across quantisations and context lengths.
Instructions
Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Hugging Face repo id, e.g. "Qwen/Qwen3-8B", or its huggingface.co URL. | |
| context | No | Tokens held in context: prompt plus conversation. | |
| kv_cache | No | KV cache precision. fp16 is what most runtimes use; q8 halves the cache. | fp16 |
| active_params_b | No | Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling. |