A
licenseA
qualityA
maintenanceChecks whether an open-weight LLM fits your GPU: VRAM for weights, KV cache and overhead at each quantisation, a tokens-per-second ceiling, the longest context that fits, and what would work instead when it doesn't. Reads any Hugging Face repo's config.json. Free, read-only, no API key.
4
6
MIT