StudioTV LLM VRAM Calculator
Server Details
Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.
Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.
If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.
- Status
- Unhealthy
- Last Tested
- Transport
- Streamable HTTP
- URL
Related MCP Connectors
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.
Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.
LLM serving capacity planner: VRAM, KV cache, GPU topology, latency, and cost.
Will a local LLM run on your hardware? GGUF quant, buy-vs-rent-vs-API cost, used-GPU prices.
Related MCP Servers
- AlicenseAqualityAmaintenanceLLM deployment planner: given a model and a GPU, answers will it fit, will it hit your SLO, and what will it cost. Sizes VRAM and KV-cache from the model's real architecture, and labels every number measured, estimated, or unknown.51,476 PyPI2MIT
- AlicenseAqualityAmaintenanceChecks whether an open-weight LLM fits your GPU: VRAM for weights, KV cache and overhead at each quantisation, a tokens-per-second ceiling, the longest context that fits, and what would work instead when it doesn't. Reads any Hugging Face repo's config.json. Free, read-only, no API key.46MIT
- AlicenseAqualityBmaintenanceEnables AI agents to search, compare, and right-size Hugging Face models, including estimating whether a model fits in available GPU VRAM before downloading it.521 PyPIMIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents and MCP clients to estimate VRAM footprint and tensor parallelization requirements for quantized llama.cpp models, returning structured outputs without external dependencies.8-
Glama MCP Gateway
Add one secure layer between your agents and this server.