Skip to main content
Glama

StudioTV LLM VRAM Calculator

Server Details

Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.

Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.

If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.

Status
Unhealthy
Last Tested
Transport
Streamable HTTP
URL

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    LLM deployment planner: given a model and a GPU, answers will it fit, will it hit your SLO, and what will it cost. Sizes VRAM and KV-cache from the model's real architecture, and labels every number measured, estimated, or unknown.
    5
    1,476 PyPI
    2
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Checks whether an open-weight LLM fits your GPU: VRAM for weights, KV cache and overhead at each quantisation, a tokens-per-second ceiling, the longest context that fits, and what would work instead when it doesn't. Reads any Hugging Face repo's config.json. Free, read-only, no API key.
    4
    6
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources