TokCalc MCP Server
Server Details
LLM capacity planning tools to estimate VRAM, compute, latency, and GPU topology for agents.
Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.
If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.
- Status
- Unhealthy
- Last Tested
- Transport
- Streamable HTTP
- URL
- Repository
- stevecrates489-commits/tokcalc
- GitHub Stars
- 1
Related MCP Connectors
GPU and LLM inference benchmarks, hardware evidence, deployment recommendations, and launch configs.
Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.
Which LLMs actually run on your GPU, and how fast. Mixture-of-experts included.
Cost-optimized LLM model routing recommendations for autonomous AI agents
Related MCP Servers
- AlicenseAqualityAmaintenanceLLM deployment planner: given a model and a GPU, answers will it fit, will it hit your SLO, and what will it cost. Sizes VRAM and KV-cache from the model's real architecture, and labels every number measured, estimated, or unknown.5277 PyPI2MIT
- AlicenseAqualityDmaintenanceEstimates GPU requirements, training/inference costs, and cloud-vs-on-prem TCO for AI workloads using deterministic calculators.121MIT
- AlicenseAqualityAmaintenanceReal-time cluster health monitoring, pre-request NLMS latency prediction, and intelligent prompt routing across multi-instance LLM backends (vLLM, Ollama, SGLang, TGI).443Apache 2.0
- AlicenseAqualityCmaintenanceEnables AI agents to search, compare, and right-size Hugging Face models, including estimating whether a model fits in available GPU VRAM before downloading it.5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.