Skip to main content
Glama
alphaparkinc

genpark-quantized-model-vram-tensor-parallel-estimator-skill

Official

Related Servers

Alternatives to genpark-quantized-model-vram-tensor-parallel-estimator-skill

No user-submitted related servers found.

    Related Servers

    • A
      license
      A
      quality
      A
      maintenance
      LLM deployment planner: given a model and a GPU, answers will it fit, will it hit your SLO, and what will it cost. Sizes VRAM and KV-cache from the model's real architecture, and labels every number measured, estimated, or unknown.
      5
      2
      MIT
    • A
      license
      Not graded
      quality
      A
      maintenance
      Gives Claude Code, Cursor and other agents structural awareness of a PyTorch model: layers, params, FLOPs, blast radius, the design linter, a full readiness/cost/deployment verdict, and a ranker for which of k candidate designs to train. Reads a .py, a .neurarch.json, 81 bundled reference architectures, or a Hugging Face repo. Offline, no API key.
      96
      1
      MIT
    • A
      license
      Not graded
      quality
      A
      maintenance
      Vendor-neutral local LLM inference benchmark and hardware-config advisor for mlx and llama.cpp. Exposes an MCP tool that measures real tokens/second on your own hardware.
      10
      Apache 2.0

    Latest Blog Posts

    MCP directory API

    We provide all the information about MCP servers via our MCP API.

    curl -X GET 'https://glama.ai/api/mcp/v1/servers/alphaparkinc/genpark-quantized-model-vram-tensor-parallel-estimator-skill'

    If you have feedback or need assistance with the MCP directory API, please join our Discord server