Skip to main content
Glama
alphaparkinc

genpark-edge-inference-latency-telemetry-skill

Official

Related Servers

Alternatives to genpark-edge-inference-latency-telemetry-skill

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables profiling and optimizing edge and on-device LLM inference by measuring TTFT, TPOT, tokens/second throughput, and P50/P90/P99 jitter while providing standard-library primitives for quantization, paged attention, speculative decoding, and prefix caching.
      7
      MIT
    • A
      license
      Not graded
      quality
      C
      maintenance
      Enables benchmarking of local LLM models (performance and quality) and sharing results to a public leaderboard via MCP tools.
      16 npm
      8
      Apache 2.0
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables MCP clients to manage LLM inference memory through virtual paged-attention KV-cache block mapping, non-contiguous physical page allocation, zero-copy fragmentation tracking, radix-trie prefix caching, INT8 quantized compute, and speculative decoding verification. It also exposes prefill and decode latency telemetry so edge deployments can be benchmarked and tuned without external dependencies.
      7
      MIT
    • A
      license
      Not graded
      quality
      A
      maintenance
      Vendor-neutral local LLM inference benchmark and hardware-config advisor for mlx and llama.cpp. Exposes an MCP tool that measures real tokens/second on your own hardware.
      10 npm
      Apache 2.0