Skip to main content
Glama
Alpha-Park

genpark-edge-inference-latency-telemetry-skill

by Alpha-Park

Related Servers

Alternatives to genpark-edge-inference-latency-telemetry-skill

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables zero-dependency INT8 symmetric quantization of LLM weights and activations on a per-tensor or per-channel basis, reporting SNR and MSE reconstruction telemetry along with TTFT/TPOT latency and jitter metrics. Also exposes edge inference primitives such as PagedAttention block allocation, radix prefix caching, and speculative decoding verification through the Model Context Protocol.
      7
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables MCP clients to manage LLM inference memory through virtual paged-attention KV-cache block mapping, non-contiguous physical page allocation, zero-copy fragmentation tracking, radix-trie prefix caching, INT8 quantized compute, and speculative decoding verification. It also exposes prefill and decode latency telemetry so edge deployments can be benchmarked and tuned without external dependencies.
      7
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables verification of draft-target speculative decoding with rejection sampling, acceptance-rate telemetry, and dynamic speedup estimation, alongside edge inference primitives such as INT8 quantization, paged KV allocation, and radix prefix caching.
      7
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Provides an MCP-compatible profiler for measuring edge and on-device LLM inference latency, including TTFT, TPOT, tokens-per-second throughput, and P50/P90/P99 jitter distributions. It enables agents and developers to run these telemetry benchmarks from MCP clients such as Claude Desktop, Cursor, and Windsurf.
      7
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables inference engines to match longest common prompt token prefixes via a radix trie, so cached KV-cache blocks are reused instead of re-prefilled. This eliminates redundant prefill computation and lowers time-to-first-token, alongside paged attention allocation, speculative decoding verification, INT8 quantization, and latency telemetry.
      7
      MIT
    • A
      license
      Not graded
      quality
      A
      maintenance
      Vendor-neutral local LLM inference benchmark and hardware-config advisor for mlx and llama.cpp. Exposes an MCP tool that measures real tokens/second on your own hardware.
      10 npm
      Apache 2.0