Skip to main content
Glama
alphaparkinc

genpark-paged-attention-kv-cache-budget-calculator-skill

Related Servers

Alternatives to genpark-paged-attention-kv-cache-budget-calculator-skill

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables MCP clients to manage LLM inference memory through virtual paged-attention KV-cache block mapping, non-contiguous physical page allocation, zero-copy fragmentation tracking, radix-trie prefix caching, INT8 quantized compute, and speculative decoding verification. It also exposes prefill and decode latency telemetry so edge deployments can be benchmarked and tuned without external dependencies.
      7
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables verification of draft-target speculative decoding with rejection sampling, acceptance-rate telemetry, and dynamic speedup estimation, alongside edge inference primitives such as INT8 quantization, paged KV allocation, and radix prefix caching.
      7
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables inference engines to match longest common prompt token prefixes via a radix trie, so cached KV-cache blocks are reused instead of re-prefilled. This eliminates redundant prefill computation and lowers time-to-first-token, alongside paged attention allocation, speculative decoding verification, INT8 quantization, and latency telemetry.
      7
      MIT