A
licenseNot graded
qualityB
maintenanceEnables profiling and optimizing edge and on-device LLM inference by measuring TTFT, TPOT, tokens/second throughput, and P50/P90/P99 jitter while providing standard-library primitives for quantization, paged attention, speculative decoding, and prefix caching.
7
MIT