A
licenseNot graded
qualityB
maintenanceEnables zero-dependency INT8 symmetric quantization of LLM weights and activations on a per-tensor or per-channel basis, reporting SNR and MSE reconstruction telemetry along with TTFT/TPOT latency and jitter metrics. Also exposes edge inference primitives such as PagedAttention block allocation, radix prefix caching, and speculative decoding verification through the Model Context Protocol.
7
MIT