benchmark_inference
Run in-memory benchmarks to measure model speed (ms/batch), throughput, and VRAM footprint. Identify performance bottlenecks for inference optimization.
Instructions
Runs synthetic in-memory benchmark to test model speed (ms/batch), throughput, and VRAM footprint.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| seq_len | No | ||
| batch_size | No | ||
| iterations | No |