request_metrics
Retrieve vLLM inference metrics: time to first token, time per output token, end-to-end latency, and total generation tokens to monitor and analyze request performance.
Instructions
[READ] vLLM TTFT / TPOT / e2e latency + generation-token totals.
Args: target: Inference target name from config; omit for the default.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| target | No |