compare_runs
Compare multiple benchmark runs side by side to see confidently-wrong and accuracy percentages per language, helping evaluate model performance across English, Urdu, and Roman Urdu.
Instructions
Side-by-side confidently-wrong % and accuracy % per language for several runs (e.g. different models).
Args: run_ids: the runs to compare.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| run_ids | Yes |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |