get_report
Generate a Markdown report for a benchmark run, covering per-language confidently-wrong rates, accuracy, confidence gaps, hedge-word checks, and example answers.
Instructions
Markdown report: confidently-wrong rate per language, accuracy, confidence when right vs wrong, hedge-word cross-check, language gaps, per-question grid and example answers.
Args: run_id: the run to report on. examples: how many confidently-wrong answers to quote. per_question: include the per-question grid.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| examples | No | ||
| per_question | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |