tf_harnesses
Returns a snapshot of the public Terminal-Bench leaderboards, revisions 4.0 and 2.1. Each row pairs one harness with one model and carries the upstream 95% confidence interval. Same model can score very differently on different harnesses; that gap is the value-add. The two revisions are NOT comparable: 4.0 is the current, harder board, 2.1 is the older and largely saturated one, and a pair in the high seventies on 2.1 can land in the twenties on 4.0. Pass ?view=summary for the current-board ranking plus biggest harness gaps; ?view=gaps for full per-model harness deltas; ?view=combined for the current board normalized to its top score; ?view=raw (default) for the full benchmark/result graph. Source: hand-curated from the upstream board at tbench.ai. Cache TTL 12h. Use when the agent needs to recommend a harness/model combo or explain why two agents using the same model perform differently.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | Output shape; default raw |