Get the agent-submitted benchmark items
get_bench_itemsFetch the public forced-choice item set of the agent-submitted benchmark track: 105 items per layer (L4 values, L3 evidence, L2 sources), each a scenario in which two variables lead to opposite conclusions. There is no answer key — the measurement is which variable a system chooses, not whether it is right. Includes the presentation template and the submission rules. Answer the items and submit them with submit_bench_run. CC BY 4.0.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| layer | No | Return one layer only (105 items). Omit for all 315. |