start_run
Launch a background benchmark that asks a target model questions in English, Urdu, and Roman Urdu, grades each answer, and flags confidently wrong responses; returns a run_id unless wait=True.
Instructions
Ask the target model every question in each language, grade each answer, and record whether it was wrong while sounding sure. Runs in the background and returns a run_id unless wait=True.
Args: target: model to test, "provider:model" (e.g. "openai:gpt-4o"). judge: grader model, "provider:model". Defaults to env SBW_JUDGE; falls back to the target (not recommended). languages: subset of ["en", "ur", "roman_ur"]; default all three. question_ids: only these question ids. categories: only these categories. limit: only the first N questions (handy for a cheap trial run, e.g. 5). temperature: sampling temperature for the target (default 0; null to use the model default). system_prompt: optional system prompt for the target. Default none, like a plain chat. max_tokens: reply length cap for the target. concurrency: parallel questions in flight (lower it if you hit rate limits). questions_file: path to a custom question bank JSON. wait: block until finished and return the report (may hit client timeouts on big runs).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | ||
| judge | No | ||
| limit | No | ||
| target | Yes | ||
| languages | No | ||
| categories | No | ||
| max_tokens | No | ||
| concurrency | No | ||
| temperature | No | ||
| question_ids | No | ||
| system_prompt | No | ||
| questions_file | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |