screen my traffic
screen_my_trafficCompare cheaper or newer models against your stored answers on real traffic, get a win-rate verdict and switch/keep recommendation. Zero configuration, but wallet billing applies.
Instructions
Would a cheaper (or newer) model hold on this workspace's own traffic? Starts a zero-config screening — the dominant logged model is the incumbent, its STORED answers the baseline, the cheaper model of each family (or the candidates you pass) the challengers — waits for it, and returns each candidate's verdict from the win-rate interval plus a switch/keep recommendation. SPENDS MONEY: judging and candidate generations bill the workspace wallet (402 when the wallet cannot cover the funds gate). Needs request logging on and logged traffic. Prefer this over create_eval for the 'is X better/cheaper' question.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Only traffic logged with this tag (one tag = one population). | |
| sample_count | No | Prompts to sample (5..500). Default: the server's screening default. | |
| wait_seconds | No | How long to wait for the run before returning its id to poll. Default 600; 0 returns immediately. | |
| candidate_models | No | Catalog model ids to test instead of the auto-picked cheaper set (max 6). An upgrade counts — anything in the catalog. |