compare_co_scientist_workflows
Compare benchmark results from single-agent, multi-agent, and co-scientist workflows to evaluate quality and traceability, and decide whether the simpler workflow should remain the default.
Instructions
Compare benchmark results across workflow types.
AUTOMATIC TRIGGERS - Call this when:
You have benchmark results for single_agent, one_session_multi_agent, or two_session_co_scientist
Deciding whether the simpler workflow should remain the default
Comparing quality and traceability across research workflows
PARAMETERS:
results: List of workflow result dicts or pre-evaluated result dicts
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| results | Yes |