Train Dpo
train_dpoFine-tune a model on chosen/rejected pairs using DPO to optimize for human preferences.
Instructions
Train on chosen/rejected pairs using Cookbook DPO. Spends credits.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | ||
| background | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||