Create a dedicated endpoint
together_create_endpointDeploy a model on dedicated GPUs. The endpoint STARTS AUTOMATICALLY and bills per minute of uptime until stopped — set inactive_timeout to auto-stop it, and use together_list_hardware for valid hardware ids. Together: POST /endpoints.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | The model to deploy. | |
| state | No | Initial state. Pass STOPPED to create without starting (and without billing). | |
| hardware | Yes | Hardware id, e.g. 1x_nvidia_a100_80gb_sxm. | |
| autoscaling | Yes | Replica bounds for autoscaling. | |
| display_name | No | Human-readable name. | |
| inactive_timeout | No | Minutes of inactivity before auto-stop; 0 disables it. | |
| availability_zone | No | Availability zone, e.g. us-central-4b. | |
| disable_speculative_decoding | No |