chimeraforge_plan
Recommend the best model, quantization, backend and GPU count for a workload, or explain why nothing fits, labeling every number's provenance.
Instructions
Recommend the best (model x quantization x backend x GPU-count) deployment for a workload, or report why nothing fits. Returns candidates with per-number provenance (measured/extrapolated/estimated/unknown). Use for: 'what GPU do I need for ', 'will fit on ', 'how many GPUs for N req/s', 'what will it cost'. Set platform (linux/windows/wsl2/macos) to the deployment OS: engines are offered only where their own docs say they run.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | online | |
| model | No | ||
| hardware | Yes | ||
| kv_quant | No | fp16 | |
| platform | No | ||
| workload | No | steady | |
| lora_rank | No | ||
| duty_cycle | No | ||
| model_size | No | 3b | |
| grid_region | No | ||
| lora_target | No | qv | |
| tpot_slo_ms | No | ||
| ttft_slo_ms | No | ||
| quality_from | No | ||
| request_rate | No | ||
| allow_network | No | ||
| allow_offload | No | ||
| gpu_overrides | No | ||
| lora_adapters | No | ||
| prompt_tokens | No | ||
| safety_target | No | ||
| context_length | No | ||
| latency_slo_ms | No | ||
| quality_target | No | ||
| tensor_parallel | No | ||
| budget_usd_month | No | ||
| reasoning_tokens | No | ||
| avg_output_tokens | No | ||
| pipeline_parallel | No | ||
| host_bandwidth_gbps | No | ||
| gpu_price_multiplier | No | ||
| prefix_cache_hit_rate | No | ||
| max_num_batched_tokens | No | ||
| unified_memory_fraction | No | ||
| carbon_intensity_g_per_kwh | No |