chimeraforge_suggest
Rank models that fit your GPU and meet latency SLOs. Find answers to 'what can I run on this hardware' using curated, local, and Hugging Face sources.
Instructions
Rank the models that actually fit and hit the SLO on a given GPU -- the inverse of planning. Use for 'what can I run on a 4090', 'best model for 12GB'. Sources: catalog (offline curated set), ollama (locally installed), hf (top Hub repos).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| source | No | catalog | |
| hardware | Yes | ||
| hf_limit | No | ||
| ollama_url | No | ||
| request_rate | No | ||
| context_length | No | ||
| latency_slo_ms | No | ||
| quality_target | No | ||
| budget_usd_month | No | ||
| avg_output_tokens | No |