generate
Run a prompt on a GPU model tier: 'chat' (fast 7B, ~$0.0015), 'llama' (budget 8B, ~$0.0008), 'reason' (35B, ~$0.004), 'think' (deep chain-of-thought 35B MoE, ~$0.004), 'code' (coder model, ~$0.0025), or 'kaspa-expert' (RAG-grounded, current Kaspa knowledge, ~$0.0015). Returns the completion text.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | chat | |
| prompt | Yes | ||
| system | No | ||
| max_tokens | No |