chimeraforge
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| chimeraforge_planA | Recommend the best (model x quantization x backend x GPU-count) deployment for a workload, or report why nothing fits. Returns candidates with per-number provenance (measured/extrapolated/estimated/unknown). Use for: 'what GPU do I need for ', 'will fit on ', 'how many GPUs for N req/s', 'what will it cost'. Set platform (linux/windows/wsl2/macos) to the deployment OS: engines are offered only where their own docs say they run. |
| chimeraforge_resolve_modelB | Resolve a model id to real params/architecture (grounds hallucinated specs). |
| chimeraforge_list_hardwareA | List known GPUs with VRAM/bandwidth/TDP/interconnect. |
| chimeraforge_compare_apiA | Compare self-hosting against the hosted APIs for a workload: sizes the cheapest feasible GPU fleet, prices the same traffic through each API model, and gives the monthly output-token volume where the two break even. Use for 'is it cheaper to self-host or use the API', 'when does a GPU pay for itself'. API prices come from a dated snapshot -- the result reports its age and flags it when stale; say so rather than quoting an old price as current. |
| chimeraforge_suggestA | Rank the models that actually fit and hit the SLO on a given GPU -- the inverse of planning. Use for 'what can I run on a 4090', 'best model for 12GB'. Sources: catalog (offline curated set), ollama (locally installed), hf (top Hub repos). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool targets a distinct resource or action: planning, hardware listing, API comparison, model resolution, and model suggestion. The only potential confusion is between plan (workload→GPU) and suggest (GPU→models), which are inverse operations, but descriptions explicitly clarify the distinction.
All names use the consistent 'chimeraforge_' prefix and snake_case. Most follow a verb_noun pattern (list_hardware, compare_api, resolve_model), with two single-verb exceptions (plan, suggest), which is a minor deviation.
Five tools is well-scoped for a GPU deployment planning server. Each tool earns its place: planning, hardware listing, API comparison, model resolution, and inverse suggestion cover the core workflow without redundancy.
The surface covers the main deployment planning lifecycle: resolving models, listing hardware, planning deployments, comparing costs, and suggesting models. Minor gaps exist, such as no explicit tool to enumerate supported backends or engines, but agents can work around this via the platform parameter in plan.