Skip to main content
Glama

chimeraforge_plan

Recommend the best model, quantization, backend and GPU count for a workload, or explain why nothing fits, labeling every number's provenance.

Instructions

Recommend the best (model x quantization x backend x GPU-count) deployment for a workload, or report why nothing fits. Returns candidates with per-number provenance (measured/extrapolated/estimated/unknown). Use for: 'what GPU do I need for ', 'will fit on ', 'how many GPUs for N req/s', 'what will it cost'. Set platform (linux/windows/wsl2/macos) to the deployment OS: engines are offered only where their own docs say they run.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoonline
modelNo
hardwareYes
kv_quantNofp16
platformNo
workloadNosteady
lora_rankNo
duty_cycleNo
model_sizeNo3b
grid_regionNo
lora_targetNoqv
tpot_slo_msNo
ttft_slo_msNo
quality_fromNo
request_rateNo
allow_networkNo
allow_offloadNo
gpu_overridesNo
lora_adaptersNo
prompt_tokensNo
safety_targetNo
context_lengthNo
latency_slo_msNo
quality_targetNo
tensor_parallelNo
budget_usd_monthNo
reasoning_tokensNo
avg_output_tokensNo
pipeline_parallelNo
host_bandwidth_gbpsNo
gpu_price_multiplierNo
prefix_cache_hit_rateNo
max_num_batched_tokensNo
unified_memory_fractionNo
carbon_intensity_g_per_kwhNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changedv0.46.0
    • addedInput schema / properties / carbon_intensity_g_per_kwh
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "number"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Carbon Intensity G Per Kwh"
      +}
    • addedInput schema / properties / grid_region
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Grid Region"
      +}
    • addedInput schema / properties / latency_slo_ms / anyOf
      Added value: +[
      +  {
      +    "type": "number"
      +  },
      +  {
      +    "type": "null"
      +  }
      +]
    • changedInput schema / properties / latency_slo_ms / default
      Previous value: -5000New value: +null
    • removedInput schema / properties / latency_slo_ms / type
      Removed value: -"number"
    • addedInput schema / properties / mode
      Added value: +{
      +  "default": "online",
      +  "title": "Mode",
      +  "type": "string"
      +}
  2. Changed2 schema fields changedv0.43.0
    • addedInput schema / properties / platform
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Platform"
      +}
    • addedInput schema / properties / unified_memory_fraction
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "number"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Unified Memory Fraction"
      +}
  3. Changed3 schema fields changedv0.34.0
    • addedInput schema / properties / gpu_overrides
      Added value: +{
      +  "anyOf": [
      +    {
      +      "additionalProperties": true,
      +      "type": "object"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Gpu Overrides"
      +}
    • addedInput schema / properties / max_num_batched_tokens
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "integer"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Max Num Batched Tokens"
      +}
    • addedInput schema / properties / quality_from
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Quality From"
      +}
  4. Changed10 schema fields changedv0.30.0
    • addedInput schema / properties / allow_offload
      Added value: +{
      +  "default": false,
      +  "title": "Allow Offload",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / gpu_price_multiplier
      Added value: +{
      +  "default": 1,
      +  "title": "Gpu Price Multiplier",
      +  "type": "number"
      +}
    • addedInput schema / properties / host_bandwidth_gbps
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "number"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Host Bandwidth Gbps"
      +}
    • addedInput schema / properties / lora_adapters
      Added value: +{
      +  "default": 0,
      +  "title": "Lora Adapters",
      +  "type": "integer"
      +}
    • addedInput schema / properties / lora_rank
      Added value: +{
      +  "default": 16,
      +  "title": "Lora Rank",
      +  "type": "integer"
      +}
    • addedInput schema / properties / lora_target
      Added value: +{
      +  "default": "qv",
      +  "title": "Lora Target",
      +  "type": "string"
      +}
    • addedInput schema / properties / safety_target
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "number"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Safety Target"
      +}
    • addedInput schema / properties / tpot_slo_ms
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "number"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Tpot Slo Ms"
      +}
    • addedInput schema / properties / ttft_slo_ms
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "number"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Ttft Slo Ms"
      +}
    • addedInput schema / properties / workload
      Added value: +{
      +  "default": "steady",
      +  "title": "Workload",
      +  "type": "string"
      +}
  5. First observed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose two real behaviors: the return shape includes candidates with per-number provenance (measured/extrapolated/estimated/unknown) and platform gating ('engines are offered only where their own docs say they run'). It does not mention computational cost/latency of the planning call or determinism, but for an analytical/recommendation tool the failure mode ('why nothing fits') and output semantics are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, followed by return semantics and use cases; every sentence earns its place with no filler. The trailing platform sentence mixes a parameter hint with a behavioral constraint, slightly disrupting the flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 35 parameters, 0% schema coverage, and no output schema, the description is materially incomplete: only one parameter is explained and no guidance covers mode, workload, SLO, cost, or quality inputs. It handles purpose, use cases, and output provenance well but leaves the dominant complexity surface undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 35 parameters, so the description must compensate and largely does not. It meaningfully explains only one parameter (platform, with its linux/windows/wsl2/macos values); the other 34 — including opaque names like tpot_slo_ms, ttft_slo_ms, kval_quant, quality_from, lora_target — get no explanation, leaving an agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (recommend/report) and a precise resource (model x quantization x backend x GPU-count deployment) plus the negative outcome ('or report why nothing fits'). An agent can distinguish this planning tool from siblings like chimeraforge_list_hardware or chimeraforge_resolve_model without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives four concrete example intents ('what GPU do I need for <model>', 'will <model> fit on <gpu>', 'how many GPUs for N req/s', 'what will it cost') that clearly map to when to call it. It lacks explicit when-not guidance or named alternatives, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.