Skip to main content
Glama
nodegrove

nodegrove VRAM: can I run it?

Official

Can I run it?

can_i_run
Read-onlyIdempotent

Check if your GPU can run an open-weight LLM: get the memory split, decode-speed ceiling, longest fitting context, and the changes that make it fit.

Instructions

Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
gpuNoA GPU from list_gpus (id or name, e.g. "rtx-4090", "4090" or "M4 Max").
modelNoA model from list_models (id or name, e.g. "llama-3.3-70b" or "Llama 3.3 70B"), or any Hugging Face repo id (e.g. "Qwen/Qwen3-8B"), read live from its config.json.
quantNoWeight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.q4
contextNoTokens held in context: prompt plus conversation.
vram_gbNoMemory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.
kv_cacheNoKV cache precision. fp16 is what most runtimes use; q8 halves the cache.fp16
architectureNoA model described by its config.json values instead of a name.
apple_siliconNovram_gb is Apple unified memory; the GPU can use about 75% of it by default.
bandwidth_gb_sNoMemory bandwidth from the maker's spec, GB/s, for a speed ceiling.
active_params_bNoParameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (read-only, idempotent, non-destructive, open-world), and with no output schema the description carries the return-value burden itself: it names the verdict values, the memory breakdown, the decode-speed ceiling, the longest fitting context, and the remediation list returned on a 'no'. That is substantial behavioral disclosure well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose sentence, then the return contents, then the accepted input forms for model and GPU. Dense but no sentence is filler; the return enumeration is longer than strictly necessary but earns its place given there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with nested architecture objects and no output schema, the description covers both the answer shape and the input domains, which is what an agent needs to call it correctly. Only minor gaps remain, such as what the speed ceiling is expressed in or how partial/unknown inputs are treated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (quant, context, kv_cache, architecture sub-fields) is already documented at the schema level. The description restates the accepted input forms for model and gpu but adds no syntax, units, or format detail the schema does not already supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific question the tool answers ('can this GPU run this open-weight LLM') and enumerates the verdicts (fits, tight, no) plus the memory split and speed ceiling it computes. That is enough for an agent to distinguish it from estimate_vram and what_fits without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear input-selection context: model can come from list_models or be any Hugging Face repo id or raw architecture, and GPU can come from list_gpus or a manual vram_gb. It does not explicitly say when to prefer this over siblings like estimate_vram or what_fits, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.