Skip to main content
Glama
nodegrove

nodegrove VRAM: can I run it?

Official

Estimate VRAM

estimate_vram
Read-onlyIdempotent

Calculate VRAM needed to run an LLM: weights, KV cache and overhead per quantisation at any context, plus the smallest GPU that fits it.

Instructions

How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoA model from list_models (id or name, e.g. "llama-3.3-70b" or "Llama 3.3 70B"), or any Hugging Face repo id (e.g. "Qwen/Qwen3-8B"), read live from its config.json.
quantNoWeight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6.
contextNoTokens held in context: prompt plus conversation.
kv_cacheNoKV cache precision. fp16 is what most runtimes use; q8 halves the cache.fp16
architectureNoA model described by its config.json values instead of a name.
active_params_bNoParameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, so the bar is lower, and the description adds genuine substance: it breaks the estimate into weights + KV cache + overhead, notes per-quantisation and per-context variation, and discloses the derived card-class output. It stops short of noting that model lookups hit Hugging Face live (rate limits/latency), which the schema mentions but the description does not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with what is computed before the input forms. Nothing is padded, though the second sentence packs three alternative input modes into one clause and could be slightly clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter tool with a nested object and no output schema, the description covers the computation, the quantisation/context levers, and the headline output (smallest common card class). It does not describe the full return shape (a per-quantisation memory list) or flag live HF fetches, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all six parameters, including enums for quant and kv_cache and the nested architecture object. The description only restates the model-source options and names four architecture fields, adding little beyond the structured data, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computation (weights + KV cache + overhead) plus the derived output (smallest common card class), which is far more than a restated title. However, it never distinguishes itself from the sibling estimate_from_hf_repo, which it partly overlaps with by accepting HF repo ids, so an agent still has to guess which estimator to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists accepted model-reference forms but gives no when-to-use guidance and names no alternatives. With siblings can_i_run, what_fits, and estimate_from_hf_repo in the same family, the agent gets no signal about when this estimator is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.