vram-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OLLAMA_BASE_URL | No | Ollama endpoint URL | http://127.0.0.1:11434 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| vram_statusB | Report the selected GPU and server-wide Ollama model residency. Includes device identity, source health/timestamps/coverage, claims, recent GPU activity, and pressure (ok|tight|degraded|thrashing|unknown). Unavailable residency is null; unavailable telemetry never establishes an empty or healthy GPU. Non-local memory pressure is a best-effort paging heuristic. |
| list_loadedA | List the models currently resident in VRAM, with claim/busy detail. |
| unloadA | Evict a single model from VRAM now (Ollama Refuses by default if |
| ensure_freeA | Free VRAM until at least Also reports |
| warmA | Load a model or refresh its keep-alive (e.g. "5m"; "-1" pins indefinitely). An already resident model requires zero additional capacity. New loads account for reservations using an approximate model size; size_verified describes that estimate, while outcome describes verified residency. Returns succeeded|refused|failed|unknown. On unknown, inspect residency and pending_until before retrying. force=True bypasses admission checks but cannot override a pending operation. Zero durations must use unload(). |
| claimA | Declare that you're using Lets other sessions see who's using a model and why before deciding to
evict it. Renew before |
| reserveA | Reserve Use this for non-Ollama GPU work (a training run, a diffusion job) so other
sessions can see the VRAM is spoken for. COOPERATIVE: this gates vram-mcp's own |
| renewC | Extend an existing claim's expiry before it lapses. |
| releaseB | Release a claim early, before its TTL would expire. |
| list_claimsB | See who's claiming what right now (all models, or one). |
| adviseB | Suggest env/config changes to keep VRAM healthy (heuristics). |
| historyA | The VRAM audit trail, newest first: who ran unload/ensure_free/warm, and
which models/processes appeared or disappeared (with a best-effort cause).
Filter by |
| trendA | Free-VRAM trend over the last Answers "was this a gradual erosion or a sudden spike?" — the question a
point-in-time
|
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 13 tools
Each tool targets a distinct operation: claim/renew/release manage claims, reserve handles capacity reservations, unload/ensure_free/warm manage model residency, while vram_status/list_loaded/list_claims/history/trend provide distinct read views. Even similar actions like unload vs ensure_free are clearly separated by scope (single model vs threshold-based eviction).
Naming uses a mix of bare verbs (renew, release, advise, unload, warm, claim, reserve), verb_noun snake_case (list_claims, list_loaded, ensure_free), and noun-only identifiers (vram_status, history, trend). The pattern is somewhat predictable by category (actions vs queries), but it lacks a single consistent convention, making it less uniform than an all-verb_noun set.
13 tools is well within the ideal 3–15 range and each tool addresses a meaningful aspect of VRAM management: claims, reservations, eviction, warming, status, history, and trend. No tool feels redundant or superfluous.
The surface covers the core lifecycle well: claim/renew/release, reserve, warm/unload/ensure_free, plus status, history, and trend. A minor gap is the lack of a dedicated list_reservations tool—list_claims seems focused on model claims, so capacity reservations may be invisible unless they appear in status/history.