Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
OLLAMA_BASE_URLNoOllama endpoint URLhttp://127.0.0.1:11434

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
vram_statusB

Report the selected GPU and server-wide Ollama model residency.

Includes device identity, source health/timestamps/coverage, claims, recent GPU activity, and pressure (ok|tight|degraded|thrashing|unknown). Unavailable residency is null; unavailable telemetry never establishes an empty or healthy GPU. Non-local memory pressure is a best-effort paging heuristic.

list_loadedA

List the models currently resident in VRAM, with claim/busy detail.

unloadA

Evict a single model from VRAM now (Ollama keep_alive=0).

Refuses by default if model has an active claim or a best-effort busy signal — pass force=True to override (busy is windowed and can lag a few seconds past a generation). by records who requested the eviction in the audit log (see history).

ensure_freeA

Free VRAM until at least gb GB is available. Skips claimed/busy models unless force=True. by records the requester in the audit log.

Also reports reserved_mb — how much of the resulting free VRAM other sessions have reserved for non-Ollama work (None if unreadable).

warmA

Load a model or refresh its keep-alive (e.g. "5m"; "-1" pins indefinitely).

An already resident model requires zero additional capacity. New loads account for reservations using an approximate model size; size_verified describes that estimate, while outcome describes verified residency. Returns succeeded|refused|failed|unknown. On unknown, inspect residency and pending_until before retrying. force=True bypasses admission checks but cannot override a pending operation. Zero durations must use unload().

claimA

Declare that you're using model for purpose.

Lets other sessions see who's using a model and why before deciding to evict it. Renew before ttl_seconds elapses if still in use — an un-renewed claim simply expires, so a crashed session never leaves a permanently-stuck claim.

reserveA

Reserve gb GB of VRAM — a claim on capacity, not on a named model.

Use this for non-Ollama GPU work (a training run, a diffusion job) so other sessions can see the VRAM is spoken for. pid is advisory. Reservations expire by TTL like claims, so a crashed session never leaves one stuck.

COOPERATIVE: this gates vram-mcp's own warm(), but vram-mcp cannot intercept an Ollama auto-load triggered by a direct /api/generate call from another process.

renewC

Extend an existing claim's expiry before it lapses.

releaseB

Release a claim early, before its TTL would expire.

list_claimsB

See who's claiming what right now (all models, or one).

adviseB

Suggest env/config changes to keep VRAM healthy (heuristics).

historyA

The VRAM audit trail, newest first: who ran unload/ensure_free/warm, and which models/processes appeared or disappeared (with a best-effort cause). Filter by model, type (action|disappeared|appeared), limit, or an ISO since floor. Answers 'what happened to model X?'.

trendA

Free-VRAM trend over the last hours, from the sampled audit log.

Answers "was this a gradual erosion or a sudden spike?" — the question a point-in-time vram_status() cannot. Returns direction, min/max/latest free MB (latest is null when the newest sample carries no reading), how many samples showed driver spill, and the raw samples.

samples holds at most the 200 most recent rows so a long window can't flood the caller's context; samples_truncated says whether older rows were dropped. Every summary figure is computed over the FULL window either way.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.5/5.0

Scored across 13 tools

Disambiguation5/5

Each tool targets a distinct operation: claim/renew/release manage claims, reserve handles capacity reservations, unload/ensure_free/warm manage model residency, while vram_status/list_loaded/list_claims/history/trend provide distinct read views. Even similar actions like unload vs ensure_free are clearly separated by scope (single model vs threshold-based eviction).

Naming Consistency3/5

Naming uses a mix of bare verbs (renew, release, advise, unload, warm, claim, reserve), verb_noun snake_case (list_claims, list_loaded, ensure_free), and noun-only identifiers (vram_status, history, trend). The pattern is somewhat predictable by category (actions vs queries), but it lacks a single consistent convention, making it less uniform than an all-verb_noun set.

Tool Count5/5

13 tools is well within the ideal 3–15 range and each tool addresses a meaningful aspect of VRAM management: claims, reservations, eviction, warming, status, history, and trend. No tool feels redundant or superfluous.

Completeness4/5

The surface covers the core lifecycle well: claim/renew/release, reserve, warm/unload/ensure_free, plus status, history, and trend. A minor gap is the lack of a dedicated list_reservations tool—list_claims seems focused on model claims, so capacity reservations may be invisible unless they appear in status/history.

Maintenance

ActivityMaintained
ResponsivenessWithin a week