Skip to main content
Glama

Local Inference Lab

Serve, size and benchmark open LLMs on one consumer GPU (built on an RTX 5070 Ti, 16 GB, CUDA 12.8).

  • Three servers, one harness. vLLM, llama.cpp and Ollama run from one docker-compose.yml; a streaming load generator measures each one on the same prompts at rising concurrency.

  • VRAM planning before launch. A planner predicts how much KV cache vLLM will allocate, how many full-length sequences fit, and how many blocks of a too-big GGUF model llama.cpp can keep on the GPU (-ngl), reading exact tensor sizes from the GGUF file.

  • GPU and server telemetry. NVML sampling (peak VRAM, utilization, power) plus each server's Prometheus metrics (KV cache fill, preemptions, the cache vLLM actually allocated) recorded with every run.

  • An MCP server for agents. The same tools, exposed over the Model Context Protocol (FastMCP), so Claude Code or Claude Desktop can check the GPU, size a deployment and run a benchmark by asking.

Demo: howardguoui.github.io/local-inference-lab: the published results as charts, and the vLLM planner running in your browser (free static page, rebuilt from results/ by inference-lab demo).

Stack: Python, asyncio + httpx, vLLM, llama.cpp, Ollama, NVML, Prometheus, FastMCP, Docker Compose, pytest, GitHub Actions

flowchart LR
    subgraph GPU["RTX 5070 Ti · 16 GB (one server at a time)"]
        V[vLLM<br/>PagedAttention<br/>FP16 / FP8 KV]
        L[llama.cpp<br/>GGUF, -ngl offload<br/>f16 / q8_0 KV]
        O[Ollama]
    end
    P[planner<br/>KV cache + offload] -. predicts .-> V & L
    B[bench<br/>async streaming load] -->|OpenAI chat API| V & L & O
    B -->|/metrics| V & L
    N[NVML sampler] --> B
    B --> R[(results/*.json<br/>latest.md)]
    M[MCP server<br/>FastMCP] --> P & B & N
    A[Claude Code /<br/>Claude Desktop] -->|MCP| M

What it measures

Metric

Why it matters

Requests in flight (measured)

Confirms each level really ran at its concurrency

Throughput (output tok/s, all streams)

What batching buys: decode is memory-bound, so serving 16 streams costs little more than 1

TTFT p50 / p95

Queueing plus prefill: what a user waits before the first word

TPOT p50

Decode speed per stream once it has started

Peak VRAM, GPU utilization, power

From NVML, sampled every 250 ms during each level

Peak KV cache fill, preemptions

From vLLM's /metrics: a full cache makes vLLM evict and recompute sequences. Current llama.cpp and Ollama don't publish a KV fill metric

KV blocks vLLM allocated

Read from vllm:cache_config_info and compared with the planner's prediction

Every prompt starts with a unique request number, so prefix caching can't make repeated prompts look free. Generation ignores end-of-sequence (ignore_eos), so every request on vLLM and llama.cpp produces exactly --max-tokens tokens. Ollama's OpenAI endpoint drops that flag, so its requests can stop early; the report shows the measured prompt and output token counts on every row, so the difference is visible. Mid-stream errors (servers send them inside a 200 response) and empty completions count as failures.

Related MCP server: gpu-mcp-server

Results

Run scripts/run_matrix.sh on the GPU machine; it benchmarks each config in turn and writes results/latest.md (one row per server config and concurrency level) plus a JSON file per run.

Server configs

Scenarios

vLLM, FP16 and FP8 KV cache (Qwen2.5-7B AWQ)

chat: 512-token prompts, 256 output tokens, concurrency 1 / 4 / 16

llama.cpp, f16 and q8_0 KV cache (Qwen2.5-7B Q4_K_M)

long: 4,096-token prompts, 256 output tokens, concurrency 8 / 16 / 32 / 48 with 144 requests per level: 48 in flight need ~210k tokens of KV cache, more than vLLM's FP16 cache holds, which is where FP8 should pay off

Ollama (qwen2.5:7b-instruct, Q4_K_M)

Optional: Qwen2.5-32B Q4_K_M through llama.cpp with the planner's -ngl

chat, concurrency 1 / 2

The engines are not configured identically, and the table should be read with that in mind: vLLM batches up to 64 sequences, while llama.cpp and Ollama run 8 parallel slots, so above concurrency 8 their extra requests queue and show up as TTFT. vLLM serves AWQ 4-bit weights, the others GGUF Q4_K_M.

Demo page

inference-lab demo writes docs/ (served by GitHub Pages from main): throughput and time-to-first-token charts for every saved run, the FP16 vs FP8 KV cache comparison, and the vLLM planner ported to planner.js. data.json carries the runs from results/, the planner's constants and model presets, and a predicted-vs-actual check: for the Qwen2.5-7B AWQ checkpoint (5.19 GiB of weights) on the 15.92 GiB card the planner predicted 143,040 FP16 and 286,096 FP8 KV tokens against the 158,048 and 272,304 vLLM allocated (−9.5% and +5.1%). tests/test_demo.py runs planner.js under Node and checks it against plan_vllm on 1,008 input combinations.

The planner

KV cache per token = 2 (K and V) × layers × KV heads × head dim × bytes per value. Both models below use grouped-query attention; Qwen2.5-7B needs 44% of Llama 3.1 8B's cache per token because it has 4 KV heads instead of 8 and 28 layers instead of 32 (0.5 × 0.875):

$ inference-lab plan kv --model qwen2.5-7b --tokens 32768
qwen2.5-7b: 28 layers, 4 KV heads x 128 dims (GQA 7:1), one 32,768-token sequence:

  engine     dtype     KiB/token      GiB
  vllm       float16        56.0     1.75
  vllm       fp8            28.0    0.875
  llama.cpp  f32           112.0      3.5
  llama.cpp  f16            56.0     1.75
  llama.cpp  q8_0           29.8     0.93
  llama.cpp  q5_1           21.0    0.656
  llama.cpp  q5_0           19.2    0.602
  llama.cpp  q4_1           17.5    0.547
  llama.cpp  q4_0           15.8    0.492      (iq4_nl is the same size)

$ inference-lab plan kv --model llama-3.1-8b --tokens 32768
  vllm       float16       128.0      4.0

vLLM: it claims --gpu-memory-utilization × VRAM, loads the weights, reserves activation and CUDA-graph memory, and splits the rest into 16-token KV blocks. The plan predicts the block count; the benchmark records the real one so the two can be compared.

$ inference-lab plan vllm --model qwen2.5-7b --weights-gib 5.2 --max-model-len 32768 --gpu-gib 16
  kv_budget_gib              7.7
  kv_tokens                  144,176
  kv_blocks                  9,011
  max_concurrent_at_max_len  4
  --kv-cache-dtype fp8 would hold about 288,352 tokens (8 full-length sequences).

llama.cpp and models bigger than VRAM: with -ngl N, llama.cpp puts the output head on the GPU first, then the last N-1 transformer blocks with their KV cache; the first blocks and the token embeddings stay in system RAM (llama-model.cpp, where i_gpu_start = n_layer + 1 - ngl). The planner reads each tensor's size from the GGUF file (--gguf-path) and finds the largest N that fits; with only a file size it estimates the split from the model's vocabulary and hidden size. Recent llama.cpp can choose N itself (--fit); the planner shows the arithmetic.

$ inference-lab plan llamacpp --model qwen2.5-32b --gguf-gib 18.5 --ctx 8192 --gpu-gib 15.5
  ngl                46
  blocks_on_gpu      45
  total_blocks       64
  vram_estimate_gib  15.28
  cpu_weights_gib    5.63
  -ngl 46: the output head and the last 45 of 64 blocks on the GPU, 5.6 GiB of weights in system RAM.
  Generation speed is then bound by RAM bandwidth. A quantized KV cache (-ctk q8_0 -ctv q8_0, with -fa on)
  frees room for more blocks.

The overhead terms (activations, CUDA graphs, llama.cpp's compute buffer) are estimates with conservative defaults, exposed as flags. Without --gpu-gib, plans use NVML: total VRAM for vLLM (it sizes from total memory), free VRAM for llama.cpp (it must fit beside whatever else is running, such as a Windows desktop).

Run it

pip install -e ".[dev]"
inference-lab gpu                                    # NVML snapshot

scripts/get_models.sh                                # GGUF for llama.cpp (vLLM and Ollama fetch their own)
docker compose --profile vllm up -d                  # one server at a time
inference-lab bench --backend vllm --concurrency 1 4 16 --label vllm-fp16kv
docker compose --profile vllm down

scripts/run_matrix.sh                                # every config, then results/latest.md
ONLY="vllm llamacpp" scripts/run_matrix.sh           # a subset (e.g. Ollama already runs on the host)
scripts/publish_results.sh                           # run the matrix and push results to GitHub

Windows: run the scripts from WSL 2 with Docker Desktop's GPU support enabled, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback in the NVIDIA Control Panel while benchmarking. Otherwise the driver spills VRAM overflow into system RAM instead of failing, and an over-sized config runs slowly but looks valid.

Use it from Claude (MCP)

claude mcp add inference-lab -e INFERENCE_LAB_HOME=/path/to/local-inference-lab -- inference-lab mcp

Claude Desktop (claude_desktop_config.json):

{
  "mcpServers": {
    "inference-lab": {
      "command": "inference-lab",
      "args": ["mcp"],
      "env": { "INFERENCE_LAB_HOME": "/path/to/local-inference-lab" }
    }
  }
}

INFERENCE_LAB_HOME tells the server where configs/backends.yaml and results/ live, since MCP clients start it from their own working directory. An editable install (pip install -e .) finds the repo without it.

Tool

Does

gpu_status

VRAM used and free, utilization, temperature, power

kv_cache_size, list_models

KV cache per storage type for a model and context length

plan_vllm_deployment

KV tokens and blocks vLLM will allocate; whether max_model_len fits

plan_llamacpp_offload

Largest -ngl for a GGUF file (exact tensor sizes) and its VRAM estimate

list_backends, backend_metrics

Which servers are up; live KV cache fill, queue depth, preemptions

run_benchmark

Benchmarks a server and saves the run; results also readable as lab://results/latest

Tests

pytest

The planner math (checked against published per-token KV sizes and llama.cpp's layer placement), GGUF reading on real files written with the gguf library, the Prometheus parser on current vLLM and llama.cpp output, the load generator against a fake streaming OpenAI-compatible server (TTFT, fixed-length output, mid-stream errors, KV cache fill and preemptions under overload), the NVML sampler with a fake driver, and the MCP server both in memory and as a stdio subprocess. CI runs them on pushes to main and on pull requests.

Layout

src/inference_lab/  models.py · gguf_layout.py · planner.py · gpu.py (NVML) · prom.py (/metrics) · bench.py
                    report.py · backends.py · paths.py · mcp_server.py (FastMCP) · cli.py
configs/            backends.yaml
scripts/            get_models.sh · run_matrix.sh
docker-compose.yml  vllm · llamacpp · ollama profiles

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Exposes NVIDIA GPU metrics (info, utilization, VRAM, temperature) via MCP tools for real-time querying from AI assistants.
    5
    27 PyPI
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server that enables AI agents to query real-time NVIDIA GPU metrics like utilization, memory, temperature, and power without external monitoring tools.
    15
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to monitor server health and capacity, diagnose outages, inspect Docker and deployment status, and perform safe, bounded recovery actions via MCP without granting unrestricted shell or SSH access.
    10
    1
    Apache 2.0
  • A
    license
    A
    quality
    C
    maintenance
    Enables MCP clients to scan local GGUF models, estimate VRAM and suggest GPU offload layers, manage llama-server lifecycle, and proxy OpenAI-format chats with idle auto-unload.
    9
    MIT