inference-lab
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@inference-labcheck GPU status and plan KV cache for qwen2.5-7b at 32k context"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Local Inference Lab
Serve, size and benchmark open LLMs on one consumer GPU (built on an RTX 5070 Ti, 16 GB, CUDA 12.8).
Three servers, one harness. vLLM, llama.cpp and Ollama run from one
docker-compose.yml; a streaming load generator measures each one on the same prompts at rising concurrency.VRAM planning before launch. A planner predicts how much KV cache vLLM will allocate, how many full-length sequences fit, and how many blocks of a too-big GGUF model llama.cpp can keep on the GPU (
-ngl), reading exact tensor sizes from the GGUF file.GPU and server telemetry. NVML sampling (peak VRAM, utilization, power) plus each server's Prometheus metrics (KV cache fill, preemptions, the cache vLLM actually allocated) recorded with every run.
An MCP server for agents. The same tools, exposed over the Model Context Protocol (FastMCP), so Claude Code or Claude Desktop can check the GPU, size a deployment and run a benchmark by asking.
Demo: howardguoui.github.io/local-inference-lab: the
published results as charts, and the vLLM planner running in your browser (free static page, rebuilt from
results/ by inference-lab demo).
Stack: Python, asyncio + httpx, vLLM, llama.cpp, Ollama, NVML, Prometheus, FastMCP, Docker Compose, pytest, GitHub Actions
flowchart LR
subgraph GPU["RTX 5070 Ti · 16 GB (one server at a time)"]
V[vLLM<br/>PagedAttention<br/>FP16 / FP8 KV]
L[llama.cpp<br/>GGUF, -ngl offload<br/>f16 / q8_0 KV]
O[Ollama]
end
P[planner<br/>KV cache + offload] -. predicts .-> V & L
B[bench<br/>async streaming load] -->|OpenAI chat API| V & L & O
B -->|/metrics| V & L
N[NVML sampler] --> B
B --> R[(results/*.json<br/>latest.md)]
M[MCP server<br/>FastMCP] --> P & B & N
A[Claude Code /<br/>Claude Desktop] -->|MCP| MWhat it measures
Metric | Why it matters |
Requests in flight (measured) | Confirms each level really ran at its concurrency |
Throughput (output tok/s, all streams) | What batching buys: decode is memory-bound, so serving 16 streams costs little more than 1 |
TTFT p50 / p95 | Queueing plus prefill: what a user waits before the first word |
TPOT p50 | Decode speed per stream once it has started |
Peak VRAM, GPU utilization, power | From NVML, sampled every 250 ms during each level |
Peak KV cache fill, preemptions | From vLLM's |
KV blocks vLLM allocated | Read from |
Every prompt starts with a unique request number, so prefix caching can't make repeated prompts look free.
Generation ignores end-of-sequence (ignore_eos), so every request on vLLM and llama.cpp produces exactly
--max-tokens tokens. Ollama's OpenAI endpoint drops that flag, so its requests can stop early; the report shows
the measured prompt and output token counts on every row, so the difference is visible.
Mid-stream errors (servers send them inside a 200 response) and empty completions count as failures.
Related MCP server: gpu-mcp-server
Results
Run scripts/run_matrix.sh on the GPU machine; it benchmarks each config in turn and writes
results/latest.md (one row per server config and concurrency level) plus a JSON file per run.
Server configs | Scenarios |
vLLM, FP16 and FP8 KV cache (Qwen2.5-7B AWQ) | chat: 512-token prompts, 256 output tokens, concurrency 1 / 4 / 16 |
llama.cpp, f16 and q8_0 KV cache (Qwen2.5-7B Q4_K_M) | long: 4,096-token prompts, 256 output tokens, concurrency 8 / 16 / 32 / 48 with 144 requests per level: 48 in flight need ~210k tokens of KV cache, more than vLLM's FP16 cache holds, which is where FP8 should pay off |
Ollama (qwen2.5:7b-instruct, Q4_K_M) | |
Optional: Qwen2.5-32B Q4_K_M through llama.cpp with the planner's | chat, concurrency 1 / 2 |
The engines are not configured identically, and the table should be read with that in mind: vLLM batches up to 64 sequences, while llama.cpp and Ollama run 8 parallel slots, so above concurrency 8 their extra requests queue and show up as TTFT. vLLM serves AWQ 4-bit weights, the others GGUF Q4_K_M.
Demo page
inference-lab demo writes docs/ (served by GitHub Pages from main): throughput and time-to-first-token
charts for every saved run, the FP16 vs FP8 KV cache comparison, and the vLLM planner ported to planner.js.
data.json carries the runs from results/, the planner's constants and model presets, and a predicted-vs-actual
check: for the Qwen2.5-7B AWQ checkpoint (5.19 GiB of weights) on the 15.92 GiB card the planner predicted
143,040 FP16 and 286,096 FP8 KV tokens against the 158,048 and 272,304 vLLM allocated (−9.5% and +5.1%).
tests/test_demo.py runs planner.js under Node and checks it against plan_vllm on 1,008 input combinations.
The planner
KV cache per token = 2 (K and V) × layers × KV heads × head dim × bytes per value. Both models below use grouped-query attention; Qwen2.5-7B needs 44% of Llama 3.1 8B's cache per token because it has 4 KV heads instead of 8 and 28 layers instead of 32 (0.5 × 0.875):
$ inference-lab plan kv --model qwen2.5-7b --tokens 32768
qwen2.5-7b: 28 layers, 4 KV heads x 128 dims (GQA 7:1), one 32,768-token sequence:
engine dtype KiB/token GiB
vllm float16 56.0 1.75
vllm fp8 28.0 0.875
llama.cpp f32 112.0 3.5
llama.cpp f16 56.0 1.75
llama.cpp q8_0 29.8 0.93
llama.cpp q5_1 21.0 0.656
llama.cpp q5_0 19.2 0.602
llama.cpp q4_1 17.5 0.547
llama.cpp q4_0 15.8 0.492 (iq4_nl is the same size)
$ inference-lab plan kv --model llama-3.1-8b --tokens 32768
vllm float16 128.0 4.0vLLM: it claims --gpu-memory-utilization × VRAM, loads the weights, reserves activation and CUDA-graph
memory, and splits the rest into 16-token KV blocks. The plan predicts the block count; the benchmark records the
real one so the two can be compared.
$ inference-lab plan vllm --model qwen2.5-7b --weights-gib 5.2 --max-model-len 32768 --gpu-gib 16
kv_budget_gib 7.7
kv_tokens 144,176
kv_blocks 9,011
max_concurrent_at_max_len 4
--kv-cache-dtype fp8 would hold about 288,352 tokens (8 full-length sequences).llama.cpp and models bigger than VRAM: with -ngl N, llama.cpp puts the output head on the GPU first, then
the last N-1 transformer blocks with their KV cache; the first blocks and the token embeddings stay in system
RAM (llama-model.cpp, where
i_gpu_start = n_layer + 1 - ngl). The planner reads each tensor's size from the GGUF file
(--gguf-path) and finds the largest N that fits; with only a file size it estimates the split from the model's
vocabulary and hidden size. Recent llama.cpp can choose N itself (--fit); the planner shows the arithmetic.
$ inference-lab plan llamacpp --model qwen2.5-32b --gguf-gib 18.5 --ctx 8192 --gpu-gib 15.5
ngl 46
blocks_on_gpu 45
total_blocks 64
vram_estimate_gib 15.28
cpu_weights_gib 5.63
-ngl 46: the output head and the last 45 of 64 blocks on the GPU, 5.6 GiB of weights in system RAM.
Generation speed is then bound by RAM bandwidth. A quantized KV cache (-ctk q8_0 -ctv q8_0, with -fa on)
frees room for more blocks.The overhead terms (activations, CUDA graphs, llama.cpp's compute buffer) are estimates with conservative
defaults, exposed as flags. Without --gpu-gib, plans use NVML: total VRAM for vLLM (it sizes from total memory),
free VRAM for llama.cpp (it must fit beside whatever else is running, such as a Windows desktop).
Run it
pip install -e ".[dev]"
inference-lab gpu # NVML snapshot
scripts/get_models.sh # GGUF for llama.cpp (vLLM and Ollama fetch their own)
docker compose --profile vllm up -d # one server at a time
inference-lab bench --backend vllm --concurrency 1 4 16 --label vllm-fp16kv
docker compose --profile vllm down
scripts/run_matrix.sh # every config, then results/latest.md
ONLY="vllm llamacpp" scripts/run_matrix.sh # a subset (e.g. Ollama already runs on the host)
scripts/publish_results.sh # run the matrix and push results to GitHubWindows: run the scripts from WSL 2 with Docker Desktop's GPU support enabled, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback in the NVIDIA Control Panel while benchmarking. Otherwise the driver spills VRAM overflow into system RAM instead of failing, and an over-sized config runs slowly but looks valid.
Use it from Claude (MCP)
claude mcp add inference-lab -e INFERENCE_LAB_HOME=/path/to/local-inference-lab -- inference-lab mcpClaude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"inference-lab": {
"command": "inference-lab",
"args": ["mcp"],
"env": { "INFERENCE_LAB_HOME": "/path/to/local-inference-lab" }
}
}
}INFERENCE_LAB_HOME tells the server where configs/backends.yaml and results/ live, since MCP clients start it
from their own working directory. An editable install (pip install -e .) finds the repo without it.
Tool | Does |
| VRAM used and free, utilization, temperature, power |
| KV cache per storage type for a model and context length |
| KV tokens and blocks vLLM will allocate; whether |
| Largest |
| Which servers are up; live KV cache fill, queue depth, preemptions |
| Benchmarks a server and saves the run; results also readable as |
Tests
pytestThe planner math (checked against published per-token KV sizes and llama.cpp's layer placement), GGUF reading on
real files written with the gguf library, the Prometheus parser on current vLLM and llama.cpp output, the load
generator against a fake streaming OpenAI-compatible server (TTFT, fixed-length output, mid-stream errors, KV cache
fill and preemptions under overload), the NVML sampler with a fake driver, and the MCP server both in memory and
as a stdio subprocess. CI runs them on pushes to main and on pull requests.
Layout
src/inference_lab/ models.py · gguf_layout.py · planner.py · gpu.py (NVML) · prom.py (/metrics) · bench.py
report.py · backends.py · paths.py · mcp_server.py (FastMCP) · cli.py
configs/ backends.yaml
scripts/ get_models.sh · run_matrix.sh
docker-compose.yml vllm · llamacpp · ollama profilesThis server cannot be deployed
Maintenance
Related MCP Connectors
Provides capabilities that let LLM agents perform a range of infrastructure management tasks.
On-demand GPU nodes for agents: create nodes, run commands, and submit jobs, billed by the minute.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Protocol-native energy infrastructure orchestration for AI data centers. Provides 46 MCP tools across 8 grid protocols (IEC-61850, DNP3, Modbus, OCPP, OpenADR, IEEE 2030.5, IEC 60870-5-104, ICCP) with 5 core API primitives: connect, dispatch, settle, comply, and intel. Enables AI agents to programmatically interact with substations, grid interfaces, and energy assets for real-time workload-grid coordination.
Related MCP Servers
- AlicenseAqualityCmaintenanceExposes NVIDIA GPU metrics (info, utilization, VRAM, temperature) via MCP tools for real-time querying from AI assistants.527 PyPIMIT
- AlicenseNot gradedqualityAmaintenanceAn MCP server that enables AI agents to query real-time NVIDIA GPU metrics like utilization, memory, temperature, and power without external monitoring tools.15Apache 2.0
- AlicenseAqualityAmaintenanceEnables AI agents to monitor server health and capacity, diagnose outages, inspect Docker and deployment status, and perform safe, bounded recovery actions via MCP without granting unrestricted shell or SSH access.101Apache 2.0
- AlicenseAqualityCmaintenanceEnables MCP clients to scan local GGUF models, estimate VRAM and suggest GPU offload layers, manage llama-server lifecycle, and proxy OpenAI-format chats with idle auto-unload.9MIT