inferwatch-mcp
Monitors Ollama inference servers by ingesting their logs to capture request-level metrics such as latency, token counts, and GPU utilization, and exposes these metrics through a dashboard and MCP tools for querying historical data.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@inferwatch-mcpWhat's the current GPU utilization and request rate for my Ollama instance?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
inferwatch
Real-time and historical metrics for locally served LLMs — Ollama and vLLM — with a browser dashboard, a settings screen, and an MCP server so an agent can query the same data.
One Python process, one SQLite file. No Docker, no Node, no Prometheus, no external services. It never sits in the request path, so it cannot slow down or break inference.
┌── Ollama ──────────────┐ ┌── vLLM ────────────────┐
│ journald / file / │ │ GET /metrics │
│ docker logs │ │ (native Prometheus) │
└──────────┬─────────────┘ └──────────┬─────────────┘
│ per-request rows │ pre-aggregated
▼ ▼
┌──────────────── SQLite (WAL) ────────────────┐
│ requests · rollups · vllm_samples/hist │
└───────┬──────────────────────────┬───────────┘
▼ ▼
dashboard :7070 MCP server (stdio)The engines are not symmetric, and the tool does not pretend otherwise
This is the central design fact, so it is worth stating plainly.
Ollama | vLLM | |
Source | its log |
|
Per-request rows | yes | no — none exist to collect |
TTFT / latency | exact, per request | histograms only |
Tokens | per request | cumulative counters |
Errors | HTTP status per request |
|
Client address | yes | no |
Percentiles | exact within retention | bucket upper bounds; means are exact |
KV / prompt cache | occupancy + eviction counts, sampled from the log | occupancy gauge, scraped |
Unique extras | prompt cache reuse, draft accept, cold-load time, KV VRAM/RAM split | preemptions, batch occupancy, waiting-by-reason |
Both tabs show GPU utilisation, VRAM, temperature and power draw, since those
are measured by nvidia-smi rather than by either engine. Temperature and power
get separate charts rather than sharing an axis, and each is aggregated the way
its unit demands: utilisation averages across cards, VRAM and watts sum,
temperature reports the hottest card. Temperature is the one series not plotted
from zero — a 33–68 °C range starting at 0 wastes most of the plot.
And a third family: image generation
SwarmUI and its ComfyUI backends are further from either of those than they are from each other. There are no tokens, no time-to-first-token and no context window; the unit of work is a generation with a duration, the models a workflow loaded, and a node that may have thrown. So it gets its own tab, its own tables and its own MCP tools rather than blank columns in someone else's.
Two sources feed it, and only one of them counts:
ComfyUI | SwarmUI journal | |
Generations | the count — stable | a timeline, never a second count |
Models | read out of the workflow graph | the name on the request line |
Errors | failing node, its class, the exception | WebAPI failures + backend stderr |
Timing | total, measured | prep vs gen, split |
Joining them would mean guessing which log line belongs to which prompt_id,
and a request for N images produces N finish lines — so a 1:1 pairing would
misattribute every batch. They are kept apart instead: generations come from
/history, and SwarmUI's prep-versus-gen split is reported as its own
aggregate. Durations here are exact; every one is measured, so unlike the
token engines there is no bucketed-percentile caveat.
Backend ports are published exactly once, when SwarmUI starts its backends
(Self-Start ComfyUI-0 on port 7821 started.). They are learned from the log,
remembered across restarts, and as a last resort probed on the conventional
range — a collector restarted mid-life resumes the journal past those lines and
would otherwise collect queue depth and no generations at all.
So they get separate dashboard tabs, separate tables and separate MCP tools. No attempt is made to reconstruct per-request rows for vLLM by differencing counters: you cannot recover which TTFT belonged to which request, and faking it would put invented rows beside real ones.
Ollama: where the numbers come from
Ollama exposes no /metrics endpoint (verified — the route is not in the
binary). With OLLAMA_DEBUG=1, the embedded llama.cpp prints a timing block per
request, which is combined with the access line and the scheduler line:
slot print_timing: id 0 | task 6763 | prompt eval time = 1254.52 ms / 55 tokens
slot print_timing: id 0 | task 6763 | eval time = 14591.31 ms / 416 tokens
[GIN] ... | 200 | 16.862061865s | 192.0.2.10 | POST "/v1/chat/completions"
time=... msg="context for request finished" runner.name=.../llama3.2:3bThat yields TTFT, prefill/decode split, token counts, decode rate, status, client, endpoint and model — for every request, from every client, without touching the request path. Two numbers fall out of the combination that neither source has alone:
queue wait = wall latency − runner time: time spent waiting rather than generating. A proxy cannot separate these.
prompt cache reuse = full prompt length − tokens actually evaluated.
OLLAMA_DEBUG=1 is required. Without it llama.cpp prints no timing lines:
request rates, statuses and GPU metrics still work, but TTFT and token counts
stay empty. The dashboard says so in a banner instead of showing zeros.
# /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_DEBUG=1"Ollama: KV and prompt cache
Ollama publishes no metrics endpoint, so there is nothing to scrape — but its
runner logs the cache state outright, and that line is collected into
ollama_cache_samples:
srv update: - cache state: 30 prompts, 8010.969 MiB (limits: 8192.000 MiB, 32768 tokens, 68068 est)
srv get_availabl: prompt cache update took 364.22 msThat is 97.8% of an 8 GiB pool, and the maintenance pass that produced it cost 364 ms inside the request that triggered it — cache pressure shows up as TTFT, which is why the cost is stored next to the occupancy rather than separately.
There are two different caches here and the tool keeps them apart:
What it is | Where it comes from | |
Prompt cache | the bounded pool of saved prompt states that lets a returning conversation skip prefill |
|
Live KV cache | the slot's own context memory, preallocated at load |
|
Prompt-cache occupancy is the closer analogue of vLLM's kv_cache_usage_perc.
Live KV occupancy is the distance to a context shift or truncation — a
request at 99.9% is one token from losing history.
Alongside the occupancy gauge, the same lines yield eviction pressure, split by
reason because the two mean different things: evict_crowded says the pool is
too small, evict_invalidated says the cached positions no longer applied. Each
eviction is a prefill somebody pays for again later.
Two caveats that shape every figure:
Sampling is ollama's, not ours. A
cache stateline appears only when ollama runs a cache update, so an empty window means no updates happened, not the cache was empty.samplesis always reported alongside, and the eviction counters are stored as deltas between samples and summed rather than turned into rates — a rate over unevenly spaced deltas would be fiction.The lines need the runner at high log verbosity. They come from llama-server, not from ollama's Go code. Ollama 0.32.x passes
--log-verbosity 4on every load, so they are present by default there;OLLAMA_DEBUG=1is the documented way to be sure of it. Note that/api/health'sdebug_loggingonly reports whetherOLLAMA_DEBUGis set on the unit, so it can readfalseon a host that is logging these lines perfectly well.
Model loads also record how the KV cache was placed:
llama_kv_cache: CUDA0 KV buffer size = 512.00 MiB
llama_kv_cache: CPU KV buffer size = 4096.00 MiBA cache that did not fit in VRAM caps decode throughput for the whole life of
the load, so that raises a kv_offload WARN event ("89% of the KV cache is
on host RAM, not VRAM") rather than being left buried in the load event's
detail.
vLLM: where the numbers come from
vLLM's native Prometheus endpoint is scraped every
collection.scrape_interval_s (default 10s). Cumulative counters are
differenced; histogram buckets are differenced per bucket; both are written
aggregated per minute, because storing every scrape would add millions of
rows a month at resolutions no chart uses.
Three details worth knowing:
vLLM's own bucket bounds are stored with the counts. Its bounds step 1ms/20ms/250ms/2.5s/40s/640s; Ollama's step 25ms/200ms/1.5s/15s/60s. Neither is a refinement of the other, so re-bucketing one into the other would require interpolating between bounds — inventing numbers. Percentiles are computed against each source's own bounds and reported as bucket upper bounds.
_sumand_countare exact, so the mean is exact. Since vLLM's buckets are coarse in the seconds range, the dashboard and MCP tools lead with the mean and label percentiles as "at most".Restarts are detected via
process_start_time_seconds(and by a counter going backwards). The interval spanning a restart is dropped rather than emitted as a bogus delta.
Inter-token latency has been spelled time_per_output_token_seconds,
inter_token_latency_seconds and request_time_per_output_token_seconds across
vLLM releases. All are collected and whichever has data is used, so this works
against old and new servers with no configuration.
Related MCP server: RHOAI Observability MCP
Install
Requires Python 3.10 or newer — not because of this code (which is 3.9-clean) but because fastapi, uvicorn, starlette and mcp all require it.
pip install git+https://github.com/floatsmyboat/inferwatch # or:
git clone https://github.com/floatsmyboat/inferwatch && cd inferwatch
python3 -m venv --upgrade-deps .venv && .venv/bin/pip install -e ".[dev]"Installing gives you two commands:
Command | What it is |
| the collector, dashboard and API ( |
| the MCP server, over stdio |
Not on PyPI yet; install from git for now.
Backfill from a journal you already have, then look at the database:
inferwatch ingest --since 2d # or: python -m inferwatch.main ingest
inferwatch statsRun it:
inferwatch serve # http://127.0.0.1:7070As a service — the unit is rendered from systemd/inferwatch.service.in for the
current user, checkout path and interpreter, so nothing is hardcoded:
./scripts/install-systemd.sh # system service (uses sudo)
sudo systemctl enable --now inferwatch
./scripts/install-systemd.sh --user # or per-user, no sudo
systemctl --user enable --now inferwatchOverride with HOST=0.0.0.0 PORT=7070 DATADIR=... ./scripts/install-systemd.sh.
Exposing it on a network
There is no authentication, and the API is not read-only. It stores no prompt or response text — only counts, timings, model names and client addresses — but anyone who can reach the port can also write:
Endpoint | What an unauthenticated caller can do |
| change any setting not pinned by env or flag |
| add, edit or remove a monitored engine — and with |
| make the server fetch an arbitrary URL, with an arbitrary bearer token |
That last one is a server-side request forgery primitive: the probe exists so a typo surfaces before a source is saved, and it will dial whatever it is given.
So the firewall is not belt-and-braces, it is the only control:
sudo ufw allow from 192.168.1.0/24 to any port 7070 proto tcp comment "inferwatch"Better still, leave server.host at 127.0.0.1 and put something that
authenticates in front of it.
Secrets
A vLLM source's api_key is stored in the sources table in plain text —
SQLite has no encryption here, and the collector needs the value to send it as a
bearer token. So prefer an indirection:
inferwatch sources add --kind vllm vllm-prod \
--set url=http://127.0.0.1:8000 --set 'api_key=${VLLM_API_KEY}'${VAR} and $VAR are expanded when the request is made, so the database holds
only the pointer and the secret stays in the environment (a systemd
EnvironmentFile= is the natural home). An unset variable resolves to empty and
logs a warning, so the request fails as a clean 401 rather than sending a
literal ${VAR} as the token.
Whatever is stored, it is masked on every read: /api/sources,
/api/config, the Settings tab, inferwatch sources list, and the MCP
list_sources and run_sql tools all report ***redacted*** instead of the
value. A reference like ${VLLM_API_KEY} is shown as itself, since knowing
which variable is referenced is useful and the reference is not the secret.
Sending ***redacted*** back in a PUT means "leave it unchanged", so editing
a source through the API cannot overwrite a key with its own mask; sending an
empty string still clears it.
Note that a literal passed as --set api_key=… is visible in this process's
command line to any local user for as long as the command runs, which is a
second reason to prefer the reference form.
Configuring what is monitored
Everything is editable from the dashboard's Settings tab, or from the CLI:
python -m inferwatch.main sources # list
python -m inferwatch.main sources add --kind vllm --name qwen \
--set url=http://127.0.0.1:8000
python -m inferwatch.main sources add --kind ollama --name box \
--set reader=file --set path=~/.ollama/logs/server.log
python -m inferwatch.main sources disable qwenOllama log readers. Not everyone runs Ollama under systemd:
reader | for | timestamp fidelity |
|
| microsecond, from journald |
|
| derived; see below |
| Ollama in a container | per-line, from |
The file reader follows like tail -F, surviving rotation (inode change) and
truncation, and persists an offset so a restart does not replay. Ollama's Go
lines carry time=, but llama.cpp's slot lines — the ones holding the token
counts — carry no timestamp, so the most recent one seen is carried forward.
Ordering, which the joins depend on, always holds; absolute precision is lower
than journald's, and a second-resolution [GIN] line is clamped forward so time
never appears to run backwards.
Settings precedence
spec default < database (Settings tab) < environment < command lineA key supplied by the environment or a flag is shown read-only in the
Settings tab with its origin, because the process was told to use it and a
browser must not override that silently. Saving is all-or-nothing, so a typo in
one field cannot leave a half-applied configuration. Changes to sources,
intervals and retention apply without a restart; server.host and
server.port are marked as needing one, and the API says so after saving.
Configuration reference
Every setting below is editable in the Settings tab, settable as an environment
variable, and some are pinnable with a flag. The tables are generated from the
code (scripts/gen-config-docs.py), so they cannot drift from what the program
actually accepts.
Generated by scripts/gen-config-docs.py — do not edit by hand.
Collection
Setting | Default | Accepts | Environment variable | Notes |
|
| 1–300 |
| How often nvidia-smi and the engine's own status endpoint are sampled. |
|
| 1–300 |
| How often each vLLM instance's /metrics endpoint is read. vLLM counters are cumulative, so this sets the resolution of every rate and histogram derived from them. |
|
|
|
| restart required — How far back to read on a first run, before any resume state exists. Accepts 7d / 6h, or a journalctl form like '-2 days'. |
|
| 10–3600 |
| How often the 1-minute and 1-hour aggregates are recomputed. |
Retention
Setting | Default | Accepts | Environment variable | Notes |
|
| 0.5–3650 |
| Per-request detail older than this is deleted. Rollups are kept indefinitely regardless, so long-range charts survive. |
|
| 0.5–3650 |
| GPU samples, engine samples and the event log are trimmed to this. |
Dashboard
Setting | Default | Accepts | Environment variable | Notes |
|
|
|
| Range selected when the dashboard is opened. |
|
| — |
| Include HEAD / and status polling in request rates. Off by default because on a polled instance it can be 90%+ of hits. |
|
| 2–600 |
| How often the open dashboard refetches. The live request feed is pushed separately and is not affected by this. |
Server
Setting | Default | Accepts | Environment variable | Notes |
|
| — |
| restart required — 0.0.0.0 exposes the dashboard on the network. There is no authentication, so restrict it at your firewall. |
|
| 1–65535 |
| restart required — Port the dashboard and API listen on. |
Source fields
Set these with --set key=value on sources add, or in the Settings tab.
Ollama (--kind ollama)
Field | Default | Required when | Notes |
|
| — | Where to read ollama's log from. Per-request metrics come from llama.cpp's debug lines, so one of these is required. One of |
|
|
| Also used to attribute GPUs: every process in this unit's cgroup is matched against nvidia-smi, so only the cards ollama actually holds appear on its pane. |
| — |
| Followed like tail -F, so rotation and truncation are handled. |
|
|
| |
|
| — | Used to poll /api/ps for resident models. |
| — | — | Optional. Resolves blob digests to model names on load events. Defaults to $OLLAMA_MODELS or ~/.ollama/models. |
SwarmUI / ComfyUI (--kind swarmui)
Field | Default | Required when | Notes |
|
| — | Polled for queue depth and backend health. Its API needs no key for a local install. |
|
| — | Optional, unlike ollama's. Adds SwarmUI's prep-vs-gen timing split, WebAPI failures that never reached a backend, and the Python stderr behind a failed generation. 'none' polls the APIs only. One of |
|
|
| Also the fallback for GPU attribution, though each self-started ComfyUI is usually resolved exactly by its own port. |
| — |
| SwarmUI's own rotated logs under Data/Logs work here. |
| — |
| |
| — | — | Optional, comma separated. Leave empty and the backend ports are discovered from the log, which is the only place SwarmUI publishes them. Required when the reader is 'none'. |
|
| — | How many /history entries to read each tick. ComfyUI keeps this in memory only, so a larger number costs little and survives a burst between polls. |
vLLM (--kind vllm)
Field | Default | Required when | Notes |
|
| — | The OpenAI-compatible server root. /metrics is read from here. |
| — | — | The unit running the ENGINE, which is what its GPUs are attributed by -- every process in that unit's cgroup is matched against nvidia-smi. Set it to the engine's unit, not a proxy in front of it: with a proxy on the URL there are no GPUs behind that port and attribution reports 'could not attribute'. If a switcher rotates flavours of the same engine, list every candidate separated by commas -- whichever is running is the one attributed, so a switch does not silently turn attribution off. The journal is also read from it for HTTP status codes, client addresses and engine errors, which /metrics does not expose. |
| — | — | Sent as a bearer token if the server requires one. Prefer an indirection like ${VLLM_API_KEY} over pasting the value: what is stored here goes into the database in plain text, and a reference keeps the secret in the environment or an EnvironmentFile instead. |
Command-line flags
Flag | Purpose | Pins the setting |
| — | — |
| systemd unit for the seeded Ollama source / for ingest | — |
| — | — |
| ollama models dir (resolves blob digests to model names) | — |
| ingest: read this log file instead of the journal | — |
| log backfill window, e.g. '-2 days' |
|
| raw request retention; rollups are kept forever |
|
| — |
|
| — |
|
| — |
|
| — |
|
| — | — |
Flags that pin a setting outrank both the environment and the Settings tab; the tab shows those keys read-only with their origin.
Scope: one Ollama source; many vLLM and SwarmUI sources
vLLM and SwarmUI rows are keyed by source throughout, so any number of those
instances can be monitored side by side. The Ollama tables (requests,
events, ps_samples) are not source-partitioned, so exactly one Ollama
source runs at a time; enabling a second logs a warning and ignores it rather
than silently blending two instances into one set of numbers. Partitioning those
tables is a schema change worth doing deliberately.
One SwarmUI source covers all of that install's ComfyUI backends, which are discovered rather than configured; the per-backend rows are keyed by backend name within the source.
Dashboard
http://127.0.0.1:7070 — four tabs: Ollama, vLLM, Images, Settings.
Which GPUs belong to which engine
A host often runs more than one engine, so plotting every card on an engine's pane would imply it uses all of them. Both panes scope their GPU charts and tiles to the cards their engine actually holds: those carry the series colour and the aggregates (VRAM, watts, hottest card) count only them, while the host's other cards stay visible in grey, labelled "other engine".
GPUs are resolved by three methods, in descending order of trust:
Method | How | When it applies |
cgroup | every pid in a unit's cgroup, intersected with | a |
↳ several units | each candidate resolved, whichever are running unioned | a switcher rotates flavours of one engine |
pids | ollama's own | Ollama, for the |
port | the listening pid → its descendants | last resort, no unit configured |
Prefer configuring unit. The port walk assumes the processes holding the GPUs
are children of whatever answers the port, and that is false for anything
non-trivial: vLLM v1 runs its engine core and workers as separate processes, and
a reverse proxy in front of the API server severs the link entirely. A cgroup is
what survives reparenting. On the development host the URL's listener is in
vllm-proxy.service while the workers are in vllm-qwen38.service, so only the
cgroup method finds them — set unit to the unit running the engine, never
the proxy.
A single unit name is not always correct for long. Here a switcher starts
exactly one of vllm-qwen38{,-w4a16,-dflash2}.service, so the name pinned in
the config went stale at the first switch and attribution fell to unknown —
which, on the vLLM pane, plots every card on the host as vLLM's own and sums
another engine's VRAM and watts into its tiles. So unit takes a comma-separated
list, and whichever candidates are running are the ones attributed:
unit = vllm-qwen38, vllm-qwen38-w4a16, vllm-qwen38-dflash2A candidate that is not running contributes nothing and is not an error — with a switcher the inactive flavours are expected to be absent. Only when none of them is running is the answer unknown, since names matching no cgroup cannot be told apart from wrong names. Two units up at once during a handover are unioned rather than raced, so the list's order never decides the answer.
Three outcomes, rendered differently, because collapsing them is how a wrong answer gets presented as a right one:
specific cards — "holds GPU 2, 3 of 4 · via cgroup:vllm-qwen38.service", with the method named so a surprising answer is diagnosable
none — attributed, and the engine is on no card: a CPU-only instance, or simply an idle one. Holding nothing is not evidence that anything else holds those cards, so no device is labelled "(other engine)" in this state — that suffix is used only once the engine demonstrably holds something. The pane names the reason instead: ollama's runner exits on keep-alive expiry, so its pane reads "no runner resident — holds none of the 4 devices". Before, every card was marked "(other engine)" and an idle ollama looked dispossessed.
unknown — no nvidia-smi, no cgroup visibility, a remote engine, or an empty result from the weakest method (far likelier to be the wrong process tree than a genuinely idle engine). Every card is shown without emphasis and the pane says it could not attribute.
A chart is not an instant. Attribution answers "which cards does this engine
hold now", but the GPU charts cover a window, and ollama's runner exits on
keep-alive expiry — on this host 87% of samples in a typical hour have no model
resident. Colouring history by the instant therefore greyed every card and
blanked the tiles while the window still contained real load: 15.5 GiB and 100%
utilisation on two of them. So the charts are scoped by what the engine held at
any point within the window, recorded per /api/ps sample, while the KPI
tiles stay on the live answer — which is the right one for "in use now". A
window reaching back before that record began is covered in part, and the pane
says which part rather than implying all of it.
Attribution is re-resolved on an interval rather than cached once, and the stored answer records when and how it was learned. A change in the number of resident models also forces a re-resolve, because the interval alone is too coarse to record faithfully: a model can load and be evicted well inside one recheck, and the samples written meanwhile would claim the engine held nothing while it was busy on two cards. An earlier version cached the first success and let a failure be merged away, so a topology change never propagated and the pane kept presenting weeks-old indices as current fact.
One filter row scopes everything below it. Every chart has a Table toggle showing the same series as numbers, so no value is reachable only by hovering. A live SSE feed drives the request ticker and the current-rate figure.
URL parameters: ?tab=vllm, ?window=6h, ?model=llama3.2:3b, ?source=name,
?nostream=1 (disables the live feed — useful for kiosk displays and screenshot
tools, which otherwise wait forever on an open stream).
API
Endpoint | Returns |
| everything the Ollama tab needs, one time slice |
| same for one vLLM instance |
| Ollama breakdowns |
| vLLM breakdowns |
| raw rows and timelines |
| Ollama prompt-cache occupancy, evictions, update cost |
| requests in flight at once, peak/mean and slot capacity |
| per-client detail: models requested, context sizes, tokens, TTFT |
| traffic by endpoint and class |
| everything the Images tab needs, one time slice |
| image breakdowns |
| settings |
| monitored engines |
| dashboard defaults, collector state |
| SSE live feed |
/api/sources/probe checks a definition before it is saved, so a typo
surfaces there rather than as silence in the charts.
MCP server
./scripts/install-mcp.sh # writes .mcp.json for this checkout (gitignored)or claude mcp add inferwatch -- /path/to/.venv/bin/python -m inferwatch.mcp_server.
Opens the same SQLite file read-only (mode=ro plus PRAGMA query_only) and
answers through the same query layer as the dashboard, so a number it reports
always matches the number on screen.
Serving it over HTTP
--transport streamable-http (or sse) serves it as a network service instead
of a subprocess:
sudo ./scripts/install-mcp-key.sh # writes /etc/inferwatch/mcp-api-key, 0640
inferwatch-mcp --transport streamable-http --port 7071Then point a client at http://127.0.0.1:7071/mcp with
Authorization: Bearer <key> (or X-API-Key: <key>, for clients that can only
set a plain header).
An API key is required, and the server refuses to start without one. Over
stdio there is nothing to protect — the client spawns the process and owns both
ends of the pipe. Over HTTP every tool here reads the metrics database, client
addresses included, and run_sql is a general read-only query tool over all of
it. So HTTP without a key fails closed with instructions rather than starting
quietly; --allow-unauthenticated exists for a socket nothing else can reach
and logs a warning each time.
The key is looked up in this order, and never lives in the repository or the metrics database:
Source | Notes |
| explicit, wins |
| the value — natural for a systemd |
| a path |
| the default |
A world-readable key file is refused, not warned about. Group-readable is
allowed, which is how you hand it to a unit (--group inferwatch). A file
containing INFERWATCH_MCP_API_KEY=… is accepted as well as a bare key, since
a key file and an EnvironmentFile look alike at 2am.
--host defaults to 127.0.0.1 and --port to 7071, and both can also be
set with INFERWATCH_MCP_HOST / INFERWATCH_MCP_PORT (the flag wins). Neither
was reachable before: the SDK's own default is 127.0.0.1:8000, which collides
with a vLLM server on the same box, and no flag or environment variable moved
it.
Binding a non-loopback address also turns off the SDK's Host-header
allow-list. That check is DNS-rebinding protection, which exists for
unauthenticated services a browser could be tricked into calling; here every
request must carry the API key, which a rebinding attacker cannot supply. Left
on, a server bound to 0.0.0.0 answers 421 Misdirected Request to every
client that addresses it by its real address. On loopback the SDK default is
untouched.
Running it on boot
sudo ./scripts/install-mcp-key.sh --rotate --group "$(id -gn)"
MCPHOST=0.0.0.0 MCPPORT=7071 ./scripts/install-systemd.sh --mcp
sudo systemctl enable --now inferwatch-mcpA separate unit from the collector on purpose: this process only reads the
database while the collector writes it, so restarting one does not interrupt
the other. It is ordered After=inferwatch.service so the database exists on a
first boot, but not bound to it — it keeps answering while the collector
restarts, its data simply stops advancing.
Bound beyond loopback the API key is required, but it crosses the wire in clear text. Restrict it at the firewall the same way as the dashboard, and put TLS in front if it ever leaves the machine.
Tool | Purpose |
| Ollama headline metrics and series |
| per-model breakdown; what is resident |
| Ollama per-request detail |
| cold loads, evictions, truncations, warnings |
| vLLM metrics, reachability, GPU attribution |
| per-device util/VRAM/temp/power |
| Ollama prompt-cache occupancy and eviction pressure, plus live KV usage |
| requests in flight at once, and how close to slot capacity |
| who is calling, for which models, at what context size |
| SwarmUI/ComfyUI throughput, durations, models, backends |
| what failed, by node class, and what never reached a backend |
| recent generations with every model the workflow loaded |
| what is monitored, and how it is configured |
| is collection working, is debug logging on |
| read-only SELECT escape hatch, with units |
Retention
Raw per-request rows (Ollama): 7 days (
retention.raw_days).GPU samples, events, prompt-cache samples, image samples and the image event timeline, vLLM rows: 30 days (
retention.sample_days).Image generations follow the raw window, being the image equivalent of a request row. There are no rollups for them, so past retention the answer is "not stored" and
complete/covers_fromsay so.rollup_1mandrollup_1h: kept indefinitely.
Rollups store fixed-bucket histograms of TTFT and latency, not pre-computed percentiles. Histograms add, so a percentile over any range is computed by summing buckets and walking to the target rank. Percentiles of percentiles would be meaningless; this is not.
Queries inside the raw window return exact percentiles; beyond it they come from
histograms and are reported as the upper bound of the containing bucket.
Every response carries exact: true|false.
Restarts and reboots
Resume state is per source — a journald cursor, a file inode+offset, or a docker timestamp — and is flushed on SIGTERM. Two things make that safe rather than merely likely:
Writes are idempotent. Every request and event row carries a dedupe_key
under a UNIQUE index, and inserts are INSERT OR IGNORE. Re-reading lines that
were already stored is a no-op, so ingest can be run repeatedly and a resume
can safely overlap.
A cursor that cannot be used is not trusted. If the journal a cursor points into was rotated away, journalctl silently repositions using the timestamp embedded in the cursor and resumes correctly. But a cursor stamped in the future (clock skew, a restored database) makes journalctl wait for entries that will not arrive, stalling collection silently; such a cursor is rejected on startup. A follow attempt yielding nothing twice running does the same.
systemctl stop completes in well under a second. systemd records
ExecMainStatus=15 alongside Result=success: uvicorn deliberately re-raises
the signal after shutting down, so exiting by SIGTERM is expected, not a crash.
Honest limitations
Attribution under parallelism (Ollama). The task id in llama.cpp's timing
lines and the access line's status never appear together, so they are joined by
arrival order. With one request in flight that is exact. When two finish before
either access line prints, nothing in the log disambiguates them — those rows are
stored as attribution='ambiguous' rather than guessed. Values: exact,
ambiguous, none (failed before reaching the runner — no model is guessed
either), orphan (timings with no access line). A 2-day backfill on the
development host scored 174 exact, 9 ambiguous, 4 orphan, with
OLLAMA_NUM_PARALLEL=1 in effect; expect a higher ambiguous share the more
requests run concurrently.
vLLM's token and request counters are not per-request aligned.
generation_tokens_total advances as tokens stream; request_success_total
advances only when a request completes. Over a short window they therefore
describe overlapping but different sets of requests, and dividing one by the
other does not give tokens per request. The API marks this with
counters_aligned: false, and the dashboard says so on the vLLM tab.
No per-request anything for vLLM. Covered above. If you need per-request detail from vLLM, its request-level logging is the only source, and it logs prompt text — which this tool deliberately never stores.
Some breakdowns reach back only as far as raw retention. Endpoint, status
code, client address and per-request identity exist only on raw request rows —
the rollups aggregate by (bucket, model, class) and carry none of them — so
by_endpoint, status_breakdown, by_client, recent_errors, slowest and
recent_requests have nothing to degrade to. Past retention the honest answer
is not stored, which is a different statement from nothing happened, so
those responses carry covers_from (the oldest surviving raw row) and
complete, and the dashboard's empty states name the cutoff instead of reading
as an idle window.
Note this bound is retention.raw_days (7 by default), not the 6-hour
RAW_WINDOW_S that decides when other queries switch to the rollups for
speed. A 7-day client breakdown is complete even though summary()['exact'] is
false for that span; conflating the two reports a full answer as a partial one.
A client's model mix is only as good as ollama's logging. Ollama names the
model on a per-request scheduler line, and when that line is absent the request
is counted with its model left null — reported as unattributed per client
rather than dropped or guessed. On this host the rate has ranged from 100% named
to 0% named on different days, so a client showing mostly unattributed is a
statement about the log, not about the client.
Image generations are counted from ComfyUI, never from SwarmUI's log.
The log describes the same work from the orchestrator's side, but a request for
N images produces N "Generated an image" lines and nothing ties either to a
prompt_id. Counting both would double-count; joining them would guess. So the
log is a timeline and the count comes from /history alone — which does mean a
generation driven through ComfyUI directly, bypassing SwarmUI, still appears
(correctly), while one that failed inside SwarmUI before reaching a backend
appears only as a log error.
ComfyUI's history is in memory, not on disk. It is polled, so a backend
restarting between polls loses whatever it had not yet reported. At the default
interval that is a few seconds' exposure; history_limit controls how much of
the ring is re-read each tick.
In-flight requests are measured; queued ones are not observable. Ollama
serves OLLAMA_NUM_PARALLEL requests per runner (an upper bound — an
architecture that cannot serve concurrently is loaded with one slot regardless,
so capacity is the summed -np across resident runners, not a configured
number). Concurrency is derived from overlapping [started_ts, ts] intervals,
so it is exact and catches a burst that begins and ends between two 5-second
samples.
What cannot be reported is a queue depth. A queued request is given no slot,
so it emits no runner lines until it starts; ollama logs no queue length; and
llama.cpp's own /metrics, which counts deferred requests, is not enabled —
ollama passes neither --metrics nor --slots to the runner. The observable
consequence of queueing is queue_ms on the request that waited, plus the time
spent with every slot busy (saturated_s). Silence on the concurrency panel is
therefore not evidence that nothing waited.
The Ollama prompt-cache gauge is sampled on ollama's schedule. A cache state line is logged only when ollama runs a cache update, so the series is
unevenly spaced and a window with no samples means "no cache updates happened",
not "the cache was empty". Its counters are therefore stored as deltas and
reported as totals, never as rates, and samples accompanies every figure.
Live KV occupancy exists only inside the raw window. It is computed from
context_tokens / n_ctx_slot on per-request rows. The rollups aggregate per
model and class rather than per slot, so beyond retention ctx_usage is null
rather than back-computed from an assumed context size. Rows written before the
n_ctx_slot column existed are skipped for the same reason.
Logs are the feed, not the archive. A journal may hold only a day or two
depending on journald.conf; the SQLite file is the historian. If logs rotate
faster than inferwatch runs, that gap is unrecoverable.
Two byte-identical events in the same microsecond collapse to one. The dedupe key for events is built from their values, so an identical warning logged twice within a microsecond keeps one row. A deliberate trade for guaranteed idempotency — dropping a repeated warning beats duplicating history.
Health-check traffic is separated, not counted. HEAD / and GET /api/ps
were 96% of requests on the development host. They are stored with
class='health' and excluded from inference rates unless
dashboard.include_health is on; requests_all always includes them.
Depends on log formats and metric names. Ollama's timing lines are debug
output, not a contract, and vLLM renames metrics between releases.
tests/test_parse.py holds verbatim fixture lines and tests/test_vllm.py a
real /metrics excerpt; if an upgrade breaks parsing, those tests fail and show
what changed.
Tests
python -m unittest discover -s tests -t .CI runs this on Python 3.10 through 3.14, plus a packaging job that builds the
wheel, asserts the dashboard HTML is inside it, and installs it into a clean
environment from an empty directory so the source tree cannot mask a packaging
mistake. No network, GPU or engine is required. Parser fixtures are verbatim real log lines and a
real /metrics excerpt. Coverage includes the correlator's join and
its ambiguous/orphan/failed cases, histogram percentiles and rollup idempotency,
the schema migration, signal-safe commits, cursor validation, file rotation and
truncation, timestamp monotonicity, counter-reset detection, and config
precedence and locking. The configuration reference in this file is generated
from the spec and a test fails if it drifts. The dashboard's JavaScript is
syntax-checked with a pure-Python parser and its formatters executed in a real
JS engine (both optional — no Node needed).
.venv/bin/python scripts/gen-config-docs.py --check # docs match the code?Layout
inferwatch/parse.py ollama log line parsers (pure, fixture-tested)
inferwatch/readers.py journald / file / docker log readers
inferwatch/collect.py correlator, GPU + model pollers, maintainer
inferwatch/vllm.py Prometheus scraper, delta and reset handling
inferwatch/parse_swarm.py SwarmUI log line parsers (pure, fixture-tested)
inferwatch/images.py SwarmUI/ComfyUI poller, history ingest, log collector
inferwatch/image_metrics.py image-generation query layer
inferwatch/gpuproc.py maps GPUs to the process tree holding them
inferwatch/vllm_metrics.py vLLM query layer
inferwatch/metrics.py ollama query layer (shared by API and MCP)
inferwatch/store.py SQLite schema, rollups, histograms, retention
inferwatch/config.py typed settings spec, precedence, source validation
inferwatch/supervisor.py builds and rebuilds collectors from the sources table
inferwatch/api.py FastAPI endpoints + SSE
inferwatch/web/index.html dashboard (single file, no CDN, no build step)
inferwatch/mcp_server.py MCP server (read-only)
inferwatch/mcp_auth.py API-key auth for the MCP HTTP transports
inferwatch/main.py serve / ingest / stats / sourcesContributing
Issues and pull requests are welcome. Two things make a change easy to accept:
python -m unittest discover -s tests -t .passes.If you touched
inferwatch/config.py, runpython scripts/gen-config-docs.pyso the README's configuration reference matches the code — a test enforces it.
Parser changes should come with a fixture line copied verbatim from real engine output, the way the existing tests do. Log formats and metric names are not contracts, and a real fixture is what makes a future break obvious.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Analytics for MCP servers. Query your tool calls, first-call success, retries and schema cost.
Analytics for MCP servers. Find out which of your tools agents get wrong. MCPulse shows you which tools AI agents retry, which come back empty, and which they never call at all. Two lines inside your own server. It never sees your arguments or your results. getmcpulse.com
Query application logs, traces, and metrics from your AI coding assistant via Foam's MCP server.
- SpanlyOAuthcom.spanly
MCP observability. Query live traffic, errors, duration, and alerts from your AI agent.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceGives AI agents real-time access to system metrics, process management, and container orchestration.-
- AlicenseNot gradedqualityDmaintenanceProvides AI assistants with direct access to Red Hat OpenShift AI observability data, enabling querying of Prometheus metrics, Alertmanager alerts, Loki logs, Grafana dashboards, and Kubernetes cluster state to troubleshoot vLLM inference workloads.5MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to monitor real-time CPU, RAM, and disk usage on the local machine.-
- FlicenseNot gradedqualityCmaintenanceEnables LLMs to monitor server health through natural language, including HTTP endpoint checks, disk/memory/CPU usage, port status, and process listings.-