FABRIC Infrastructure Metrics MCP Server
OfficialREADME.md
# FABRIC Infrastructure Metrics MCP Server
[](https://www.python.org/downloads/)
[](./LICENSE)
[](https://modelcontextprotocol.io/)
A **[Model Context Protocol (MCP)](https://modelcontextprotocol.io/)** server that lets LLM clients (Claude Desktop, Claude Code, VS Code Copilot, Chatbox, etc.) **ask natural-language questions about** the infrastructure metrics of the **[FABRIC Testbed](https://fabric-testbed.net)** — a nation-wide programmable network research infrastructure.
### Key features
- **Grounded by construction** — every tool runs a **fixed PromQL query authored in a FABRIC Grafana dashboard**. The model never writes PromQL; it picks a tool, a variant, and parameters.
- **17 tools** — 14 metric topics (CPU, memory, disk, filesystem, network, sockets, netstat, system, time, temperature, link traffic, link status, port traffic) plus 3 discovery tools.
- **Two modes** — a local stdio process using your own CILogon session, or one served deployment holding a Grafana service-account token.
- **Friendly targets** — `amst` means every node at that site, `cern worker 2` resolves to `cern-w2`; the server maps them to real instances.
- **Bounded responses** — series, samples, and total points are capped and downsampled to fit an LLM context without losing the time window.
- **Production-ready** — CredMgr token verification, per-caller rate limiting, structured access logs, Prometheus metrics, and a startup credential probe.
---
## Table of contents
| Getting started | Reference | Operations |
|:---|:---|:---|
| [Quick install](#quick-install) | [Tools reference](#tools-reference) | [Deploy with Docker Compose (Server Mode)](#deploy-with-docker-compose-server-mode) |
| [MCP client configuration](#mcp-client-configuration) | [Variants, parameters & results](#variants-parameters--results) | [Environment variables](#environment-variables) |
| [Local mode setup](#local-mode-setup) | [Quick tool examples](#quick-tool-examples) | [Testing](#testing) |
| [Remote mode setup](#remote-mode-setup) | [Architecture](#architecture) / [Repo layout](#repo-layout) | [Logging](#logging) / [Monitoring & metrics](#monitoring--metrics-server-mode-only) |
| [Local vs Remote — which to use?](#local-vs-remote--which-to-use) | [Rebuilding the catalog](#rebuilding-the-catalog) | [Security notes](#security-notes) |
---
## Quick install
> **Prerequisites:** Python 3.11+, a FABRIC token from [portal.fabric-testbed.net](https://portal.fabric-testbed.net) (Experiments → Manage Tokens), and — for local mode only — a browser for the one-time Grafana login.
```bash
# Remote mode: talk to the deployed server (no Grafana credential needed)
curl -fsSL https://raw.githubusercontent.com/fabric-testbed/fabric_metrics_mcp/main/install.sh | bash -s -- --remote
# Local mode: run the server yourself against your own Grafana session
curl -fsSL https://raw.githubusercontent.com/fabric-testbed/fabric_metrics_mcp/main/install.sh | bash -s -- --local
```
The installer:
1. Creates a virtual environment and installs the package.
2. Writes `run_stdio.sh` (local) or `run_remote.sh` (remote) with your paths baked in.
3. Prints the MCP client configuration to paste.
4. Runs a smoke test against the server.
> **Note:** Remote mode needs nothing but a FABRIC token — the deployed server holds its own Grafana credential. Prefer it unless you are developing the server itself.
> **Manual setup:** see [Local mode setup](#local-mode-setup) or [Remote mode setup](#remote-mode-setup).
---
## MCP client configuration
Both helper scripts speak stdio, so any MCP client can launch them:
- `run_stdio.sh` — runs the server locally
- `run_remote.sh` — bridges stdio to the deployed HTTPS server, injecting your token
> **Important:** keep the server name **short**. The model sees `mcp__<server-name>__<tool>`, and OpenAI-compatible stacks reject function names over 64 characters. `fabric_metrics_list_param_options` is already 33; a long server name silently drops the longest tools.
#### Claude Code CLI
```bash
claude mcp add fabric-metrics -- /path/to/run_remote.sh
```
#### Claude Desktop
```json
{
"mcpServers": {
"fabric-metrics": {
"command": "/path/to/run_remote.sh"
}
}
}
```
#### VS Code
```json
{
"servers": {
"fabric-metrics": {
"type": "stdio",
"command": "/path/to/run_remote.sh"
}
}
}
```
#### Chatbox
```json
{
"name": "fabric-metrics",
"command": "/path/to/run_remote.sh",
"env": {
"FABRIC_TOKEN_LOCATION": "/path/to/id_token.json",
"FABRIC_MCP_URL": "https://<your-host>/mcp"
}
}
```
---
## Tools reference
Every topic tool takes `variant` (which measurement), the dashboard's parameters, and an optional time window. Call `fabric_metrics_list_topics()` to route a question.
### Node metrics
| Tool | Description |
|:-----|:------------|
| `fabric_metrics_cpu` | CPU utilization, load average, core count *(9 variants)* |
| `fabric_metrics_memory` | RAM/swap used, paging, page faults, **OOM kills** *(12 variants)* |
| `fabric_metrics_memory_internals` | Slab, vmalloc, hugepages, writeback, NFS *(13 variants)* |
| `fabric_metrics_disk` | Throughput, IOPS, latency, %busy, queue depth *(8 variants)* |
| `fabric_metrics_filesystem` | Free space and inodes *(3 variants)* |
| `fabric_metrics_network` | Bandwidth, packets, errors, drops, softnet, conntrack *(16 variants)* |
| `fabric_metrics_sockets` | TCP/UDP/RAW socket counts *(5 variants)* |
| `fabric_metrics_netstat` | Established connections, TCP errors, ICMP, retransmits *(11 variants)* |
| `fabric_metrics_system` | Uptime, forks, context switches, entropy, systemd *(11 variants)* |
| `fabric_metrics_time` | Clock synchronization and NTP drift *(4 variants)* |
| `fabric_metrics_temperature` | Hottest nodes, above-cutoff, by rack *(3 variants)* |
### Switch and link metrics
| Tool | Description |
|:-----|:------------|
| `fabric_metrics_link_traffic` | Rack-to-rack dataplane throughput *(1 variant)* |
| `fabric_metrics_link_status` | Busiest/idlest links, per-link rates, alerts *(6 variants)* |
| `fabric_metrics_port_traffic` | Per switch-port bandwidth *(1 variant)* |
### Discovery
| Tool | Description |
|:-----|:------------|
| `fabric_metrics_list_topics` | Every tool, its variants, source dashboard, and parameters |
| `fabric_metrics_list_param_options` | Allowed values for a parameter, resolved live |
| `fabric_metrics_list_sites` | FABRIC site/rack codes that have node data |
> **Note:** `variant` is an **argument**, never a tool name. There is no `oom_killer` tool — it is `fabric_metrics_memory(variant="oom_killer", ...)`.
---
## Authentication
There are **two independent boundaries**. The caller's token is never forwarded to Grafana.
| Boundary | How it works |
|:---|:---|
| **Caller → MCP server** | Server mode reads the `Authorization: Bearer` header **only** and verifies the signature against CredMgr's JWKS. Local mode reads `$FABRIC_TOKEN_LOCATION`. |
| **MCP server → Grafana** | A Grafana service-account token (`GRAFANA_TOKEN`), or a replayed CILogon browser session for local development. |
**Get a token:** [portal.fabric-testbed.net](https://portal.fabric-testbed.net) → Experiments → Manage Tokens, or `fabric-cli tokens create`. Identity tokens last about four hours.
> **Important:** this server *terminates* authentication rather than proxying it — it queries Grafana with a credential broader than any caller's. `FABRIC_VERIFY_TOKENS=1` is therefore mandatory in any reachable deployment; an unverified JWT payload is just base64.
---
## Architecture
```
"How busy is the CPU on cern-w5?"
│
▼ LLM picks tool + variant + params
fabric_metrics_cpu(variant="busy", node="cern-w5")
│
▼ HTTPS + Authorization: Bearer <FABRIC token>
┌─────────────────────────────────────────────────┐
│ nginx (TLS) → fabric-metrics-mcp │
│ ├─ verify token (CredMgr) │
│ ├─ resolve node → instance │
│ └─ substitute into the panel's│
│ FIXED PromQL │
└─────────────────────────────────────────────────┘
│
▼ Authorization: Bearer <Grafana service token>
Grafana → datasource proxy → Mimir / Prometheus
│
▼
→ 0.8 % (grounded, real data)
```
- The catalog of dashboards, panels, and their queries ships **inside the package**, so the running server has no dependency on the repo layout or on Grafana's dashboard API.
- Responses are capped at 60 series and 3000 total samples; beyond that the step widens to fit while keeping the full window.
---
## Repo layout
```
fabric_metrics_mcp/
__main__.py # entry point: transport, middleware, prompt registration
config.py # all configuration, read once from the environment
registry.py # topic -> variant -> panel, loaded from the shipped catalog
query.py # time parsing, substitution, Grafana calls, response budget
tools.py # one tool per topic + discovery tools
session.py # the SERVER's Grafana credential (token or cookies)
system.md # served as the `fabric-metrics-system` MCP prompt
auth/resolver.py # the CALLER's FABRIC token
middleware/ # access log, rate limiting
data/annotated/ # per-dashboard panels: fixed PromQL, units, legends
data/groups/ # the topic -> variant taxonomy
deploy/
central.env.example # one config file for the served deployment
local.env.example # one config file for a local Docker run
deploy.sh # mode selector + preflight, wraps docker compose
DEPLOY.md # step-by-step deployment runbook
docker-compose.yml # central: mcp-server + optional nginx (tls profile)
nginx/default.conf # TLS termination, bearer-only /mcp, restricted /metrics
tests/ # offline unit tests (no network, no credentials)
step2_build_catalog.py # rebuild: scrape dashboards
step3_annotate.py # rebuild: authored descriptions
step3_5_group.py # rebuild: topic/variant taxonomy
test_server.py # end-to-end smoke test + LLM evaluation harness
```
---
## Environment variables
Configuration is read once at startup. Deployment settings live in `deploy/central.env` (see [Deploy with Docker Compose](#deploy-with-docker-compose-server-mode)).
| Var | Default | Purpose |
|-----|---------|---------|
| `FABRIC_LOCAL_MODE` | `0` | `1` = stdio + token file; `0` = HTTP + bearer header |
| `FABRIC_MCP_TRANSPORT` | `stdio` if local, else `http` | FastMCP transport |
| `FABRIC_VERIFY_TOKENS` | `0` if local, else `1` | Verify caller signatures against CredMgr's JWKS |
| `FABRIC_CREDMGR_HOST` | `cm.fabric-testbed.net` | JWKS host |
| `FABRIC_TOKEN_AUDIENCE` | *empty* | Expected `aud` claim; unchecked when unset |
| `FABRIC_TOKEN_LOCATION` | *empty* | Local mode only: path to the caller's token file |
| `GRAFANA_TOKEN` | *empty* | Grafana service-account token. **Required for a served deployment** |
| `GRAFANA_TOKEN_FILE` | *empty* | Path to a file holding that token; preferred over inline |
| `GRAFANA_STATE_FILE` | `./grafana_state.json` | Replayed CILogon session, used only when no token is set |
| `FABRIC_METRICS_BASE_URL` | `https://infrastructure-metrics.fabric-testbed.net/grafana` | Grafana API base. Use Grafana's internal address when deployed alongside it |
| `FABRIC_METRICS_DS_UID` | `P83FD23C85A64357C` | Mimir datasource UID |
| `FABRIC_METRICS_TOOL_PREFIX` | `fabric_metrics_` | Prefix on registered tool names; `""` for bare names |
| `HOST` / `PORT` | `0.0.0.0` / `8000` | Listen address inside the container |
| `FORWARDED_ALLOW_IPS` | *empty* | Reverse proxy whose `X-Forwarded-*` headers to trust. **Required behind nginx**, or the app emits `http://` redirects and clients lose their session. Never `*` |
| `RATE_LIMIT` | `60/minute` | Fixed-window limit per caller |
| `RATE_LIMIT_ENABLED` | `0` if local, else `1` | Toggle the limiter |
| `RATE_LIMIT_TRUSTED_PROXIES` | *empty* | Proxies allowed to assert the real client via `X-Real-IP`, as `/32`. See [Security notes](#security-notes) |
| `LOG_LEVEL` / `LOG_FORMAT` | `INFO` / `text` | `LOG_FORMAT=json` for structured logs |
| `METRICS_ENABLED` | `0` if local, else `1` | Expose `/metrics` for Prometheus |
| `METRICS_CLIENT_IP_LABELS` | `0` | Per-IP Prometheus series; unbounded on a public endpoint |
Deployment-only variables, consumed by `docker-compose.yml` rather than the app:
| Var | Default | Purpose |
|-----|---------|---------|
| `MCP_BIND` / `MCP_PORT` | `127.0.0.1` / `8000` | Host-side publish address for the app |
| `NGINX_BIND` | `0.0.0.0` | Address nginx listens on. Set to the public IP on a multi-homed host |
| `HTTP_PORT` / `HTTPS_PORT` | `80` / `443` | Host-side nginx ports |
| `SSL_CERT` / `SSL_KEY` | `./ssl/fullchain.pem`, `./ssl/privkey.pem` | Certificate and key; any path, any filename |
| `PUBLIC_HOSTNAME` | *empty* | Checked against the certificate's SANs by the preflight |
> **Note:** the server also publishes its operating instructions as an MCP prompt named `fabric-metrics-system`.
---
## Deploy with Docker Compose (Server Mode)
The full runbook — including network prerequisites and troubleshooting — is in [`DEPLOY.md`](./DEPLOY.md). Summary:
### Step 1: Confirm the Grafana credential works from this host
```bash
export GRAFANA_TOKEN='glsa_...'
curl -s -o /dev/null -w '%{http_code}\n' \
-H "Authorization: Bearer $GRAFANA_TOKEN" \
https://infrastructure-metrics.fabric-testbed.net/grafana/api/user
```
`200` means the token reaches Grafana. A `302` to `/login` means an auth proxy intercepts it — point `FABRIC_METRICS_BASE_URL` at Grafana's internal address instead.
### Step 2: Place your TLS certificates
```bash
sudo mkdir -p /etc/ssl/fabric-metrics
# copy the certificate (FULL CHAIN, not the leaf alone) and key into it
sudo chmod 600 /etc/ssl/fabric-metrics/privkey.pem
```
> **Important:** ask your CA for the **full chain**. Browsers fetch missing intermediates themselves, so a leaf-only certificate looks fine in a browser and fails every command-line client.
### Step 3: Create the `deploy/central.env` file
```bash
cp deploy/central.env.example deploy/central.env
chmod 600 deploy/central.env
```
At minimum set:
```ini
GRAFANA_TOKEN=glsa_...
FABRIC_METRICS_BASE_URL=<whatever Step 1 proved>
PUBLIC_HOSTNAME=<your hostname>
NGINX_BIND=<your public IP>
SSL_CERT=/etc/ssl/fabric-metrics/fullchain.pem
SSL_KEY=/etc/ssl/fabric-metrics/privkey.pem
FORWARDED_ALLOW_IPS=172.32.240.10
RATE_LIMIT_TRUSTED_PROXIES=172.32.240.10/32
```
### Step 4: Start without TLS and test through a tunnel
```bash
./deploy.sh central up -d --build
./deploy.sh central logs --tail 40
```
Look for `Grafana credential OK: authenticated to ... as '<service-account>'`. Then from your workstation:
```bash
ssh -N -L 8000:127.0.0.1:8000 <user>@<host> # leave running
python test_server.py --http http://127.0.0.1:8000/mcp
```
This proves the app and the Grafana credential work with no certificate involved.
### Step 5: Start the services with TLS
```bash
./deploy.sh central --profile tls up -d
```
Two containers run:
- **`fabric-metrics-mcp`** — the MCP server, published on loopback only
- **`fabric-metrics-nginx`** — TLS termination, reverse proxy to `:8000`
The preflight prints the certificate's subject and expiry, refuses a mismatched certificate/key pair, and warns if `PUBLIC_HOSTNAME` is absent from the SANs.
### Step 6: Verify
```bash
./deploy.sh central --profile tls ps # both "running", app "(healthy)"
dig +short <your-hostname> # must resolve to this host
curl -s -o /dev/null -w '%{http_code}\n' --max-time 10 https://<your-host>/mcp
# 406 = up (the MCP endpoint rejects a plain GET). Never use -I here.
export FABRIC_TOKEN_LOCATION=$HOME/id_token.json
python test_server.py --http https://<your-host>/mcp
```
| Service | Port | Exposure |
|--------|------|----------|
| `fabric-metrics-mcp` | 8000 | loopback only |
| `fabric-metrics-nginx` | 80, 443 | `NGINX_BIND` |
| `/metrics` | via nginx | restricted to the monitoring network |
### Configuration files
- [`deploy/central.env.example`](./deploy/central.env.example) — every setting, one file
- [`docker-compose.yml`](./docker-compose.yml) — services; nothing hardcoded
- [`nginx/default.conf`](./nginx/default.conf) — TLS and proxy rules
- [`deploy.sh`](./deploy.sh) — mode selector, preflight, passes everything else to `docker compose`
NGINX requirements, if you front this with your own proxy:
- `proxy_set_header Authorization $http_authorization` — the caller's token must arrive untouched
- `proxy_buffering off` and a long `proxy_read_timeout` — MCP streams responses
- `proxy_set_header X-Real-IP $remote_addr` — `$remote_addr` overwrites; `X-Forwarded-For` appends and must never key a rate limit
- **No `auth_request`** on `/mcp` — it is bearer-token-only, verified by the app
---
## Rebuilding the catalog
The shipped catalog is a snapshot of five Grafana dashboards. Rebuild it when FABRIC changes a dashboard:
```bash
pip install -e '.[build]'
python step2_build_catalog.py # scrape dashboards -> catalog/
python step3_annotate.py # authored descriptions -> data/annotated/
python step3_5_group.py # topic/variant taxonomy -> data/groups/
```
- `step2` extracts panels, queries, units, legends, and template variables for the five featured dashboards, keyed by UID.
- `step3` holds hand-authored descriptions; these are the source of truth for routing and are what the model reads.
- `step3_5` assigns every panel to exactly one `(topic, variant)` and fails loudly on duplicates or omissions.
> **Note:** panel IDs are referenced by number, so a deleted panel breaks the rebuild with a `KeyError` rather than silently vanishing.
---
## Local mode setup
Runs the server on your machine against **your own** Grafana session, so it sees exactly what your account sees. No service token.
> **Quick install:** `curl -fsSL .../install.sh | bash -s -- --local`
### Step 1: Create a Python virtual environment
```bash
python3 -m venv .venv && source .venv/bin/activate
```
### Step 2: Install the package
```bash
pip install -e '.[local]'
```
### Step 3: Capture your Grafana session
```bash
python get_grafana_session.py
```
Opens a browser, waits for you to complete CILogon, and saves `grafana_state.json`. It expires in hours to days.
### Step 4: Get your FABRIC token
Download `id_token.json` from the portal, or `fabric-cli tokens create`.
### Step 5: Configure the script
```bash
export FABRIC_TOKEN_LOCATION=$HOME/id_token.json
```
| Var | Default | Purpose |
|-----|---------|---------|
| `FABRIC_TOKEN_LOCATION` | `./id_token.json` | Your FABRIC token |
| `GRAFANA_STATE_FILE` | `./grafana_state.json` | Your saved Grafana session |
| `FABRIC_LOCAL_MODE` | set to `1` by `run_stdio.sh` | stdio transport, token from file |
### Step 6: Test
```bash
python test_server.py
```
### Step 7: Configure your MCP client
Point it at `run_stdio.sh` — see [MCP client configuration](#mcp-client-configuration).
---
## Remote mode setup
Talks to the deployed server over HTTPS. Nothing runs locally except a stdio bridge.
> **Quick install:** `curl -fsSL .../install.sh | bash -s -- --remote`
### Step 1: Install prerequisites
```bash
# node (for npx mcp-remote) and jq
brew install node jq # or: apt install nodejs npm jq
```
### Step 2: Set up the venv
Only needed to run `test_server.py`; MCP clients need just the bridge script.
```bash
python3 -m venv .venv && source .venv/bin/activate && pip install -e .
```
### Step 3: Create your token
Download `id_token.json` from the portal, or `fabric-cli tokens create`.
### Step 4: Get the helper script
`run_remote.sh` from this repo. It reads your token, checks it has not expired, and bridges stdio to HTTPS.
### Step 5: Configure the script
```bash
export FABRIC_TOKEN_LOCATION=$HOME/id_token.json
export FABRIC_MCP_URL=https://<your-host>/mcp
```
| Var | Default | Purpose |
|-----|---------|---------|
| `FABRIC_TOKEN_LOCATION` | `$PWD/id_token.json` | Your FABRIC token. Use an **absolute** path |
| `FABRIC_MCP_URL` | the FABRIC deployment | Server endpoint, **no trailing slash** |
> **Important:** use an absolute path for `FABRIC_TOKEN_LOCATION`. A relative `./id_token.json` resolves against whatever directory your MCP client happens to start in.
### Step 6: Test
```bash
python test_server.py --http https://<your-host>/mcp
```
### Step 7: Configure your MCP client
Point it at `run_remote.sh` — see [MCP client configuration](#mcp-client-configuration).
---
## Local vs Remote — which to use?
| | Local mode | Remote mode |
|:---|:---|:---|
| **Grafana credential** | your own CILogon session | the server's service token |
| **Expires** | hours to days | not your problem |
| **Needs a browser** | yes, once per session | no |
| **Transport** | stdio | HTTPS + bearer |
| **Sees** | what your account sees | what the service account sees |
| **Good for** | developing the server | everyday use |
> **Recommendation:** use remote mode unless you are changing the server itself. It needs only a FABRIC token, and the session never expires out from under you.
---
## Variants, parameters & results
### Naming the target
`node` accepts what a person would say; the server resolves it. **Pass the user's target as-is.**
| User says | Pass | Resolves to |
|:---|:---|:---|
| "at amst" | `amst` | **all** nodes at that site |
| "amst worker 2" | `amst worker 2` | that one worker |
| "the head node at uky" | `uky head node` | that site's head node |
| a full instance | as given | that instance |
Head node is `hn` — one per site; there is no `h1`. Workers are `w1`…`wN`. Site codes are opaque: present them as-is rather than expanding them into institution names.
### Time ranges
`start` / `end` take natural language, like Grafana's time picker: `"past 5 minutes"`, `"last 30 min"`, `"now-1h"`, `"2 hours ago"`, `"today 6am"`, `"2026-01-13 15:00"`.
- No time given → last 5 minutes. Gauge panels evaluate *at* `end`, so they return the current value regardless.
- `step` is auto-chosen for ~120 points, never finer than the 30 s scrape interval.
- `mode` (`instant` | `range`) overrides the panel's default query type.
### Reading results
Each result carries `panel`, `unit`, `description`, `resolved_queries`, `params_used`, and `series`. Each series has a `name` from the dashboard's legend.
- Honour `unit`. A `percent` panel is already a percentage.
- A range series carries `values`, **not** timestamped pairs: `values[i]` is the sample at `window.start + i * window.step`, and `null` is a gap — not zero.
- `truncated: true` means series were dropped (>60). `downsampled: true` means the step widened to fit the budget; the window is unchanged.
- Zero series is a real answer, not a reason to invent one.
- `error: "invalid_param_value"` means a fixed-option parameter got a value it can never match. The message names the valid values.
---
## Quick tool examples
**CPU utilization on one node**
```jsonc
{ "name": "fabric_metrics_cpu",
"arguments": { "variant": "busy", "node": "amst-w1" } }
```
**Compare memory across a whole site**
```jsonc
{ "name": "fabric_metrics_memory",
"arguments": { "variant": "ram_used", "node": "cern" } }
```
**The three hottest nodes in the fleet**
```jsonc
{ "name": "fabric_metrics_temperature",
"arguments": { "variant": "hottest", "params": { "toplimit": "3" } } }
```
**Routing hints for confusable requests**
| Question | Tool and variant |
|:---|:---|
| OOM kills, page faults | `memory` — *not* `memory_internals`, despite the name |
| slab, vmalloc, hugepages, writeback | `memory_internals` |
| softnet / netdev budget | `network`, variants `softnet_packets` / `softnet_quota` |
| "how busy is the disk" | `disk` variant `io_util`; "how slow" → `wait_time` |
| "hottest node" | `temperature` variant `hottest` — the one real fleet-wide ranking |
---
## System prompt
The server publishes its operating instructions as an MCP prompt, so a client can load them without hardcoding anything:
```python
@mcp.prompt(name="fabric-metrics-system")
def fabric_metrics_system_prompt():
return SYSTEM_TEXT
```
Tool names inside it track `FABRIC_METRICS_TOOL_PREFIX`, so the served prompt never goes stale.
---
## Logging
* Every HTTP request is logged with method, path, status, duration, a request id, and the caller's identity.
* Every tool call logs start, completion with duration, sanitised parameters, and errors.
* `LOG_FORMAT=json` emits one JSON object per line for ingestion.
* Tokens are never logged — every helper is redacted by construction.
```json
{"ts": "2026-09-16T15:45:19+0000", "level": "INFO", "logger": "fabric_metrics.access",
"msg": "POST /mcp -> 200 in 1.98ms", "request_id": "f4ea4c72f16f",
"user_email": "user@example.edu", "client_ip": "203.0.113.7", "duration_ms": 1.98}
```
> **Important:** with a shared Grafana service account, Grafana's own logs cannot tell you which person ran which query. These access logs are the only per-user record — ship them somewhere durable.
---
## Testing
```bash
pip install -e '.[test]'
pytest # offline unit tests, no network or credentials
```
The suite covers the pure layers where a wrong answer is easiest to hide: PromQL substitution (including the regex and `$__all` forms that must keep working), the parameter-injection guard, time-window parsing, the request timeout, the startup credential probe, and catalog loading.
End-to-end, against a real server:
```bash
python test_server.py # local stdio
python test_server.py --http https://<host>/mcp # a deployment
python test_server.py --eval --limit 20 # drive an LLM over test_questions.md
```
`--eval` needs an OpenAI-compatible endpoint (`FABRIC_LLM_BASE_URL`, `FABRIC_LLM_API_KEY`, `FABRIC_LLM_MODEL`) and writes `report.html` grading which tool and variant the model chose.
---
## Monitoring & Metrics (Server Mode Only)
```
MCP server ──/metrics──▶ Prometheus ──▶ Grafana
(:8000) (scrape) (dashboards)
```
- Enabled by default in server mode; needs the `[metrics]` extra.
- `/metrics` is restricted to the monitoring network by `nginx/default.conf`.
| Metric | Type | Labels | Description |
|--------|------|--------|-------------|
| `mcp_requests_total` | counter | `method`, `path`, `status` | HTTP requests |
| `mcp_request_duration_seconds` | histogram | `method`, `path` | Request latency |
| `mcp_tool_calls_total` | counter | `tool`, `status` | Tool invocations |
| `mcp_tool_duration_seconds` | histogram | `tool` | Tool latency |
| `mcp_rate_limit_hits_total` | counter | `key_type` | Rejected by the limiter |
### Disabling metrics
```ini
METRICS_ENABLED=0
```
### Production considerations
- **Per-IP labels are off by default** — one Prometheus series per source address is unbounded on a public endpoint. `METRICS_CLIENT_IP_LABELS=1` opts in deliberately.
- **Rate-limit state is per-process and in memory** — fine for one container; a multi-replica deployment needs a shared store, and two replicas double the effective limit.
- **The healthcheck is a TCP probe** — it cannot detect an expired Grafana credential. Watch for `Grafana credential NOT working` in the logs.
---
## Security notes
* **The caller's token is never forwarded.** This server queries Grafana with its own credential, so it terminates authentication rather than proxying it.
* **`FABRIC_VERIFY_TOKENS=1` is mandatory when reachable.** Without CredMgr verification the gate is decorative — a JWT payload is base64 anyone can write.
* **Server mode never falls back to a local token file.** A served process must not answer someone else's request with the operator's credential.
* **Parameters cannot escape their label matcher.** Values carrying a quote or backslash are rejected before substitution; regex metacharacters stay legal because dashboards legitimately use them.
* **Every Grafana request has a finite timeout.** Without one a hung upstream pins a worker thread until the process restarts.
* **There is no per-user authorization.** Every authenticated FABRIC user sees all sites' metrics — the same access they already have by logging into the dashboards.
### Rate limiting behind a proxy
The limiter's key is its bucket, so it may only come from inputs a caller cannot forge — anything a caller controls is something they can rotate for a fresh bucket per request. But it must still distinguish callers: behind a reverse proxy the socket peer is the proxy for every request, so keying on it alone puts everyone in one bucket and turns `RATE_LIMIT` into a service-wide cap.
`RATE_LIMIT_TRUSTED_PROXIES` resolves this by naming the proxies allowed to assert the real client address:
1. List **only** your reverse proxy, as a `/32`. Any host inside a wider range could forge the header.
2. Use `X-Real-IP`, which nginx sets from `$remote_addr` and therefore overwrites. Never key on `X-Forwarded-For` — nginx *appends* to whatever the client sent, so its left-most entry stays caller-controlled even on a trusted hop.
3. Leave it empty when clients reach the server directly.
#### Per-user limiting: what it would take
Keying on a verified JWT subject would give true per-user limits regardless of address. It is not enabled because:
1. The Starlette helper that reads claims at the edge has no verifier hook, so claims there are unverified.
2. An unverified claim is attacker input, not identity — a payload can be hand-written with no signing key.
3. Verification happens in the tool layer, after the limiter has already run.
4. Address keying at the edge is the safe default; the cost is that callers behind one NAT share a bucket.
---
## License
[MIT](./LICENSE)
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues