MCP Prometheus
๏ปฟ# MCP Prometheus + Loki ๐
Prometheus + Loki ๊ธฐ๋ฐ ๋ชจ๋ํฐ๋ง์ฉ MCP ์๋ฒ์
๋๋ค.
์ํธ๋ฆฌํฌ์ธํธ๋ `main.py`์
๋๋ค.
## Quick Start ๐
```powershell
cd d:\MCPTools
uv sync
uv run python mcp_prometheus/main.py
```
## ํ๋ก์ ํธ ๊ตฌ์กฐ ๐งฉ
```text
mcp_prometheus/
main.py
core/
config.py
runtime.py
server.py
time_utils.py
domain/
checks.py
infra/
loki_client.py
prom_client.py
tools/
catalog.py
alerts_runner.py
checks_runner.py
loki_query.py
promql.py
utils/
query_utils.py
summarize.py
```
## Tools ์์ฝ ๐ ๏ธ
| Tool | ๋ชฉ์ | ๋น๊ณ |
|---|---|---|
| `list_checks` | ๋ฑ๋ก๋ ์ฒดํฌ ๋ชฉ๋ก ์กฐํ | `id`, `name`, `description` ๋ฐํ |
| `list_environments` | ํ๊ฒฝ๋ณ Prometheus URL ์กฐํ | `prod/dev_test/dr` |
| `list_servers` | ์ต๊ทผ up ๊ธฐ์ค ์๋ฒ ๋ชฉ๋ก ์กฐํ | `(instance, job)` ๊ธฐ์ค ์ค๋ณต ์ ๊ฑฐ |
| `list_process_groups` | ํ๋ก์ธ์ค ๊ทธ๋ฃน ๋ชฉ๋ก ์กฐํ | `process_monitoring` ๊ธฐ์ค |
| `list_loki_environments` | ํ๊ฒฝ๋ณ Loki URL ์กฐํ | `prod/dev_test` |
| `list_loki_hosts` | ์ต๊ทผ ๋ก๊ทธ ๊ธฐ์ค host ํ๋ณด ์กฐํ | ๊ธฐ๋ณธ ์ต๊ทผ 1์๊ฐ |
| `list_loki_apps` | ์ต๊ทผ ๋ก๊ทธ ๊ธฐ์ค app ํ๋ณด ์กฐํ | ๊ธฐ๋ณธ ์ต๊ทผ 1์๊ฐ |
| `find_logs` | ๊ตฌ์กฐํ๋ Loki ๋ก๊ทธ ์กฐํ | `loki_environment`, `log_env`, `host`, `app` ํ์ |
| `get_alerts` | Prometheus ํ์ฑ Alert ์กฐํ | `/api/v1/alerts` ๊ธฐ๋ฐ, ๋ผ๋ฒจ/์ํ ํํฐ ์ง์ |
| `run_check` | ๋จ์ผ ์ฒดํฌ ์คํ | ๊ธฐ๋ณธ ๊ถ์ฅ |
| `run_all_checks` | ์ ์ฒด ์ฒดํฌ ๋ณ๋ ฌ ์คํ | `step=5m` ๊ณ ์ |
| `run_promql` | ์ฌ์ฉ์ PromQL ์ง์ ์คํ | `approved=True` ํ์ |
## Loki Tool ์
๋ ฅ ๊ฐ์ด๋ ๐ชต
### ํ๊ฒฝ ๊ตฌ๋ถ
- `loki_environment`: ์ด๋ค Loki ์๋ฒ๋ก ๋ถ์์ง ์ ํ (`prod`, `dev_test`)
- `log_env`: Loki ๋ก๊ทธ ๋ผ๋ฒจ `env` ๊ฐ (`prod`, `DEV`, `TEST`)
์ด ๋์ ๊ฐ์ ์๋ฏธ๊ฐ ์๋๋๋ค.
์๋ฅผ ๋ค์ด `dev_test` Loki ์๋ฒ ์์ `DEV`์ `TEST` ๋ก๊ทธ๊ฐ ํจ๊ป ์์ ์ ์์ผ๋ฏ๋ก, `find_logs` ํธ์ถ ์ `log_env`๋ ๋ช
์์ ์ผ๋ก ๋ฃ์ด์ผ ํฉ๋๋ค.
### Discovery Tool
- `list_loki_hosts`
- `list_loki_apps`
๊ณตํต ๊ท์น:
- ๊ธฐ๋ณธ ๊ธฐ๊ฐ์ ์ต๊ทผ 1์๊ฐ
- ์ ๋ ์๊ฐ ์กฐํ ์ `start_time_utc_iso`, `end_time_utc_iso` ์ฌ์ฉ
- ๊ฒฐ๊ณผ๋ ์ค๋ณต ์ ๊ฑฐ๋ ํ๋ณด ๋ชฉ๋ก ๋ฐํ
### `find_logs`
ํ์:
- `loki_environment`
- `log_env`
- `host`
- `app`
์ ํ:
- ๊ธฐ๊ฐ: `hours`, `minutes`, `days`
- ์ ๋ ์๊ฐ: `start_time_utc_iso`, `end_time_utc_iso`
- ์ข
๋ฃ ์คํ์
: `end_offset_minutes`, `end_offset_hours`, `end_offset_days`
- ํํฐ: `contains`, `level`
- ๊ฐ์ ์ ํ: `limit`
์๋ต:
- ์์ฑ๋ LogQL
- UTC ๋ฒ์
- `line_count`
- `logs[]` (`timestamp`, `timestamp_jakarta`, `labels`, `line`)
## `run_check` ์
๋ ฅ ๊ฐ์ด๋ ๐งญ
### ํ์
- `check_id`
### ๊ธฐ๊ฐ
- ์๋: `hours`, `minutes`, `days`
- ์ ๋: `start_time_utc_iso`, `end_time_utc_iso`
- ์ข
๋ฃ ์คํ์
: `end_offset_minutes`, `end_offset_hours`, `end_offset_days`
### ํ๊ฒ ํํฐ
- `server_name`
- `instance` (์: `host-or-ip:9100`)
ํํฐ ๊ท์น:
- `server_name`์ `instance`๋ฅผ ํจ๊ป ์ฃผ๋ฉด AND ์ ์ฉ
- ํ๋๋ง ์ฃผ๋ฉด ํด๋น ๋ผ๋ฒจ๋ง ์ ์ฉ
## `run_promql` ๊ฐ๋๋ ์ผ ๐
- `approved=False`: ์คํํ์ง ์๊ณ ํ์ธ ๋ฉ์์ง ๋ฐํ
- `approved=True`: ์คํ
๋ชจ๋:
- `instant=True` -> `/api/v1/query`
- `instant=False` -> `/api/v1/query_range`
## ์ฌ์ฉ ์์ ๐
### 1) ํน์ ์๋ฒ CPU ํ๊ท (์ต๊ทผ 24์๊ฐ)
```json
{
"check_id": "cpu_avg_pct",
"hours": 24,
"instance": "10.23.12.11:9100",
"environment": "prod"
}
```
### 2) ํน์ ์๋ฒ ๋์คํฌ ์ฌ์ฉ๋ฅ (mountpoint๋ณ)
```json
{
"check_id": "disk_used_pct_by_mount",
"hours": 24,
"server_name": "CMS AP #1",
"environment": "prod"
}
```
### 3) ์ฌ์ฉ์ PromQL ์คํ (instant)
```json
{
"promql": "up",
"approved": true,
"instant": true,
"environment": "prod"
}
```
### 4) Loki host ํ๋ณด ์กฐํ
```json
{
"loki_environment": "dev_test",
"log_env": "DEV",
"app": "finast",
"hours": 1
}
```
### 5) Loki ๋ก๊ทธ ์กฐํ
```json
{
"loki_environment": "prod",
"log_env": "prod",
"host": "cms-ap-01",
"app": "cms",
"hours": 1,
"contains": "timeout",
"limit": 200
}
```
## CHECKS Catalog โ
> Source: `domain/checks.py` (`CHECKS`)
### System / Resource
- `cpu_avg_pct`: CPU average usage (%) by instance/server_name
- `cpu_peak_pct`: window peak CPU usage (%) over selected range
- `mem_used_pct`: memory used ratio (%)
- `mem_swap_used_pct`: swap used ratio (%)
- `load15_avg`: 15-minute load average
- `cpu_iowait_pct`: CPU iowait ratio (%)
### Disk / Filesystem
- `disk_used_pct_by_mount`: filesystem used (%) by mountpoint/device (0-100 scale)
- `disk_used_top5_pct`: top 5 filesystem usage (%)
- `disk_inodes_used_pct`: inode usage (%)
- `fs_readonly`: readonly filesystem indicator (1=readonly)
- `disk_io_busy_pct`: disk I/O busy ratio (%)
### Availability
- `up`: target liveness (1=up, 0=down)
### Network / TCP
- `net_in_bytes`: inbound throughput (bytes/sec)
- `net_out_bytes`: outbound throughput (bytes/sec)
- `net_errs_per_sec`: RX+TX network errors per second
- `tcp_retrans_per_sec`: TCP retransmit segments per second
- `tcp_established`: established TCP connections
- `tcp_time_wait`: TIME_WAIT TCP sockets
- `tcp_inuse`: in-use TCP sockets
- `tcp_orphan`: orphan TCP sockets
### Process Monitoring
- `proc_cpu_pct`: process group CPU usage (%)
- `proc_mem_bytes`: process group memory usage (bytes)
- `proc_count`: process group process count
### PostgreSQL
- `pg_up`: PostgreSQL exporter up state (1=up, 0=down)
- `pg_qps`: PostgreSQL transactions/sec (commit + rollback)
- `pg_cache_hit_pct`: PostgreSQL buffer cache hit ratio (%)
- `pg_active_conn`: active PostgreSQL connections
## ํ๊ฒฝ ๋ณ์ ์์ฝ โ๏ธ
```env
PROM_ENV_URLS={"prod":"http://...:9090","dev_test":"http://...:9090","dr":"http://...:9090"}
PROM_URL=http://...:9090
PROM_BEARER_TOKEN=
PROM_TIMEOUT_SEC=15
LOKI_ENV_URLS={"prod":"http://...:3100","dev_test":"http://...:3100"}
LOKI_URL=http://...:3100
LOKI_BEARER_TOKEN=
LOKI_TIMEOUT_SEC=15
ALERT_WARN_PCT=85
ALERT_CRIT_PCT=95
ALERT_SUSTAIN_MINUTES=5
PROM_MAX_SAMPLES_PER_SERIES=5000
PROM_MAX_PARALLEL_CHECKS=6
```
ํ๊ฒฝ ์ ํ ์ฐ์ ์์:
1. `environment`
2. `env_hint`
3. `PROM_URL` fallback
Loki ํ๊ฒฝ ์ ํ ์ฐ์ ์์:
1. `loki_environment`
2. `LOKI_URL` fallback
## ์ด์ ํ ๐ก
- ๋ฆฌํฌํธ ์ถ๋ ฅ ์ `%` ๋จ์๋ฅผ ๋ช
ํํ ํ๊ธฐํ์ธ์.
- ๋จ์ผ ์๋ฒ ์ ๊ฒ์ `instance` ๋๋ `server_name` ํํฐ๋ฅผ ์ฌ์ฉํ์ธ์.
- `disk_used_pct_by_mount` ๊ฐ์ 0~100 ์ค์ผ์ผ์
๋๋ค. (`0.8` = `0.8%`)
- Loki ์กฐํ ์ ์๋ `list_loki_hosts` ๋๋ `list_loki_apps`๋ก ํ๋ณด๊ฐ์ ๋จผ์ ํ์ธํ๋ ํธ์ด ์์ ํฉ๋๋ค.
TDQS
Scored across 12 tools
Most tools have distinct purposes, such as get_alerts for fetching alerts and run_check for executing specific checks. However, list_loki_apps, list_loki_environments, and list_loki_hosts are ambiguous without descriptions, potentially causing confusion about their specific functions and differences.
Tool names follow a highly consistent verb_noun pattern throughout, such as list_servers, run_promql, and find_logs. All tools use snake_case without any deviations, making the naming predictable and easy to understand.
With 12 tools, the count is well-suited for a Prometheus monitoring server, covering key operations like listing resources and running queries. It's slightly on the higher side but remains reasonable for the domain, though some tools like the Loki-related ones might be redundant or under-specified.
The toolset provides good coverage for querying and listing operations in Prometheus, including alerts, checks, and servers. However, there are notable gaps, such as missing tools for creating or managing alerts, configuring checks, or handling Loki data beyond listing, which limits full lifecycle management in the monitoring domain.