Skip to main content
Glama
temmeik
by temmeik
README.md
<div align="center">

# ⚡ Vigilo

**Your self-hosted AI on-call engineer.**
Uptime monitoring that doesn't just *tell* you something broke — it investigates *why* and reports back, using your own local LLM.

[![License: MIT](https://img.shields.io/badge/License-MIT-34d399.svg)](LICENSE)
[![Node](https://img.shields.io/badge/node-%E2%89%A518-34d399.svg)](https://nodejs.org)
[![Docker](https://img.shields.io/badge/docker-compose%20up-34d399.svg)](#-quick-start)
[![Zero dependencies](https://img.shields.io/badge/npm%20dependencies-0-34d399.svg)](#-why-vigilo)

**Monitoring · AI incident reports · MCP-native · Public status page · 100% self-hosted**

[Quick start](#-quick-start) · [How the agent works](#-how-the-agent-works) · [MCP](#-plug-it-into-claude-mcp) · [Screenshots](#-screenshots)

</div>

---

![Vigilo dashboard](screenshots/dashboard.png)

## 🔥 Why Vigilo

PagerDuty charges per user. StatusPage costs $99/mo. n8n is a behemoth for a simple "watch my stuff" job.
And when your API dies at 3am, none of them tell you **why**.

Vigilo is one small self-hosted service that:

- 📡 **Monitors** HTTP(s) endpoints and TCP ports on a schedule, with retries, keyword checks and expected status codes
- 🚨 **Opens incidents** and pings you (webhook / Telegram / ntfy) the moment something goes down
- 🤖 **Investigates with an AI agent** — it re-runs checks, probes related endpoints, inspects history, checks the blast radius across your other monitors, then writes a structured report: *severity, probable cause, recommended next steps*
- 🧠 **Runs on your local LLM** (Ollama, LM Studio) or any OpenAI-compatible API — no data ever leaves your machine if you don't want it to
- 🔌 **Is MCP-native** — Vigilo exposes its own MCP server (`/mcp`), so Claude (or any MCP client) can ask *"is anything down?"*; and it can consume remote MCP servers as extra tools for the agent
- 🌍 **Publishes a status page** at `/status` — the $99/mo SaaS, self-hosted and free

## 🚀 Quick start

**No build step. No npm install. Zero dependencies.**

```bash
git clone https://github.com/temmeik/vigilo.git
cd vigilo
node server.js
# ➜ Dashboard:   http://localhost:4242
# ➜ Status page: http://localhost:4242/status
```

or with Docker:

```bash
docker compose up -d          # set ADMIN_PASSWORD in .env first (see .env.example)
```

First launch asks you to create an admin password. Then hit **Load demo monitors** to see it alive in 10 seconds.

### Turn on the AI (local, 2 minutes)

```bash
# 1. Install Ollama → https://ollama.com
ollama pull llama3.2
# 2. In Vigilo: Settings → AI Agent → preset "Ollama (local)" → Test connection → Save
```

Any OpenAI-compatible API works too: OpenAI, Groq, OpenRouter, LM Studio, vLLM — just paste the base URL and key.

## 🤖 How the agent works

When a monitor goes down, the agent gets a **tool loop** (max N iterations, you decide):

| Tool | What it does |
|---|---|
| `http_request` | Probe the failing URL, health endpoints, dependencies |
| `get_monitor_history` | See when degradation started and how it trended |
| `list_monitors` | Check the blast radius — is the whole host down or one service? |
| `trigger_check` | Force a fresh check to confirm |
| `submit_report` | Deliver the structured verdict |
| *your MCP tools* | Any tool from MCP servers you connect in Settings |

The report lands in the incident timeline and in your notifications:

![Agent incident report](screenshots/incident-agent.png)

## 🔌 Plug it into Claude (MCP)

Vigilo *is* an MCP server. Add it to your MCP client config:

```json
{
  "mcpServers": {
    "vigilo": {
      "url": "http://localhost:4242/mcp",
      "headers": { "Authorization": "Bearer <your-api-key>" }
    }
  }
}
```

Then just ask Claude: *"anything down right now?"* — it will call `get_status_summary` and `get_incidents`. Tools exposed: `list_monitors`, `get_monitor`, `get_incidents`, `trigger_check`, `get_status_summary`.

*(API key is in Settings → API & MCP access.)*

## 🌍 Status page

![Status page](screenshots/status-page.png)

Toggle any monitor to **public** and it appears on `/status` with a 90-day uptime bar, uptime percentages and incident history. No login required. Link it in your app's footer.

## 📸 Screenshots

| Dashboard | Agent settings |
|---|---|
| ![Dashboard](screenshots/dashboard.png) | ![Settings](screenshots/settings-agent.png) |

## 🧱 Architecture

- **Zero npm dependencies** — the entire backend is plain Node.js (≥18): scheduler, check engine, agent loop, MCP client *and* server, SSE hub, auth (scrypt + session cookies), JSON storage with atomic flushes
- **One process** — `node server.js` is everything; data lives in `./data/vigilo.json`
- **Live UI** — hand-rolled vanilla-JS SPA, Server-Sent Events push every state change instantly, offline-friendly (no CDN, no fonts, no tracking)
- **State machine** — heartbeat → consecutive failures → incident → agent escalation → auto-resolve on recovery

## 🗺️ Roadmap

- [ ] SQLite storage backend (optional, for big fleets)
- [ ] DNS / ICMP / Docker container monitors
- [ ] Agent "suggested fix" actions (restart container via MCP tool)
- [ ] Multi-user + read-only team seats
- [ ] Status page custom domains & theming

## 🛠️ Development

```bash
node --watch server.js   # auto-restart on changes
npm run mock-llm         # fake OpenAI-compatible LLM on :4243 for testing the agent
npm test                 # 28-assertion e2e suite (needs server + mock-llm running)
```

## ⚖️ License

MIT — do whatever you want, a star is appreciated ⭐

---

<div align="center">
<sub>Built for indie hackers and small teams who can't justify PagerDuty — but still get paged at 3am.</sub>
</div>