Vigilo
by temmeik
README.md
<div align="center">
# ⚡ Vigilo
**Your self-hosted AI on-call engineer.**
Uptime monitoring that doesn't just *tell* you something broke — it investigates *why* and reports back, using your own local LLM.
[](LICENSE)
[](https://nodejs.org)
[](#-quick-start)
[](#-why-vigilo)
**Monitoring · AI incident reports · MCP-native · Public status page · 100% self-hosted**
[Quick start](#-quick-start) · [How the agent works](#-how-the-agent-works) · [MCP](#-plug-it-into-claude-mcp) · [Screenshots](#-screenshots)
</div>
---

## 🔥 Why Vigilo
PagerDuty charges per user. StatusPage costs $99/mo. n8n is a behemoth for a simple "watch my stuff" job.
And when your API dies at 3am, none of them tell you **why**.
Vigilo is one small self-hosted service that:
- 📡 **Monitors** HTTP(s) endpoints and TCP ports on a schedule, with retries, keyword checks and expected status codes
- 🚨 **Opens incidents** and pings you (webhook / Telegram / ntfy) the moment something goes down
- 🤖 **Investigates with an AI agent** — it re-runs checks, probes related endpoints, inspects history, checks the blast radius across your other monitors, then writes a structured report: *severity, probable cause, recommended next steps*
- 🧠 **Runs on your local LLM** (Ollama, LM Studio) or any OpenAI-compatible API — no data ever leaves your machine if you don't want it to
- 🔌 **Is MCP-native** — Vigilo exposes its own MCP server (`/mcp`), so Claude (or any MCP client) can ask *"is anything down?"*; and it can consume remote MCP servers as extra tools for the agent
- 🌍 **Publishes a status page** at `/status` — the $99/mo SaaS, self-hosted and free
## 🚀 Quick start
**No build step. No npm install. Zero dependencies.**
```bash
git clone https://github.com/temmeik/vigilo.git
cd vigilo
node server.js
# ➜ Dashboard: http://localhost:4242
# ➜ Status page: http://localhost:4242/status
```
or with Docker:
```bash
docker compose up -d # set ADMIN_PASSWORD in .env first (see .env.example)
```
First launch asks you to create an admin password. Then hit **Load demo monitors** to see it alive in 10 seconds.
### Turn on the AI (local, 2 minutes)
```bash
# 1. Install Ollama → https://ollama.com
ollama pull llama3.2
# 2. In Vigilo: Settings → AI Agent → preset "Ollama (local)" → Test connection → Save
```
Any OpenAI-compatible API works too: OpenAI, Groq, OpenRouter, LM Studio, vLLM — just paste the base URL and key.
## 🤖 How the agent works
When a monitor goes down, the agent gets a **tool loop** (max N iterations, you decide):
| Tool | What it does |
|---|---|
| `http_request` | Probe the failing URL, health endpoints, dependencies |
| `get_monitor_history` | See when degradation started and how it trended |
| `list_monitors` | Check the blast radius — is the whole host down or one service? |
| `trigger_check` | Force a fresh check to confirm |
| `submit_report` | Deliver the structured verdict |
| *your MCP tools* | Any tool from MCP servers you connect in Settings |
The report lands in the incident timeline and in your notifications:

## 🔌 Plug it into Claude (MCP)
Vigilo *is* an MCP server. Add it to your MCP client config:
```json
{
"mcpServers": {
"vigilo": {
"url": "http://localhost:4242/mcp",
"headers": { "Authorization": "Bearer <your-api-key>" }
}
}
}
```
Then just ask Claude: *"anything down right now?"* — it will call `get_status_summary` and `get_incidents`. Tools exposed: `list_monitors`, `get_monitor`, `get_incidents`, `trigger_check`, `get_status_summary`.
*(API key is in Settings → API & MCP access.)*
## 🌍 Status page

Toggle any monitor to **public** and it appears on `/status` with a 90-day uptime bar, uptime percentages and incident history. No login required. Link it in your app's footer.
## 📸 Screenshots
| Dashboard | Agent settings |
|---|---|
|  |  |
## 🧱 Architecture
- **Zero npm dependencies** — the entire backend is plain Node.js (≥18): scheduler, check engine, agent loop, MCP client *and* server, SSE hub, auth (scrypt + session cookies), JSON storage with atomic flushes
- **One process** — `node server.js` is everything; data lives in `./data/vigilo.json`
- **Live UI** — hand-rolled vanilla-JS SPA, Server-Sent Events push every state change instantly, offline-friendly (no CDN, no fonts, no tracking)
- **State machine** — heartbeat → consecutive failures → incident → agent escalation → auto-resolve on recovery
## 🗺️ Roadmap
- [ ] SQLite storage backend (optional, for big fleets)
- [ ] DNS / ICMP / Docker container monitors
- [ ] Agent "suggested fix" actions (restart container via MCP tool)
- [ ] Multi-user + read-only team seats
- [ ] Status page custom domains & theming
## 🛠️ Development
```bash
node --watch server.js # auto-restart on changes
npm run mock-llm # fake OpenAI-compatible LLM on :4243 for testing the agent
npm test # 28-assertion e2e suite (needs server + mock-llm running)
```
## ⚖️ License
MIT — do whatever you want, a star is appreciated ⭐
---
<div align="center">
<sub>Built for indie hackers and small teams who can't justify PagerDuty — but still get paged at 3am.</sub>
</div>
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues