ai-sre
Provides tools for inspecting Docker container states and health as part of system reconnaissance across cluster nodes.
Provides tools for inspecting PM2 process clusters, service state, and process health across cluster nodes.
Provides fast JSON telemetry streaming for Three.js WebGL 3D Arc Reactor HUD interfaces.
Allows dispatching incident alerts and notifications via WhatsApp Bot worker integration.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ai-srecheck system health across all nodes"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
š”ļø AI-SRE: Autonomous Site Reliability Engineering & Observability Daemon
Autonomous Infrastructure Reliability, Security Reconnaissance, and Self-Healing Engine for Distributed Cloud & Edge Nodes.
š Overview
AI-SRE is a lightweight, agentic Site Reliability Engineering daemon engineered to safeguard multi-node infrastructure. Unlike traditional heavy APM suites, AI-SRE pairs autonomous background telemetry collection with asymmetric cryptographic authentication (Ed25519) and native Model Context Protocol (MCP) interfaces, enabling AI assistants (such as FRIDAY / Antigravity) to inspect, triage, and remediate system anomalies autonomously.
Related MCP server: Nibble
ā” Core Capabilities
Zero-Trust Cryptographic Communication
All inter-node telemetry requests (
/api/*) are signed using Ed25519 asymmetric private keys and verified against authorized public keys (keys/*.pub).Replay protection with strict 60-second timestamp freshness windows.
Full-Spectrum System Telemetry & Reconnaissance
Metrics Engine (
src/metrics.js): Real-time CPU pressure, memory utilization (RSS / Heap / Swap), disk I/O, and network bandwidth.Reconnaissance Engine (
src/recon.js): PM2 process cluster inspection, systemd service health, Docker container states, and listening socket audits.
Autonomous SRE Patrol & Self-Healing Hub
Distributed multi-node patrol orchestrator (
patrol-hub.js).Automated memory leak detection and graceful service restarts.
Circuit-breaker cooldown preventing alert spamming.
Native MCP Server (Model Context Protocol)
Exposes production-ready tools for AI agents:
sre_get_system_health: Real-time health check across cluster nodes.sre_get_metrics: CPU, memory, and disk telemetry.sre_get_recon: Service state, PM2 processes, and open ports.sre_run_patrol: Execute comprehensive multi-node patrol.sre_send_alert: Dispatch incident notification.
Telemetry Bridge for Holographic HUD & Notifications
Fast JSON telemetry streaming for Three.js WebGL 3D Arc Reactor interfaces.
Direct incident alerting via WhatsApp Bot worker integration.
šļø Architecture
āāāāāāāāāāāāāāāāāāāāāāāāāā
ā Friday AI / Agent ā
ā (Antigravity / MCP) ā
āāāāāāāāāāāāā¬āāāāāāāāāāāāā
ā MCP Tools
ā¼
āāāāāāāāāāāāāāāāāāāāāāāāāā
ā AI-SRE Hub ā
ā (vm-maskii :3400) ā
āāāāāāāāāāāāā¬āāāāāāāāāāāāā
ā
āāāāāāāāāāāāāāāāāāāāāāāāāāā“āāāāāāāāāāāāāāāāāāāāāāāāāā
ā Ed25519 Signed Request (Tailscale Private Mesh) ā
ā¼ ā¼
āāāāāāāāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāāāā
ā AI-SRE Node (lucky) ā ā AI-SRE Node (atcs) ā
ā - 14 PM2 Processes ā ā - KVM Hypervisor ā
ā - MySQL & Nginx ā ā - Ant Media ā
āāāāāāāāāāāāāāāāāāāāāāāāāā āāāāāāāāāāāāāāāāāāāāāāāāāāš Getting Started
Prerequisites
Node.js 18.0.0 or higher
Linux (Ubuntu / Debian / RHEL)
Installation
# Clone the repository
git clone https://github.com/lhermawan/ai-sre.git
cd ai-sre
# Install dependencies
npm install
# Copy environment configuration
cp .env.example .envKeypair Generation (Ed25519)
Generate asymmetric authentication keypairs for the agent and hub:
mkdir -p keys
ssh-keygen -t ed25519 -N "" -f keys/friday_agent.key -C "friday-agent-auth"Export authorized public keys to the node's keys/ directory (keys/*.pub).
š§ CLI & MCP Usage
Running Locally
# Start daemon
npm start
# Run system metrics check
node cli.js metrics
# Run service reconnaissance
node cli.js recon
# Execute full cluster patrol
node patrol-hub.jsConnecting to MCP (Claude Desktop / Antigravity)
Add to your MCP configuration (mcp_config.json):
{
"mcpServers": {
"ai-sre": {
"command": "node",
"args": ["/path/to/ai-sre/mcp-server.js"]
}
}
}š Security Best Practices
Never commit
.envor private keys (keys/*.key).Keep all inter-node communication confined to private overlay networks (e.g., Tailscale / WireGuard).
Enforce periodic key rotation for Ed25519 agents.
š License
This project is licensed under the ISC License.
Authored by Maskii Studio / Friday AI-SRE Team.
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP-native AI SRE: ask what's broken in production, get a reviewed GitHub fix PR.
- mttrlyOAuthcom.mttrly
AI-powered incident management and server monitoring via MCP.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Protocol-native energy infrastructure orchestration for AI data centers. Provides 46 MCP tools across 8 grid protocols (IEC-61850, DNP3, Modbus, OCPP, OpenADR, IEEE 2030.5, IEC 60870-5-104, ICCP) with 5 core API primitives: connect, dispatch, settle, comply, and intel. Enables AI agents to programmatically interact with substations, grid interfaces, and energy assets for real-time workload-grid coordination.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to manage system updates, application installations, and remote host orchestration through MCP tools.3MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to investigate production incidents by exposing service health, logs, and deployment data through MCP tools.15 npm-
- AlicenseAqualityAmaintenanceEnables AI agents to monitor server health and capacity, diagnose outages, inspect Docker and deployment status, and perform safe, bounded recovery actions via MCP without granting unrestricted shell or SSH access.101Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to inspect Kubernetes pods and cluster events, query Prometheus and Loki, and diagnose pod health with suggested actions through MCP tools.MIT