Triage MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Triage MCP Servercheck the health status of the service"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ฉบ Triage โ a self-healing ops MCP for any Dockerized service
Let an AI agent (or a human) check, diagnose, and recover a service โ without a host shell.
Most "give the agent ops powers" setups are bad: you either hand the model a raw shell (now it can roam the whole box and conflate unrelated subsystems), or you wire up dashboards a model can't read. Triage is the third option:
A small MCP server that exposes a handful of health/diagnose/recover tools. Each returns raw evidence AND a plain-English translation, a suggested action, and whether the fix is safe to auto-apply. The agent acts through tools โ it never touches the host directly.
The policy that makes it safe
Class | Tools | Behaviour |
Auto-fix safe |
| An agent may run these on its own and report after. Infra only โ no data touched. |
Ask before risky |
| Anything that could lose data / change external state. Dry-run unless |
Can't self-fix | (reported) | Diagnosed and handed to the human with exact steps โ never faked. |
The dual raw + layman output is the differentiator: the agent gets structured data to act on, and the
human gets a sentence they can actually understand ("Postiz's API engine isn't running โ the known cold-boot hiccup. I'll restart it.").
Related MCP server: docker-mcp-server
Tools
Tool | Kind | What it does |
| read | Containers + configured in-container processes + optional dependency ping. |
| read | Health check matched to a runbook โ issues with raw + plain-English + action + |
| read | Raw service log tail. |
| safe | Restart one in-container process (pm2). |
| safe | Recreate the service container from compose. No volumes/data touched. |
| risky | Dry-run by default; runs the configured risky command only on |
Configure (zero code changes)
Everything is env-driven โ point it at any compose-managed service:
TRIAGE_COMPOSE=/path/to/docker-compose.yaml # compose file
TRIAGE_SERVICE=app # the main container/service name
TRIAGE_LABEL="My App" # friendly name used in messages
TRIAGE_PROCS=backend,worker # optional: in-container processes to watch
TRIAGE_PROC_MGR=pm2 # "pm2" | "none"
TRIAGE_DB_PING="docker exec app-db pg_isready" # optional: rc 0 = dependency healthy
TRIAGE_RISKY_CMD="" # optional: a guarded recovery (clear a queue, etc.)
TRIAGE_RISKY_DESC="clear the stuck job queue"
TRIAGE_PORT=9500See .env.example.
Run
pip install -r requirements.txt
python3 triage.py # serves an MCP over streamable-http on TRIAGE_PORTRegister it with your agent runtime (any MCP client). For an always-on host service, use the included
launchd template com.triage.ops.plist (macOS) โ adapt to systemd on Linux.
Hard-won lessons baked in
Agents in a container can't see host processes. Give them status tools, not a shell. With shell access a model conflates unrelated subsystems and reports false negatives. Tools keep it honest.
Two reports, always. Structured
rawfor the agent to branch on; a one-sentencelaymanfor the human. A health check the human can't read is half a tool.Encode the safe/risky boundary in the tool, not the prompt. "Don't clear the queue without asking" in a system prompt is a suggestion; a
confirm=true-gated dry-run is a guarantee.docker compose ps --format jsonvaries by version (NDJSON vs single array) โ handle both.Recover โ restart. A dead process needs a restart; an unhealthy container needs a recreate. Separate tools so the agent escalates correctly.
Built by
Built by KodeKing ยท author Fazal Shah. We build local, private, multi-agent AI systems for teams who can't send their data to the cloud. Issues and PRs welcome.
License
MIT โ see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
- emisarOAuthdev.emisar
Let AI operate servers without SSH. Choose actions, approve risky changes, and audit every step.
Hosting for AI agents: your AI client deploys Docker apps to live HTTPS URLs over MCP.
Watchdog for unattended AI agents: alerts, evidence checks and a verifiable proof per run.
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceEnables AI assistants to interact with Docker containers through safe, permission-controlled access to inspect, manage, and diagnose containers, images, and compose services with built-in timeouts and AI-powered analysis.-
- AlicenseBqualityCmaintenanceEnables AI agents to manage Docker containers, images, Compose stacks, health checks, and logs through a unified MCP interface, ensuring containers stay running with self-healing capabilities.3166 npm5MIT
- AlicenseAqualityCmaintenanceEnables AI agents to monitor Docker containers, system resources, ports, and logs in real-time.5MIT
- AlicenseAqualityBmaintenanceEnables AI agents to safely monitor Linux VPS health, inspect Docker containers, retrieve service logs, and execute whitelisted recovery actions via MCP.31248 npmMIT