runner-fleet-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@runner-fleet-mcpshow me the runner fleet status and SLOs"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Runner Fleet Reliability Platform
A portfolio platform for operating ephemeral GitHub Actions runner fleets with SLOs, capacity policy, cost evidence, Prometheus metrics, incident simulation, OpenTofu, Kubernetes, and a read-only MCP interface.
Status
Version 0.1.0 is a local engineering scaffold. It does not yet represent a production deployment, live AWS operation, or paying-customer system.
No cloud resources are created by the repository unless an operator explicitly runs an OpenTofu apply.
Current capabilities
Deterministic simulations for baseline, burst, capacity loss, image pull failure, and GitHub API degradation
Queue, startup, cleanup, and cost SLO evaluation
Advisory capacity recommendations with explicit maximums
SQLite report history and incident summaries
Prometheus fleet metrics
FastAPI control and evidence API
Read-only MCP tools for fleet status, SLOs, incidents, capacity, and cost
Tests and synthetic examples
Safety boundary
MCP tools are intentionally read-only. They can explain fleet evidence and recommend capacity, but they cannot create clusters, scale runners, apply OpenTofu, or mutate Kubernetes resources.
Local development
python -m venv .venv
.\.venv\Scripts\python -m pip install -e ".[dev]"
.\.venv\Scripts\python -m pytest -q
.\.venv\Scripts\runner-fleet-simulate --scenario baseline --jobs 20
.\.venv\Scripts\runner-fleet-apiOpen:
API documentation:
http://127.0.0.1:8080/docsPrometheus metrics:
http://127.0.0.1:8080/metrics
Run the MCP server over stdio:
.\.venv\Scripts\runner-fleet-mcpRun the complete deterministic demonstration:
.\scripts\run-demo.ps1The demo compares a healthy baseline with a capacity-loss incident and writes
portfolio evidence under evidence/generated.
Current synthetic evidence:
Baseline: queue p95 10 seconds, no startup failures, all SLOs pass
Burst: queue and maximum-age SLOs fail
Capacity loss: queue, startup-success, and maximum-age SLOs fail
Image pull failure: queue and startup-success SLOs fail
GitHub API degradation: queue and startup-success SLOs fail
Initial SLOs
Signal | Target |
Queue-to-start p95 | 90 seconds or less |
Runner startup success | 99% or greater |
Runner cleanup success | 100% |
Maximum queue age | 300 seconds or less |
These are portfolio targets. They become claims only after a real deployment produces supporting measurements.
Roadmap
Validate the two kind-cluster development environment.
Deploy Actions Runner Controller and Prometheus.
Apply the temporary AWS EKS lab after explicit cost approval.
Add controlled chaos experiments against real Kubernetes resources.
Capture a cloud deployment, load test, incident timeline, cost report, and complete teardown.
Record a two-minute technical demonstration.
Infrastructure
infra/bootstrap: encrypted S3 state bucket and DynamoDB lock tableinfra/aws: VPC, public lab subnets, EKS 1.36, Spot managed nodes, ECR, CloudWatch control-plane logs, budget, and optional GitHub OIDC roledeploy/kind: primary and recovery local clustersdeploy/arc: controller and runner scale-set valuesdeploy/helm/runner-fleet-platform: API, metrics, disruption budget, ServiceMonitor, and SLO alerts
OpenTofu applies are intentionally absent from CI. CI formats and validates the configuration only.
Documentation
Independence
This is a personal, unofficial clean-room project based only on public documentation and synthetic data. It contains no employer source code, customer data, internal configurations, support cases, or credentials.
This server cannot be deployed
Maintenance
Related MCP Connectors
Read-only MCP access to a documented IT fleet: state, changes, posture. 15 tools.
Provides read access to your GKE and Kubernetes resources.
Read-only MCP access to sessions, funnels, campaigns, errors, live visitors, and anomalies.
Read-only access to a Lumin project's logs, metrics, uptime checks, alerts and infrastructure.