runner-fleet-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@runner-fleet-mcpshow me the runner fleet status and SLOs"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Runner Fleet Reliability Platform
A portfolio platform for operating ephemeral GitHub Actions runner fleets with SLOs, capacity policy, cost evidence, Prometheus metrics, incident simulation, OpenTofu, Kubernetes, and a read-only MCP interface.
Status
Version 0.1.0 is a local engineering scaffold. It does not yet represent a production deployment, live AWS operation, or paying-customer system.
No cloud resources are created by the repository unless an operator explicitly runs an OpenTofu apply.
Related MCP server: Claude Ops Investigator
Current capabilities
Deterministic simulations for baseline, burst, capacity loss, image pull failure, and GitHub API degradation
Queue, startup, cleanup, and cost SLO evaluation
Advisory capacity recommendations with explicit maximums
SQLite report history and incident summaries
Prometheus fleet metrics
FastAPI control and evidence API
Read-only MCP tools for fleet status, SLOs, incidents, capacity, and cost
Tests and synthetic examples
Safety boundary
MCP tools are intentionally read-only. They can explain fleet evidence and recommend capacity, but they cannot create clusters, scale runners, apply OpenTofu, or mutate Kubernetes resources.
Local development
python -m venv .venv
.\.venv\Scripts\python -m pip install -e ".[dev]"
.\.venv\Scripts\python -m pytest -q
.\.venv\Scripts\runner-fleet-simulate --scenario baseline --jobs 20
.\.venv\Scripts\runner-fleet-apiOpen:
API documentation:
http://127.0.0.1:8080/docsPrometheus metrics:
http://127.0.0.1:8080/metrics
Run the MCP server over stdio:
.\.venv\Scripts\runner-fleet-mcpRun the complete deterministic demonstration:
.\scripts\run-demo.ps1The demo compares a healthy baseline with a capacity-loss incident and writes
portfolio evidence under evidence/generated.
Current synthetic evidence:
Baseline: queue p95 10 seconds, no startup failures, all SLOs pass
Burst: queue and maximum-age SLOs fail
Capacity loss: queue, startup-success, and maximum-age SLOs fail
Image pull failure: queue and startup-success SLOs fail
GitHub API degradation: queue and startup-success SLOs fail
Initial SLOs
Signal | Target |
Queue-to-start p95 | 90 seconds or less |
Runner startup success | 99% or greater |
Runner cleanup success | 100% |
Maximum queue age | 300 seconds or less |
These are portfolio targets. They become claims only after a real deployment produces supporting measurements.
Roadmap
Validate the two kind-cluster development environment.
Deploy Actions Runner Controller and Prometheus.
Apply the temporary AWS EKS lab after explicit cost approval.
Add controlled chaos experiments against real Kubernetes resources.
Capture a cloud deployment, load test, incident timeline, cost report, and complete teardown.
Record a two-minute technical demonstration.
Infrastructure
infra/bootstrap: encrypted S3 state bucket and DynamoDB lock tableinfra/aws: VPC, public lab subnets, EKS 1.36, Spot managed nodes, ECR, CloudWatch control-plane logs, budget, and optional GitHub OIDC roledeploy/kind: primary and recovery local clustersdeploy/arc: controller and runner scale-set valuesdeploy/helm/runner-fleet-platform: API, metrics, disruption budget, ServiceMonitor, and SLO alerts
OpenTofu applies are intentionally absent from CI. CI formats and validates the configuration only.
Documentation
Independence
This is a personal, unofficial clean-room project based only on public documentation and synthetic data. It contains no employer source code, customer data, internal configurations, support cases, or credentials.
Available Tools
5 toolsget_cost_estimateB
Return the latest synthetic runner-compute estimate and exclusions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full behavioral burden and largely fails it: it does not say what 'synthetic' means, how fresh the estimate is, what the 'exclusions' are, whether auth is needed, or how the result is shaped. The only behavioral hint is the word 'latest', implying a cached/precomputed value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the verb leads and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read tool this is minimal but workable; still, with no annotations and no output schema, the description is the only contract and it does not explain the returned estimate's structure, units, or what is excluded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a clear verb ('Return') and a specific resource ('latest synthetic runner-compute estimate and exclusions'), which separates it from get_fleet_status, get_slo_report, and list_recent_incidents. However, 'synthetic' is unexplained jargon and it never acknowledges recommend_runner_capacity, the closest sibling, so sibling differentiation is only partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the related alternative recommend_runner_capacity. An agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fleet_statusB
Return the latest read-only runner fleet snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden of behavioral disclosure. It does declare a meaningful trait – 'read-only' – and implies a point-in-time 'snapshot' that is 'latest', adding some value. But it omits auth/permission requirements, rate limits, freshness guarantees, and what the snapshot contains, which are important for a zero-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler; the verb and resource are front-loaded. Every word earns its place for a zero-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (no params, no nested objects) and the description is minimal but not wrong. However, with no output schema and no annotations, it should describe what the snapshot returns or when it is stale, and it does neither. Adequate but with a clear gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There are no parameter semantics to clarify, and the description does not need to compensate for any schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('runner fleet snapshot'), making the tool's function immediately clear. It also adds a scope qualifier ('latest', 'read-only'). However, it does nothing to distinguish this from siblings like get_slo_report or list_recent_incidents beyond the resource name itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the alternatives in the sibling set (e.g., recommend_runner_capacity, get_slo_report). No prerequisites, no triggering context, and no exclusions are stated. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_slo_reportB
Return current SLO checks and breaches without changing infrastructure.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does disclose one meaningful behavioral trait: it is non-mutating ('without changing infrastructure'). However, it says nothing about auth requirements, rate limits, freshness of the data, or what the report contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the resource and the key safety constraint with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, read-only reporting tool with no output schema and no annotations, the description covers the essentials but leaves the shape of the 'report' and any scoping (all SLOs? which services?) unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so parameter semantics is inherently uncomplicated; the baseline for a 0-param tool is 4. The description adds nothing param-related, but there is nothing to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and resource ('current SLO checks and breaches'), which clearly separates it from siblings like get_fleet_status, list_recent_incidents, and get_cost_estimate. It doesn't explicitly name a sibling it is not, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives no explicit when-to-use guidance, no prerequisites, and does not reference any alternative tool. The intended use (checking SLO status) must be inferred from the resource name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recent_incidentsC
List recent synthetic SLO-breach reports and failure evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'List' weakly implies a read, but nothing states side effects, ordering, recency window, or pagination, and the term 'synthetic' is ambiguous about whether the data is real or mock.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the action and object front-loaded and no filler. It is well-sized for the tool, though brevity is partly achieved by omitting necessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But with no annotations, an undocumented parameter, an undefined recency window, and no differentiation from get_slo_report, the description is only minimally adequate for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'limit' parameter (default 10) has no documentation anywhere. The description's 'recent' hints at a recency bound but never defines it or explains how limit interacts with it, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource ('List recent ... reports and failure evidence') makes the operation clear. However, it does not distinguish this tool from the close sibling get_slo_report, and the qualifier 'synthetic' is left unexplained, so the agent cannot tell which of the two SLO-oriented tools to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, and no mention of alternatives such as get_slo_report or get_fleet_status. The agent gets no signal about which sibling surfaces SLO breaches versus which surfaces fleet or cost data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_runner_capacityB
Calculate advisory runner capacity; this tool never scales resources.
| Name | Required | Description | Default |
|---|---|---|---|
| queued_jobs | Yes | ||
| current_runners | Yes | ||
| maximum_runners | No | ||
| average_job_seconds | No | ||
| target_queue_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| capped | Yes | |
| rationale | Yes | |
| desired_runners | Yes | |
| additional_runners | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses a key behavioral trait: the tool is advisory and never scales resources, making its non-mutating nature clear. It does not cover other behavioral details such as permissions or determinism, but the critical safety property is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. It delivers purpose and a critical constraint immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, so the description need not explain them. However, with no annotations, 0% schema description coverage, and five parameters, the description is thin on parameter semantics and usage context. It is minimally viable but leaves clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for five parameters, so the description must compensate. It adds no meaning whatsoever for queued_jobs, current_runners, maximum_runners, average_job_seconds, or target_queue_seconds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Calculate') and resource ('runner capacity'), and explicitly distinguishes the tool from any scaling action. It does not, however, differentiate itself from siblings like get_fleet_status or get_cost_estimate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when not to use it ('never scales resources'), which is useful, but it does not state when to use it or name alternative tools for related tasks. Usage context must be inferred from the purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
get_cost_estimate - First observed
get_fleet_status - First observed
get_slo_report - First observed
list_recent_incidents - First observed
recommend_runner_capacity
TDQS
Scored across 5 tools
Each tool targets a distinct concern (fleet status, SLO checks, capacity advice, incidents, cost). The main overlap is between get_slo_report (current breaches) and list_recent_incidents (recent breach reports), but descriptions clarify the temporal and evidentiary distinction.
All names use snake_case with a clear verb_noun pattern (get_fleet_status, list_recent_incidents, recommend_runner_capacity). The verbs get/list/recommend are appropriate and consistently ordered.
Five tools is well-scoped for a read-only advisory server covering fleet status, SLOs, capacity, incidents, and cost. No tool feels redundant or missing at this count.
The read-only advisory surface covers the main observability areas (status, SLOs, capacity, incidents, cost). Minor gaps exist for drill-down details (e.g., per-runner information or historical trends), but core agent workflows are supported.
Maintenance
Related MCP Connectors
Read-only MCP access to a documented IT fleet: state, changes, posture. 15 tools.
Provides read access to your GKE and Kubernetes resources.
Read-only MCP access to sessions, funnels, campaigns, errors, live visitors, and anomalies.
Read-only access to a Lumin project's logs, metrics, uptime checks, alerts and infrastructure.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables safe, read-only interaction with Kubernetes clusters, allowing users to list resources and fetch logs without any create/update/delete operations.116Apache 2.0
- FlicenseNot gradedqualityBmaintenanceEnables read-only Kubernetes incident investigation through MCP tools for listing pods, describing resources, fetching logs, and searching runbooks.1-
- AlicenseNot gradedqualityCmaintenanceEnables read-only inspection of Google Ads accounts through MCP, including account inventory, reporting, metadata, change history, and safe GAQL queries.Apache 2.0
- FlicenseAqualityBmaintenanceEnables read-only inspection of pipeline jobs, pods, configurations, and logs from the dev EKS cluster and log store.8-