Skip to main content
Glama
jjoseph456

runner-fleet-mcp

by jjoseph456

Runner Fleet Reliability Platform

A portfolio platform for operating ephemeral GitHub Actions runner fleets with SLOs, capacity policy, cost evidence, Prometheus metrics, incident simulation, OpenTofu, Kubernetes, and a read-only MCP interface.

Status

Version 0.1.0 is a local engineering scaffold. It does not yet represent a production deployment, live AWS operation, or paying-customer system.

No cloud resources are created by the repository unless an operator explicitly runs an OpenTofu apply.

Related MCP server: Claude Ops Investigator

Current capabilities

  • Deterministic simulations for baseline, burst, capacity loss, image pull failure, and GitHub API degradation

  • Queue, startup, cleanup, and cost SLO evaluation

  • Advisory capacity recommendations with explicit maximums

  • SQLite report history and incident summaries

  • Prometheus fleet metrics

  • FastAPI control and evidence API

  • Read-only MCP tools for fleet status, SLOs, incidents, capacity, and cost

  • Tests and synthetic examples

Safety boundary

MCP tools are intentionally read-only. They can explain fleet evidence and recommend capacity, but they cannot create clusters, scale runners, apply OpenTofu, or mutate Kubernetes resources.

Local development

python -m venv .venv
.\.venv\Scripts\python -m pip install -e ".[dev]"
.\.venv\Scripts\python -m pytest -q
.\.venv\Scripts\runner-fleet-simulate --scenario baseline --jobs 20
.\.venv\Scripts\runner-fleet-api

Open:

  • API documentation: http://127.0.0.1:8080/docs

  • Prometheus metrics: http://127.0.0.1:8080/metrics

Run the MCP server over stdio:

.\.venv\Scripts\runner-fleet-mcp

Run the complete deterministic demonstration:

.\scripts\run-demo.ps1

The demo compares a healthy baseline with a capacity-loss incident and writes portfolio evidence under evidence/generated.

Current synthetic evidence:

  • Baseline: queue p95 10 seconds, no startup failures, all SLOs pass

  • Burst: queue and maximum-age SLOs fail

  • Capacity loss: queue, startup-success, and maximum-age SLOs fail

  • Image pull failure: queue and startup-success SLOs fail

  • GitHub API degradation: queue and startup-success SLOs fail

See the generated summary.

Initial SLOs

Signal

Target

Queue-to-start p95

90 seconds or less

Runner startup success

99% or greater

Runner cleanup success

100%

Maximum queue age

300 seconds or less

These are portfolio targets. They become claims only after a real deployment produces supporting measurements.

Roadmap

  1. Validate the two kind-cluster development environment.

  2. Deploy Actions Runner Controller and Prometheus.

  3. Apply the temporary AWS EKS lab after explicit cost approval.

  4. Add controlled chaos experiments against real Kubernetes resources.

  5. Capture a cloud deployment, load test, incident timeline, cost report, and complete teardown.

  6. Record a two-minute technical demonstration.

Infrastructure

  • infra/bootstrap: encrypted S3 state bucket and DynamoDB lock table

  • infra/aws: VPC, public lab subnets, EKS 1.36, Spot managed nodes, ECR, CloudWatch control-plane logs, budget, and optional GitHub OIDC role

  • deploy/kind: primary and recovery local clusters

  • deploy/arc: controller and runner scale-set values

  • deploy/helm/runner-fleet-platform: API, metrics, disruption budget, ServiceMonitor, and SLO alerts

OpenTofu applies are intentionally absent from CI. CI formats and validates the configuration only.

Documentation

Independence

This is a personal, unofficial clean-room project based only on public documentation and synthetic data. It contains no employer source code, customer data, internal configurations, support cases, or credentials.

Available Tools

5 tools
get_cost_estimateB

Return the latest synthetic runner-compute estimate and exclusions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full behavioral burden and largely fails it: it does not say what 'synthetic' means, how fresh the estimate is, what the 'exclusions' are, whether auth is needed, or how the result is shaped. The only behavioral hint is the word 'latest', implying a cached/precomputed value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb leads and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read tool this is minimal but workable; still, with no annotations and no output schema, the description is the only contract and it does not explain the returned estimate's structure, units, or what is excluded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a clear verb ('Return') and a specific resource ('latest synthetic runner-compute estimate and exclusions'), which separates it from get_fleet_status, get_slo_report, and list_recent_incidents. However, 'synthetic' is unexplained jargon and it never acknowledges recommend_runner_capacity, the closest sibling, so sibling differentiation is only partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the related alternative recommend_runner_capacity. An agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_fleet_statusB

Return the latest read-only runner fleet snapshot.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears the full burden of behavioral disclosure. It does declare a meaningful trait – 'read-only' – and implies a point-in-time 'snapshot' that is 'latest', adding some value. But it omits auth/permission requirements, rate limits, freshness guarantees, and what the snapshot contains, which are important for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler; the verb and resource are front-loaded. Every word earns its place for a zero-parameter read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (no params, no nested objects) and the description is minimal but not wrong. However, with no output schema and no annotations, it should describe what the snapshot returns or when it is stale, and it does neither. Adequate but with a clear gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There are no parameter semantics to clarify, and the description does not need to compensate for any schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('runner fleet snapshot'), making the tool's function immediately clear. It also adds a scope qualifier ('latest', 'read-only'). However, it does nothing to distinguish this from siblings like get_slo_report or list_recent_incidents beyond the resource name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the alternatives in the sibling set (e.g., recommend_runner_capacity, get_slo_report). No prerequisites, no triggering context, and no exclusions are stated. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_slo_reportB

Return current SLO checks and breaches without changing infrastructure.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does disclose one meaningful behavioral trait: it is non-mutating ('without changing infrastructure'). However, it says nothing about auth requirements, rate limits, freshness of the data, or what the report contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the resource and the key safety constraint with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, read-only reporting tool with no output schema and no annotations, the description covers the essentials but leaves the shape of the 'report' and any scoping (all SLOs? which services?) unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so parameter semantics is inherently uncomplicated; the baseline for a 0-param tool is 4. The description adds nothing param-related, but there is nothing to add.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('current SLO checks and breaches'), which clearly separates it from siblings like get_fleet_status, list_recent_incidents, and get_cost_estimate. It doesn't explicitly name a sibling it is not, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives no explicit when-to-use guidance, no prerequisites, and does not reference any alternative tool. The intended use (checking SLO status) must be inferred from the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recent_incidentsC

List recent synthetic SLO-breach reports and failure evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'List' weakly implies a read, but nothing states side effects, ordering, recency window, or pagination, and the term 'synthetic' is ambiguous about whether the data is real or mock.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the action and object front-loaded and no filler. It is well-sized for the tool, though brevity is partly achieved by omitting necessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with no annotations, an undocumented parameter, an undefined recency window, and no differentiation from get_slo_report, the description is only minimally adequate for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter (default 10) has no documentation anywhere. The description's 'recent' hints at a recency bound but never defines it or explains how limit interacts with it, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource ('List recent ... reports and failure evidence') makes the operation clear. However, it does not distinguish this tool from the close sibling get_slo_report, and the qualifier 'synthetic' is left unexplained, so the agent cannot tell which of the two SLO-oriented tools to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, and no mention of alternatives such as get_slo_report or get_fleet_status. The agent gets no signal about which sibling surfaces SLO breaches versus which surfaces fleet or cost data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommend_runner_capacityB

Calculate advisory runner capacity; this tool never scales resources.

ParametersJSON Schema
NameRequiredDescriptionDefault
queued_jobsYes
current_runnersYes
maximum_runnersNo
average_job_secondsNo
target_queue_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
cappedYes
rationaleYes
desired_runnersYes
additional_runnersYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses a key behavioral trait: the tool is advisory and never scales resources, making its non-mutating nature clear. It does not cover other behavioral details such as permissions or determinism, but the critical safety property is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It delivers purpose and a critical constraint immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, so the description need not explain them. However, with no annotations, 0% schema description coverage, and five parameters, the description is thin on parameter semantics and usage context. It is minimally viable but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for five parameters, so the description must compensate. It adds no meaning whatsoever for queued_jobs, current_runners, maximum_runners, average_job_seconds, or target_queue_seconds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Calculate') and resource ('runner capacity'), and explicitly distinguishes the tool from any scaling action. It does not, however, differentiate itself from siblings like get_fleet_status or get_cost_estimate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when not to use it ('never scales resources'), which is useful, but it does not state when to use it or name alternative tools for related tasks. Usage context must be inferred from the purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedget_cost_estimate
    • First observedget_fleet_status
    • First observedget_slo_report
    • First observedlist_recent_incidents
    • First observedrecommend_runner_capacity

TDQS

B3.4/5.0

Scored across 5 tools

Disambiguation4/5

Each tool targets a distinct concern (fleet status, SLO checks, capacity advice, incidents, cost). The main overlap is between get_slo_report (current breaches) and list_recent_incidents (recent breach reports), but descriptions clarify the temporal and evidentiary distinction.

Naming Consistency5/5

All names use snake_case with a clear verb_noun pattern (get_fleet_status, list_recent_incidents, recommend_runner_capacity). The verbs get/list/recommend are appropriate and consistently ordered.

Tool Count5/5

Five tools is well-scoped for a read-only advisory server covering fleet status, SLOs, capacity, incidents, and cost. No tool feels redundant or missing at this count.

Completeness4/5

The read-only advisory surface covers the main observability areas (status, SLOs, capacity, incidents, cost). Minor gaps exist for drill-down details (e.g., per-runner information or historical trends), but core agent workflows are supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables read-only Kubernetes incident investigation through MCP tools for listing pods, describing resources, fetching logs, and searching runbooks.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables read-only inspection of Google Ads accounts through MCP, including account inventory, reporting, metadata, change history, and safe GAQL queries.
    Apache 2.0