Skip to main content
Glama
AIops-tools

cicd-aiops

runner_health_rca

Flag problematic CI runners and queue bottlenecks: identifies offline, stale, or paused runners, long-queued jobs, and tag saturation, returning counts to guide RCA.

Instructions

[READ] Flag offline/stale/paused runners, long-queued jobs, tag saturation.

The flagship capacity RCA: pulls the runner fleet, flags each runner that is offline, stale (no contact for stale_contact_min minutes) or paused, lists jobs queued past queue_sec, and computes per-tag saturation (queued jobs vs online unpaused runners). Every flag carries its numbers. Pass 'runners' / 'queued_jobs' for pure analysis, or a target to pull the fleet live.

Args: stale_contact_min: Minutes since last contact at which a runner is stale. queue_sec: Seconds a job may wait before being flagged (default 300). saturation_ratio: Flagged queued jobs per online runner at which a tag is saturated (default 2.0). runners: Injected rows {id, description, status, paused, online, tags, contactedAt}; skips the live pull. queued_jobs: Injected rows {id, name, queuedDurationSec, createdAt, tags}. target: Server target name from config; omit for the default.

Returns dict: {runnersEvaluated, flaggedRunners, longQueuedJobs, saturatedTags, thresholds, note}.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
targetNo
runnersNo
queue_secNo
queued_jobsNo
saturation_ratioNo
stale_contact_minNo
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explains the analysis algorithm (flagging offline/stale/paused runners, queued jobs, tag saturation) and the return dict structure. With no annotations, it carries the burden well but omits potential behaviors like authentication requirements or error scenarios.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a concise summary, followed by a brief algorithm explanation and structured Args list. Could be slightly more concise but every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all input parameters and high-level output structure, but lacks details on exact return fields (e.g., 'note' content) and error handling. No output schema provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description's Args section provides clear explanations for each parameter (e.g., 'stale_contact_min: Minutes since last contact at which a runner is stale'). Adds significant meaning beyond the schema's property titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with '[READ] Flag offline/stale/paused runners, long-queued jobs, tag saturation' which clearly states the verb and resource. Positions itself as 'the flagship capacity RCA', distinguishing from sibling tools like list_runners and runner_detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context: explains when to pass 'runners'/'queued_jobs' for analysis vs a target for live pull. However, lacks explicit when-not-to-use or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/AIops-tools/CICD-AIops'

If you have feedback or need assistance with the MCP directory API, please join our Discord server