Skip to main content
Glama

cluster_health

Check Kafka cluster health in one call: get controller, broker details, and all problem partitions with offline, replication, ISR, or error issues.

Instructions

Check the health of the cluster this endpoint serves in one call: cluster id, controller, every broker (id, host, port, rack, and how many partitions it leads), a summary count, and every partition that is not fully healthy.

A problem partition lists its issues:

  • offline: no leader, so nothing can be produced to or consumed from it.

  • under_replicated: fewer in-sync replicas than replicas, so a broker is down or falling behind.

  • under_min_isr: fewer in-sync replicas than the topic's min.insync.replicas, so producers using acks=all fail with NOT_ENOUGH_REPLICAS.

  • error: the broker returned an error for the partition.

healthy is true only when there are no problems. Use search to limit the check to topics whose name contains a substring (case-insensitive); internal topics are skipped unless include_internal is true. min_isr_unknown lists topics whose min.insync.replicas the broker did not report, so under_min_isr was not judged for them.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
searchNoOptional case-insensitive substring that a topic name must contain to be checked. Omit to check every topic. Brokers and the controller are always reported.
include_internalNoOptional. When true, also check Kafka's internal topics such as __consumer_offsets. Defaults to false. An unhealthy internal topic breaks every consumer group, so include them when groups fail across the board.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
brokersYes
healthyYes
summaryYes
problemsYes
warningsYes
cluster_idYes
controllerYes
min_isr_unknownNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely succeeds: it discloses the return shape, the meaning of each problem type, the rule that internal topics are skipped unless include_internal is true, and the caveat that min_isr_unknown topics were not judged for under_min_isr. It omits operational details like required permissions or cost/latency, which is the only notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Content is front-loaded (purpose first, then the issue taxonomy, then parameter behavior) and the bulleted enumeration of offline/under_replicated/under_min_isr/error is efficient and skimmable. It is somewhat long, but nearly every sentence carries distinct diagnostic meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value structure is covered, yet the description still supplies the interpretive layer an agent needs: the healthy flag's condition and the semantics of each reported problem. Combined with the two fully documented parameters, nothing required to invoke or interpret the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description nevertheless adds semantic value: it clarifies that search matches a case-insensitive substring and that brokers and the controller are always reported regardless of the filter, and ties include_internal to the consequence of an unhealthy internal topic.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Check the health of the cluster') and enumerates exactly what one call returns: cluster id, controller, per-broker details, a summary count, and unhealthy partitions. It is clearly distinguishable from siblings like compare_clusters, list_clusters, or describe_topic, which serve different scopes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to use the tool ('check the health ... in one call') and gives concrete conditions for the optional filters, e.g. use search to limit to topics by substring and include_internal 'when groups fail across the board'. It does not explicitly route the agent away from compare_clusters or describe_topic, but the usage context is clear enough to act on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.