Skip to main content
Glama
sanjay-amu

kafka-sentinel-mcp

by sanjay-amu
README.md
# kafka-sentinel-mcp

[![CI](https://github.com/sanjay-amu/kafka-sentinel-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/sanjay-amu/kafka-sentinel-mcp/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/kafka-sentinel-mcp.svg)](https://pypi.org/project/kafka-sentinel-mcp/)
[![Python](https://img.shields.io/pypi/pyversions/kafka-sentinel-mcp.svg)](https://pypi.org/project/kafka-sentinel-mcp/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

**Give AI agents safe, read-only eyes on your Kafka clusters.**

An [MCP (Model Context Protocol)](https://modelcontextprotocol.io) server that exposes Kafka cluster health, consumer lag, partition state, and replay-readiness as structured tools — so LLM agents (Claude, or any MCP client) can diagnose streaming incidents without ever being able to break anything.

Built by an engineer who spent a decade running Kafka-based financial messaging at 99.999% availability, and got tired of every "AI + Kafka" demo assuming write access to production.

## Why this exists

When a consumer group stalls at 3 a.m., the questions are always the same: Is it lag? A stuck partition? A rebalance storm? An offset reset gone wrong? These are pattern-matching questions — exactly what LLM agents are good at — but no operator will hand an agent admin rights on a production cluster.

`kafka-sentinel-mcp` draws a hard line: **every tool is read-only by design**, enforced at the client-config level (no admin operations are even imported). The agent can observe, correlate, and recommend; a human executes.

## Tools

| Tool | What it returns |
|---|---|
| `list_topics` | All non-internal topics with partition count and replication factor — start here if you don't know a topic name |
| `list_consumer_groups` | All consumer group IDs with state — start here if you don't know a group name |
| `cluster_health` | Broker count, controller status, under-replicated / offline partition counts |
| `consumer_lag` | Per-group, per-topic, per-partition lag with committed vs end offsets |
| `topic_audit` | Replication factor, min.insync.replicas, retention, and flags configs that violate durability best practice |
| `partition_state` | Leaders, ISR shrinkage, skew across brokers |
| `replay_readiness` | For a group + topic: earliest available offsets vs committed, i.e., "can we still replay what we missed?" |
| `incident_snapshot` | One-call bundle of all the above, timestamped — designed for pasting into a postmortem |

## Quick start

```bash
pip install kafka-sentinel-mcp   # (or: uv tool install)

# Run against your cluster (read-only credentials!)
KAFKA_BOOTSTRAP=localhost:9092 kafka-sentinel-mcp
```

Add to Claude Desktop / any MCP client:

```json
{
  "mcpServers": {
    "kafka-sentinel": {
      "command": "kafka-sentinel-mcp",
      "env": { "KAFKA_BOOTSTRAP": "broker1:9092,broker2:9092" }
    }
  }
}
```

Then ask your agent: *"Why is the payments-consumer group falling behind, and can we still replay from where it stalled?"*

## Security posture

- **Read-only by construction:** no produce, no topic/config mutation, no offset commits, no ACL ops. The mutation APIs are never imported, and [a test in CI](tests/test_server.py) greps the server source on every run to keep it that way.
- The observer consumer runs with `enable.auto.commit=False` and never commits — [verified against a real broker](tests/test_integration.py), not just asserted.
- Supports SASL/SSL; credentials are read from the environment only and never logged.
- Every tool call is logged with its parameters for audit.
- **Least privilege:** run with a principal that has only `Describe` on the cluster and topics, and `Describe` on consumer groups. When an ACL denies an operation the tool returns a structured result rather than a stack trace:

  ```json
  {
    "error": "permission_denied",
    "operation": "list_consumer_groups",
    "detail": "...",
    "hint": "The Kafka principal in use lacks the ACL required for this operation. ..."
  }
  ```

  The agent can then tell the operator which ACL is missing instead of appearing broken. Non-authorization failures are deliberately *not* swallowed — they propagate, because silently degrading on an unrelated error would hide real problems.

## Testing

```bash
pip install -e ".[dev]"

pytest -m "not integration"   # fast, fully mocked — no Docker needed
pytest -m integration         # starts a real Kafka via testcontainers (needs Docker)
pytest                        # both
```

The unit suite mocks librdkafka entirely and covers tool logic. The integration suite starts an actual broker, produces real records, and asserts the tools return correct lag, ISR state, durability flags, and replay-readiness — including that the observer leaves no committed offsets behind. Both run in CI.

## Status

Early but tested. See [ROADMAP.md](ROADMAP.md). Issues and PRs welcome — especially war stories about what you wish an agent could have told you during an incident.

## Citing this work

If you reference this project in academic work, see [CITATION.cff](CITATION.cff), or use the "Cite this repository" button on GitHub.

## License

MIT

<!-- mcp-name: io.github.sanjay-amu/kafka-sentinel-mcp -->

TDQS

B3.4/5.0

Scored across 8 tools

Disambiguation4/5

Most tools target clearly distinct concerns: health, lag, audit, partition state, replay readiness. However, cluster_health and incident_snapshot overlap in scope since the snapshot bundles the health data, and topic_audit vs partition_state both touch replication/durability concerns at the partition level, creating minor confusion.

Naming Consistency4/5

Tool names use a consistent noun-based pattern (adjective_noun) like cluster_health, consumer_lag, topic_audit, partition_state — not strictly verb_noun but internally consistent. Minor deviation is having two 'list_' verbs mixed in with the descriptive noun pattern, though this reads as intentional since they are discovery operations.

Tool Count4/5

Eight tools for a Kafka operational monitoring server is a reasonable, focused scope. Each tool maps to a distinct operational concern with no obvious redundancy, and the incident_snapshot aggregation is a valuable convenience rather than bloat.

Completeness4/5

The surface covers the core read-only operational concerns an agent would need: discovery (list_topics, list_consumer_groups), health, lag, audit, partition state, and replay/retention analysis. Minor gaps exist — there's no consumer group reset or topic config mutation (read-only by design), and no offset trimming tool, but for an operational sentinel these are reasonable absences.

Maintenance

ActivitySlowing
ResponsivenessNo issues