kafka-dataops-mcp
by Aguantar
README.md
# kafka-dataops-mcp
mcp-name: io.github.Aguantar/kafka-dataops-mcp
A DataOps-focused Kafka MCP server with consumer lag diagnosis and broker health monitoring. Diagnosis logic is based on actual CDC pipeline operational experience.
## Features
- **`kafka_consumer_lag`** — Consumer group lag with incident-pattern diagnosis
- **`kafka_topic_info`** — Topic details with ISR/replication health checks
- **`kafka_broker_status`** — Cluster health: brokers, controller, under-replicated partitions
- **`kafka_list_topics`** — Topic catalog with built-in descriptions
### Diagnosis based on real incidents
The diagnosis logic is not generic — it's based on actual operational experience:
- **Flink crash detection**: "no active members" + growing lag = likely Flink Job failure (based on a 50-hour outage caused by MySQL DELETE → Debezium tombstone → Flink NPE)
- **Checkpoint vs consumer group**: warns that Kafka consumer group reset alone is insufficient for Flink — checkpoints must be deleted first
- **ClusterIdMismatch**: detects missing brokers and suggests Docker volume conflict as root cause
- **ISR monitoring**: ISR < min.insync.replicas = write failures (critical)
## Installation
```bash
pip install kafka-dataops-mcp
```
## Usage with Claude Code
Add to your `.mcp.json`:
```json
{
"mcpServers": {
"kafka": {
"command": "kafka-dataops-mcp",
"env": {
"KAFKA_BOOTSTRAP_SERVERS": "localhost:9092"
}
}
}
}
```
## Environment Variables
| Variable | Default | Description |
|----------|---------|-------------|
| `KAFKA_BOOTSTRAP_SERVERS` | `localhost:9092` | Kafka bootstrap servers |
| `KAFKA_COMMAND_TIMEOUT` | `10` | Command timeout in seconds |
## License
MIT
TDQS
A3.9/5.0
Scored across 4 tools
Disambiguation5/5
Each tool targets a distinct aspect of Kafka operation: cluster health, consumer lag, topic listing, and topic details. No functional overlap.
Naming Consistency4/5
All tools start with 'kafka_' but mix verb-noun ('list_topics') and noun phrases ('broker_status'). Still predictable and readable.
Tool Count4/5
Four tools is reasonable for a monitoring-focused server. Could expand with more actions but not too sparse.
Completeness3/5
Covers key diagnostic needs (health, lag, topic details) but lacks topic creation or consumer group management. Adequate for observability.
Maintenance
ActivityInactive
ResponsivenessNo issues