PrismSRE
Provides tools for diagnosing Kubernetes clusters, enabling analysis of pod status, deployment definitions, logs, and events to troubleshoot issues like CrashLoopBackOff and OOMKilled.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PrismSREdiagnose why my nginx pod is in CrashLoopBackOff in the default namespace"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ PrismSRE
The next-generation, AI-powered Site Reliability Engineer for your Kubernetes Clusters.
PrismSRE is a production-grade Kubernetes troubleshooting system that acts as an autonomous AI agent. It seamlessly bridges the gap between raw cluster metrics/logs and actionable SRE insights. Powered by the Google Agent Development Kit (ADK), Model Context Protocol (MCP), and a beautiful Glassmorphism Dashboard, PrismSRE provides immediate, intelligent diagnostics for your Kubernetes workloads.
โจ Features
๐ง Autonomous Diagnostics: Powered by Google's Gemini models, capable of analyzing
CrashLoopBackOff,OOMKilled, and stuck rollouts.๐ก๏ธ Secure by Design: Employs the Model Context Protocol (FastMCP) to enforce strict read-only access to the Kubernetes cluster. The AI agent operates outside the direct execution context.
๐จ Glassmorphism UI: A breathtaking, dependency-free, single-file HTML dashboard using Vanilla JS and Tailwind CSS.
โก Real-time Context Gathering: Automatically fetches pod status, deployment definitions, and tail logs through MCP tools without requiring raw shell access.
โ๏ธ Cloud Agnostic: Compatible with GKE, K3s, Minikube, and standard Kubernetes distributions.
Related MCP server: K8s MCP
๐๏ธ Architecture
For a deep dive into the system design, security boundaries, and component interaction, please see the Architecture Documentation.
๐ Getting Started
Prerequisites
Python 3.11+
A running Kubernetes cluster (GKE, K3s, Minikube, etc.)
kubectlconfigured and authenticated to your clusterA Google Gemini API Key
Local Development
Clone the repository:
git clone https://github.com/barbaria888/PrismSRE.git cd PrismSREInstall dependencies:
pip install -r requirements.txtConfigure Environment Variables:
cp .env.example .envAdd your
GOOGLE_API_KEYto the.envfile.Run the Dashboard Server:
uvicorn app:app --reload --host 0.0.0.0 --port 8000Navigate to
http://localhost:8000in your browser.
โธ๏ธ Running in Your Own Cluster
To deploy PrismSRE as a long-running service inside your Kubernetes cluster, follow these steps.
1. Create the Secret
The agent requires your Gemini API key to operate. We provide a compatible secret manifest.
Edit secret.yaml with your actual base64/plaintext key, then apply:
kubectl apply -f secret.yaml2. Containerize the Application
Build and push the Docker image to your container registry:
# Example Dockerfile included in the project or write a simple one for FastAPI
docker build -t your-registry/prismsre:latest .
docker push your-registry/prismsre:latest3. Deploy to Kubernetes
You can deploy the application using standard Kubernetes manifests. Ensure you grant the necessary RBAC permissions (read-only access to Pods, Deployments, and Logs).
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: prismsre-sa
namespace: default
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: prismsre-reader
rules:
- apiGroups: ["", "apps"]
resources: ["pods", "pods/log", "deployments", "events"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: prismsre-reader-binding
subjects:
- kind: ServiceAccount
name: prismsre-sa
namespace: default
roleRef:
kind: ClusterRole
name: prismsre-reader
apiGroup: rbac.authorization.k8s.io
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: prismsre
namespace: default
spec:
replicas: 1
selector:
matchLabels:
app: prismsre
template:
metadata:
labels:
app: prismsre
spec:
serviceAccountName: prismsre-sa
containers:
- name: prismsre
image: your-registry/prismsre:latest
ports:
- containerPort: 8000
envFrom:
- secretRef:
name: kubeops-ai-secret
---
apiVersion: v1
kind: Service
metadata:
name: prismsre-service
spec:
type: ClusterIP
selector:
app: prismsre
ports:
- protocol: TCP
port: 80
targetPort: 8000Apply the deployment:
kubectl apply -f deployment.yaml(Note: If you want external access, configure an Ingress or change the Service type to LoadBalancer).
๐ก๏ธ Security Considerations
No Root Access: The agent operates strictly with
ClusterRoleread-only permissions.No Direct Shell: Uses the Model Context Protocol to execute predefined tools, preventing Prompt Injection attacks that try to execute arbitrary bash commands.
๐ License
This project is licensed under the MIT License.
This server cannot be deployed
Maintenance
Related MCP Connectors
The Google GKE MCP server is a managed Model Context Protocol server that provides AI applications with tools to manage Google Kubernetes Engine (GKE) clusters and Kubernetes resources. It exposes a structured, discoverable interface that allows AI agents to interact with GKE and Kubernetes APIs, enabling them to inspect cluster configurations, retrieve Kubernetes resource YAMLs, monitor operations like cluster upgrades, diagnose issues, and optimize costsโall without needing to parse text output or use complex kubectl commands.
- emisarOAuthdev.emisar
Let AI operate servers without SSH. Choose actions, approve risky changes, and audit every step.
MCP-native AI SRE: ask what's broken in production, get a reviewed GitHub fix PR.
Fail-closed policy guardrails for AI agents running kubectl, terraform, helm, and argocd.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI assistants to interact with Kubernetes clusters through natural language, supporting core Kubernetes operations, monitoring, security, and diagnostics.108 npm956MIT
- FlicenseNot gradedqualityCmaintenanceEnables interaction with Kubernetes clusters through 32 specialized tools for managing resources, deployments, and services. Provides both CLI and web interfaces for real-time Kubernetes operations powered by Google Gemini.5-
- AlicenseBqualityDmaintenanceAI-powered Kubernetes diagnostics that analyzes pod crashes, logs, and cluster health to provide root cause analysis and actionable solutions for common issues like CrashLoopBackOff, OOM kills, and connection errors.811 npm1MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server for Kubernetes that provides full resource coverage and advanced troubleshooting tools via HTTP chunked streaming. It enables users to manage clusters and diagnose complex issues like pod crashloops through specialized prompts and standard kubectl-like operations.4,455 npmMIT