Claude Ops Investigator
Provides tools for investigating Kubernetes incidents, including listing pods, describing pods, retrieving pod logs, fetching recent namespace events, and checking top pod resource usage.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Claude Ops InvestigatorInvestigate high CPU usage in pod my-app-xyz in namespace production"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Claude Ops Investigator
An MCP-based Kubernetes incident investigation tool with two supported agent harnesses — Claude Code and IBM Bob Shell — combining live cluster signals, Prometheus, log search, runbooks, evidence memory, OpenJ9 JVM troubleshooting (GC logs, thread dumps), and structured incident reports.
Claude Ops Investigator helps engineers investigate Kubernetes incidents safely by combining read-only operational tools, external evidence storage, compact investigation memory, and human-controlled remediation boundaries.
The goal is not to give an AI unrestricted production access. The goal is to expose narrow, auditable, read-only interfaces that help engineers gather evidence, form hypotheses, rule out causes, and produce reliable incident reports faster.
What this project provides
Narrow MCP-style tools instead of generic
kubectlRead-only live Kubernetes investigation
Per-harness project instructions (CLAUDE.md, AGENTS.md, Bob mode rules)
MCP resources, tools, and prompts
Skills, slash commands, and scoped project rules
Coordinator/subagent-style investigation workflows
Structured tool errors
Hooks and gates for destructive actions
Structured incident-report output
Human escalation for risky or ambiguous actions
A JVM specialist (
jvm-analyst) backed by a separatejvm-troubleshooterMCP server: GC, heap, memory pools, threads, from Prometheus/JMX ExporterReal GC logs and thread dumps: human-run capture scripts plus read-only analysis tools (true per-pause GC stats, stuck vs busy threads, thread churn)
Prometheus through Grafana's datasource proxy, no port-forward needed
Related MCP server: mcp-kubernetics
Safety rule
Start read-only. Do not give Claude unrestricted shell, kubectl, Helm, or production mutation permissions.
Allowed operations in this scaffold:
kubectl getkubectl describekubectl logskubectl top
Blocked operations include:
kubectl deletekubectl applykubectl patchkubectl scalekubectl rollout restarthelm upgradekubectl exec
Agents never exec into a pod. The only kubectl exec in the project lives
in three human-run scripts: scripts/capture-gclog.sh,
scripts/capture-javacore.sh and scripts/cleanup-javacores.sh. They refuse
to run without an interactive terminal, and the Claude Code shell hook blocks
agents from running them. Agents only analyze what you capture (see
GC logs and thread dumps).
Quick start
cd claude-ops-investigator
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev,mcp]" # core package + MCP server
pip install -e "mcp-servers/jvm-troubleshooter[dev]" # second MCP server (JVM)Install both packages into the Python your harness uses to launch MCP
servers. Claude Code started from this shell uses .venv. IBM Bob may use a
different interpreter (e.g. a pyenv Python), so run the two pip install
lines with that interpreter too. Otherwise the server can't import its
package, fails to start, and Bob silently drops it. Bob's MCP log
(~/Library/Application Support/IBM Bob/logs/<session>/…/IBM Bob MCP.log on
macOS) shows the ModuleNotFoundError.
Check Kubernetes access:
kubectl config current-context
kubectl auth can-i get pods -n si
kubectl auth can-i get pods/log -n siRun a read-only snapshot:
python -m claude_ops.main investigate --namespace si --service event-data --since-minutes 60Run tests (two independent suites):
pytest # core: tests/
(cd mcp-servers/jvm-troubleshooter && pytest) # jvm-troubleshooterSlash commands
Both harnesses expose the same three commands: a read-only investigation
command, a standalone JVM health-check command, and a separate, autonomous
fix-proposal command. /investigate-incident and /investigate-jvm are
both strictly read-only; /propose-fix is the only one that can ever open a
(draft) PR, and none of the read-only commands trigger it on their own.
Claude Code slash commands
The main interactive workflow is the /investigate-incident slash command:
/investigate-incident namespace=<namespace> service=<service> symptom="<specific symptom>" since_minutes=<minutes>symptom is required — this workflow is symptom-driven, not a generic
service health check. If no symptom is given, Claude asks for one before
investigating anything.
Examples:
/investigate-incident namespace=si service=multi-system-processor symptom="readiness probe failures during recent rollout" since_minutes=60
/investigate-incident namespace=si service=event-data symptom="KafkaConsumerCommitRateLow alert fired" since_minutes=120
/investigate-incident namespace=si service=multi-system-processor symptom="OOMKilled restarts observed" since_minutes=180What it does:
Mints an
investigation_idand createsruns/<investigation_id>/scratchpad/, then delegates to theincident-coordinatorsubagent (falls back to a single agent, with that fallback stated explicitly, only if subagents aren't available)The coordinator reads the service catalog and runbook catalog, then routes the symptom to the narrowest relevant specialist subagents:
k8s-evidence-collector— pod listing, describe, live logs, namespace events, current resource usageprometheus-analyst— restart counts/increase, CPU, memory, HTTP error rate, latency p95log-analyst— historical IBM Cloud Logs search (errors, probe failures, arbitrary text) spanning restarts/deploymentsrunbook-analyst— matches the symptom against local runbooksjvm-analyst(conditional) — GC pause/throughput, heap, memory-pool/ native-memory, and thread signals via the separatejvm-troubleshooterMCP server; only routed to when the symptom is GC-, heap-, memory-pressure-, or OOM-flavored on a known OpenJ9/IBM Semeru service
Every specialist stores raw evidence as an
evidence_refand hands back only summaries/findings — the coordinator never gathers evidence directly. Each also writes a concise markdown scratchpad (scope, tools called, key findings, evidence_refs, unknowns/gaps, decisions, handoff summary) to its assignedruns/<investigation_id>/scratchpad/wave<N>-<subagent-name>.mdfile — never raw log/metric bodies, only summaries and evidence_refs. The coordinator maintains its own running Structured Finding Brief atcoordinator-brief.mdin the same directory, and passes each subagent that brief plus any relevant prior scratchpad paths in its task prompt.incident-reporterruns last, synthesizing all subagents' findings (never its own) into a single evidence-grounded, schema-valid reportThe final output includes a "Subagent usage audit" table: which subagent ran, what it did, which tools/evidence_refs/scratchpad path it used, and its result
What it does not do:
Does not mutate Kubernetes resources
Does not restart pods
Does not apply fixes
Does not run destructive commands
Does not fetch raw evidence detail unless needed
/investigate-jvm
A lightweight, standalone JVM health check against one OpenJ9/IBM Semeru service — GC behavior, heap, memory pools, threads, allocation rate, leak trend — with no coordinator/subagent delegation:
/investigate-jvm namespace=<namespace> service=<service> lookback_minutes=<minutes>It starts from jvm-troubleshooter's metrics (a one-call snapshot, then
focused tools and charts) and names the pods that stand out. When metrics
aren't enough, it tells you which pod to capture a GC log or thread dumps
from, and which script to run. It then analyzes what you saved with the
jvm_* tools, and reminds you to clean up the dumps afterwards. See
GC logs and thread dumps.
Use this for a quick point check; use /investigate-incident (which routes
to jvm-analyst automatically for a GC/heap/OOM-flavored symptom) when the
finding needs to sit alongside other specialists' evidence in one incident
report. See .claude/commands/investigate-jvm.md
for the full workflow this command follows.
Bob Shell slash commands
Bob Shell (.bob/) is a parallel harness that talks to the same MCP tool
layer and exposes the same three commands.
/investigate-incident
Read-only, same behavior and args as the Claude Code version above:
/investigate-incident namespace=<namespace> service=<service> symptom="<specific symptom>" since_minutes=<minutes>An orchestrator mode decomposes the incident and delegates to the same
specialist roles (k8s-evidence-collector, prometheus-analyst,
log-analyst, runbook-analyst, plus jvm-analyst when the symptom is
GC-, heap-, memory-pressure-, or OOM-flavored on a known JVM workload), then
hands off to incident-reporter for a schema-valid, evidence-grounded report — see .bob/commands/investigate-incident.md
and AGENTS.md for the full workflow.
/investigate-jvm
Same lightweight, standalone JVM health check as the Claude Code version above, including the GC-log and thread-dump hand-off, no mode switch required:
/investigate-jvm namespace=<namespace> service=<service> lookback_minutes=<minutes>See .bob/commands/investigate-jvm.md for the full workflow. Bob may also
auto-generate a skill version of it under .bob/skills/ from an older copy
of the command. If its answers never mention the capture scripts or jvm_*
tools, start your request with "follow .bob/commands/investigate-jvm.md".
/propose-fix
A separate, autonomous command. It only fires when an incident report traces the cause to a named application-code location (an exception class, stack trace, or file/function reference — not an infra/operational finding); if that gate isn't met, no fix is proposed. When it does fire, it:
Works only in the target service's own existing local git checkout (looked up from
data/service_catalog.json) — it never clones a repo.Requires a clean working tree first — any uncommitted changes to tracked files stop it immediately, nothing is stashed or discarded.
Opens a draft-only PR, with an AI-disclosure line and a human-review checklist in the description. It never opens a non-draft PR.
Args:
/propose-fix namespace=<namespace> service=<service> symptom="<symptom>" since_minutes=<minutes>
/propose-fix investigation_id=<id>|latestOptional: dry_run=true (locates the code and narrates the proposed fix to
a scratchpad file without branching, committing, pushing, or opening a PR)
and base_branch=<branch> (defaults to whatever branch is already checked
out in the local checkout if omitted).
/investigate-incident never triggers /propose-fix — they are separate
commands, and a code change only ever happens when /propose-fix is run
explicitly.
Environment for optional tools
Prometheus:
PROMETHEUS_URLPROMETHEUS_AUTO_PORT_FORWARDPROMETHEUS_PF_SERVICEPROMETHEUS_PF_NAMESPACEor, via Grafana:
GRAFANA_URL,GRAFANA_DATASOURCE_UID, andGRAFANA_API_TOKENorGRAFANA_SESSION_COOKIE(see below)
IBM Cloud Logs:
IBM_LOGS_ENDPOINTIBM_CLOUD_API_KEY
Copy .env.example to .env and fill in local values. Never commit .env.
The MCP server loads it automatically at startup so these tools have access
without any secrets going into .mcp.json.
Querying Prometheus through Grafana instead of a port-forward
If you can log into Grafana but can't (or don't want to) kubectl port-forward to Prometheus, point the Prometheus tools at Grafana's
datasource proxy instead. Both MCP servers (claude-ops-investigator and
jvm-troubleshooter) support it; put these in .env, never in .mcp.json:
GRAFANA_URL— e.g.https://grafana.example.com. Setting it switches to Grafana mode and takes precedence overPROMETHEUS_URL. Must behttpsunless it'slocalhost.GRAFANA_DATASOURCE_UID— the Prometheus datasource's uid (Grafana → Connections → Data sources → the datasource; its URL ends in/edit/<uid>).One credential:
GRAFANA_API_TOKEN(a Viewer-role service account token, sent asAuthorization: Bearer; preferred) orGRAFANA_SESSION_COOKIE(thegrafana_sessioncookie value from your logged-in browser, via DevTools → Application → Cookies; expires with your session). The token wins if both are set.
Queries then go to
<GRAFANA_URL>/api/datasources/proxy/uid/<uid>/api/v1/query[_range] — the
same read-only PromQL calls, just proxied. The credential is redacted from
any error the tools return and never stored as evidence.
prom_ensure_connection checks reachability through the proxy in this mode
and never starts a port-forward.
Finding the right datasource:
A name is not a UID. Grafana returns 404 if you put the datasource's name (e.g.
my-cluster-prom) where its UID goes. List them withGET <GRAFANA_URL>/api/datasourcesusing the same credential, then take theuidfield.Pick the datasource that matches your cluster. Several datasources can carry the same namespace from different clusters. Choose the one whose pod names match
kubectl get podson your current context.Check
JVM_LABEL_KEYagainst the real series. It's set per server in.mcp.json/.bob/mcp.json. A wrong value returns empty results, not an error. Our JMX Exporter series are labeled byjob; confirm withcount by (job) (java_lang_GarbageCollector_CollectionCount).
Quick check:
curl -s -H "Authorization: Bearer $GRAFANA_API_TOKEN" \
"$GRAFANA_URL/api/datasources/proxy/uid/$GRAFANA_DATASOURCE_UID/api/v1/query?query=up" | head -c 300No-token local tests
These exercise the tools and structured error paths without any real Prometheus, IBM Cloud, or Kubernetes credentials:
python -m pytest
python scripts/mcp_smoke_client.pyDirect tool checks:
python - <<'PY'
from claude_ops.tools.prometheus_preflight import ensure_prometheus
import json
print(json.dumps(ensure_prometheus(), indent=2))
PY
python - <<'PY'
from claude_ops.tools.ibm_logs_tools import ibm_logs_search_errors
import json
print(json.dumps(ibm_logs_search_errors("si", "multi-system-processor", limit=1), indent=2))
PYLocal environment
The MCP server needs environment variables for the optional Prometheus and
IBM Cloud Logs tools (PROMETHEUS_URL, IBM_LOGS_ENDPOINT,
IBM_CLOUD_API_KEY, etc.). Configure them locally with a .env file — it is
gitignored and loaded automatically, no secrets ever need to go in
.mcp.json.
cp .env.example .env
# edit .env with your local values
source .venv/bin/activate
claudesrc/claude_ops/mcp/server.py calls load_dotenv() at startup, so the MCP
server picks up .env automatically when Claude Code launches it — no
manual export needed. Missing .env is fine; tools that need a variable
that still isn't set return a structured config error instead of failing
silently.
Harness hooks (safety gate + audit trail)
.claude/settings.json wires four read-only Claude Code hooks under
.claude/hooks/. They're a harness-level safety net and audit trail that sit
alongside the application-level guardrails (src/claude_ops/hooks.py,
schemas/incident_report_schema.py) — none of them call Kubernetes,
Prometheus, IBM Cloud Logs, or the Claude API; they only inspect the JSON
Claude Code already passes them on stdin, and the only files they write are
JSONL audit logs under runs/ (gitignored, like the rest of that directory).
Hook | Event | What it does |
|
| Denies raw shell |
|
| Appends |
|
| Appends |
|
| If the last assistant message looks like an incident report (mentions "Subagent usage audit", "incident report", or |
Disabling hooks locally
Two ways, from least to most surgical:
Disable everything: add
"disableAllHooks": trueto.claude/settings.local.json(gitignored, personal — never commit this to the project's shared.claude/settings.json).Disable just these four: set
CLAUDE_OPS_HOOKS_DISABLED=1in your shell environment before launchingclaude. Each script checks this at the top and no-ops immediately — no audit lines written, no shell command blocked, no report validated.
Recommended first live use
Use a non-production namespace first.
python -m claude_ops.main investigate --namespace si --service multi-system-processor --since-minutes 120Then paste the generated JSON snapshot into Claude/Claude Code and ask it to produce an incident report using the schema in src/claude_ops/schemas/incident_report_schema.py.
MCP client/server map
In this project:
Claude Code = MCP client
src/claude_ops/mcp/server.py = local MCP server
.mcp.json = project-level MCP client configuration for Claude CodeStart the MCP server manually for a quick syntax check:
python -m claude_ops.mcp.serverFor Claude Code, keep .mcp.json in the project root. Claude Code reads the config and launches the server over STDIO.
Optional smoke test:
pip install -e ".[dev,mcp]"
python scripts/mcp_smoke_client.pyThe MCP server exposes:
Resources:
ops://runbook-catalogops://service-catalog
Tools:
k8s_list_podsk8s_describe_podk8s_get_pod_logsk8s_get_recent_namespace_eventsk8s_top_podsrunbook_searchprom_query_instantprom_get_pod_restart_countsprom_get_pod_restart_increaseprom_get_pod_cpu_usageprom_get_pod_memory_usageprom_get_http_error_rateprom_get_latency_p95prom_ensure_connectionibm_logs_searchibm_logs_search_errorsibm_logs_search_probe_failuresibm_logs_search_textjvm_get_gc_log_events— verbose GC events fromkubectl logs(GC log on stderr)jvm_analyze_gc_log— GC log files underruns/gclogs/jvm_analyze_javacore— one javacore underruns/javacores/jvm_compare_javacores— a series of javacores: stuck vs busy threads, thread churnevidence_get_detailevidence_store_external— archive another server's result (e.g.jvm-troubleshooter's) as evidence with anevidence_ref
Prompt:
investigate_incident
Second MCP server: jvm-troubleshooter
A dedicated, independently-testable MCP server for OpenJ9/IBM Semeru JVM
internals (GC, heap, memory pools, threads) lives at
mcp-servers/jvm-troubleshooter/, with its own pyproject.toml, src/,
tests/ and README.md.
Repository layout
src/claude_ops/ core application: tool layer, evidence store, safety gate (hooks.py),
CLI (claude_ops.main), and its MCP front-end (claude_ops/mcp/server.py)
data/ runs/ artifacts/ runbooks + service catalog; investigation scratchpads, captures,
audit logs; archived evidence (runs/ and artifacts/ are gitignored)
mcp-servers/jvm-troubleshooter/ standalone add-on MCP server; imports nothing from claude_ops
scripts/ human-run capture/cleanup scripts, MCP smoke client
.claude/ .bob/ the two agent harnesses: agents/modes, commands, rules, hooks
docs/ guides (Bob harness, GC logs and thread dumps)
dashboards/ Grafana dashboard JSON (JVM troubleshooting) and its generatorThe asymmetry is deliberate. claude_ops is more than an MCP server: the
CLI and the safety gate share its tool layer, and it relies on the repo-level
data/, runs/ and artifacts/. mcp-servers/ holds servers that stand on
their own. Any Java team can install jvm-troubleshooter without the rest of
this project, which is why it carries its own small copies of errors.py and
the Prometheus/Grafana endpoint resolver instead of importing them.
How the two servers work together
They're combined in the harness, not in code:
.mcp.json(Claude Code) and.bob/mcp.json(Bob) register both servers, so agents see one toolbox, namespaced per server (mcp__jvm-troubleshooter__…,mcp__claude-ops-investigator__…).jvm-analystis the one agent whose tool allowlist spans both servers. The coordinator routes GC-, heap-, memory- or OOM-flavored symptoms on JVM services to it.Every finding in a report needs an
evidence_reffromclaude-ops-investigator's evidence store, andjvm-troubleshooterhas no store of its own. Sojvm-analystarchives each result it cites throughevidence_store_external, andincident-reportercites it like any other evidence. Thejvm_*GC-log and javacore tools live inclaude-ops-investigatorand produce refs directly.
Configuration
Install the package into the Python your harness launches it with (see Quick start). The
PYTHONPATH=${PWD}/…in the MCP configs isn't reliable everywhere; Bob, for example, doesn't apply it.JVM_LABEL_KEY(in each MCP config'senvblock) must match the Prometheus label your JMX Exporter series use to name the service. A wrong value returns empty results, not an error.Grafana settings (
GRAFANA_*) and credentials go in the repo-root.env, which this server also loads at startup. SettingGRAFANA_URLtakes precedence over thePROMETHEUS_URLin the MCP config.
It's backed by Prometheus (or Thanos Query) scraping OpenJ9 JVMs via the
standard Prometheus JMX Exporter, a different metric-naming convention than
this project's own prom_* tools assume, so it ships its own PromQL.
Its 27 tools are documented in full, including caveats on what each one
can't see, in
mcp-servers/jvm-troubleshooter/README.md:
GC: activity, pause stats, throughput, behavior over time.
Heap and memory: heap status and trend, memory-pool breakdown, native memory, fragmentation, allocation rate, leak indicator, and memory vs the container limit (headroom, heap / non-heap / direct / native split).
Threads: thread status, and thread trend with threads started per second (churn) and deadlocked threads.
Process and runtime: process CPU and file descriptors, class loading trend, JVM version and uptime per pod.
Time windows: an explicit start/end incident window, a baseline comparison (avg/p50/p95/p99 and % change), CPU-vs-GC correlation.
Cross-signal: GC-memory correlation, before/after deploy comparison, a one-call incident snapshot, and four charts (heap trend, GC behavior, heap vs GC, thread trend). Each returns a PNG, saves it to
runs/charts/, and includes Mermaid blocks so the chart shows inline in Bob. Run its own test suite (independent of this project'spytestinvocation, which only looks at the roottests/) with:
cd mcp-servers/jvm-troubleshooter
pytestFor a quick, standalone JVM health check outside the full incident-
investigation flow, see /investigate-jvm under Slash commands above.
Grafana dashboard
dashboards/jvm-troubleshooting.json
("JVM troubleshooting (OpenJ9)") shows the same JVM signals the
jvm-troubleshooter tools read, on one Grafana board. Use it to watch a
service yourself, or to see what the agent reports at a glance.
Import
In Grafana: Dashboards → New → Import → Upload dashboard JSON file.
Pick the Data source: the Prometheus or Thanos datasource that scrapes your JVMs.
Pick the Namespace, the Service and the Pods (All, or a subset).
If Service stays empty, change Service label to the label your JMX Exporter series use to name the service. It's the same setting as the MCP server's
JVM_LABEL_KEY:job(default),apporservice.
The board's URL keeps the selected variables, so it can be shared as a link to the same view.
What's on it (24 panels; the ⓘ on each panel says how to read it)
Section | Panels |
At a glance | JVM pods, lowest memory headroom, highest heap used, highest GC overhead, threads started per second, deadlocked threads |
Memory | Container memory vs limit, memory headroom, heap used vs max, tenured (old gen) after GC, memory pools, non-heap and direct buffers |
Garbage collection | GC overhead per pod, collections per minute (scavenge / global), average pause by collector, heap used vs GC frequency |
Threads | Live threads, threads started per second, daemon threads |
CPU, file descriptors, class loading | Process CPU (with the CPU limit when set), CPU vs GC overhead, open file descriptors (% of max), classes loaded and unloaded |
Runtime | JVM version and uptime per pod |
Reading the key panels
Lowest memory headroom / Memory headroom: how far each pod's container working set is below its memory limit. Under 10% is OOMKilled risk, even if the heap looks fine; heap max (
-Xmx) is not the container limit.Tenured after GC: old-gen usage right after the last collection, i.e. what survived GC. Flat or falling is healthy. A floor that keeps rising over many hours suggests objects are being retained (a leak candidate). Use a long time range, 12–24 h.
Heap used vs GC frequency: heap returning to the same baseline while GC rises and falls means load. A rising floor together with climbing GC frequency means a leak or an undersized heap.
Threads started per second: a high rate with a flat live-thread count is churn: pool threads expiring and being recreated. To find which pool, capture a javacore series and compare it (see GC logs and thread dumps).
CPU vs GC overhead: lines moving together with low GC overhead mean load drives both. Rising GC overhead with CPU means GC is driving the CPU.
Average GC pause: a 5-minute average, which hides the worst pauses. For real per-pause max and p99, use the GC log tools.
Requirements
OpenJ9 / IBM Semeru JVMs scraped by the Prometheus JMX Exporter java agent (its
java_lang_*,jvm_*andprocess_*metrics).For the memory-limit panels, cAdvisor and kube-state-metrics in the same Prometheus. Container series are matched to the JVM pods on
namespace,podandcontainer, so other workloads in the namespace never appear.
Changing it: edit
dashboards/generate_jvm_troubleshooting.py,
not the JSON, then run python dashboards/generate_jvm_troubleshooting.py.
tests/test_dashboard.py fails when the committed JSON is out of date. More
detail is in dashboards/README.md.
GC logs and thread dumps
Step-by-step guide, with example output and troubleshooting:
docs/jvm-gc-logs-and-thread-dumps.md.
The Prometheus-based JVM tools give counts and 5-minute averages. For the
real thing, claude-ops-investigator adds three read-only tools, and two
scripts that only a human runs:
jvm_get_gc_log_events(namespace, pod_name, since_minutes, previous)reads the pod's logs (kubectl logs, the same read-only verb ask8s_get_pod_logs) and parses OpenJ9 verbose GC XML: true per-pause p50/p95/p99/max, scavenge vs global, the longest pauses with timestamps and triggers (allocation failure,System.gc(), concurrent kickoff), percolate/copy-failed events and GC warnings. Prerequisite: verbose GC has to be switched on. For a Liberty service, add-verbose:gcto itsjvm.options(e.g. thejvmoptions-<service>-configConfigMap) and roll the pods. OpenJ9 writes it natively to stderr, so it lands inkubectl logsand IBM Cloud Logs without Liberty's logging in the way. Expect roughly a few MB of extra log volume per pod per hour. Until it's on, the tool returns abusinesserror saying so.scripts/capture-gclog.sh [-m max_files] <namespace> <pod> [container]— human-run only, for JVMs that write their GC log to a file (-Xverbosegclog, or-Xloggcas in our Liberty images; it also follows a stderr redirected to a file). It reads the JVM's real setting from its command line andOPENJ9_JAVA_OPTIONS/IBM_JAVA_OPTIONS/JAVA_TOOL_OPTIONS/JDK_JAVA_OPTIONS(including%pid/%seqpatterns and rotated files), also checks/tmp/verbosegc.*.txtand/opt/ibm/*verbosegc.*.txt, and streams the newest files (default 10) intoruns/gclogs/<pod>-<utc>/. It changes nothing in the pod. If the JVM only has-verbose:gc, it tells you to usejvm_get_gc_log_eventsinstead. If no verbose GC is configured at all, it says so: OpenJ9 can't switch it on at runtime, so that needs ajvm.optionschange and a restart. Same safeguards as the javacore script.jvm_analyze_gc_log(path)gives the same analysis asjvm_get_gc_log_eventsfor a file or a whole capture directory underruns/gclogs/(rotated files are analyzed together). IBM GCMV remains the tool for very large logs.scripts/capture-javacore.sh [-n count] [-i seconds] <namespace> <pod> [container]— human-run only. Taking a thread dump needskubectl exec(jcmd <pid> Dump.java, falling back to SIGQUIT if the image has nojcmd, then reading the javacore file back), which agents are never allowed to do. The script refuses to run without an interactive terminal, shows the kubectl context and target pod, and saves each javacore toruns/javacores/<pod>-<utc>[-n].txt(gitignored)..claude/hooks/block_unsafe_shell.pyalso denies it from the agent's Bash tool. The JVM keeps running; application threads pause briefly while each dump is written. Use-n 3 -i 10to take a series: threads sitting in the same frame in every dump are stuck, not just busy. Run it withbash scripts/capture-javacore.sh …; it also prints the IBM TMDA command for the saved files.scripts/cleanup-javacores.sh [-a] [-y] <namespace> <pod> [container]— human-run only. Javacores stay in the pod until it restarts, so remove them once they've been analyzed. By default it deletes exactly the dumpscapture-javacore.shtook from that pod (recorded underruns/javacores/.in-pod/);-atargets everyjavacore*.txtin the JVM's dump folders. It lists the files and asks once before deleting (-yskips that), and only ever deletesjavacore*.txt. The capture script, the javacore tools' summaries andjvm-analystall remind you to run it.jvm_analyze_javacore(path)analyzes a captured javacore (only files underruns/javacores/): threads by state, largest thread pools (digits collapsed, e.g.Default Executor-thread-#), deadlocks, most-contended lock owners, blocked/parked threads, common stacks, and hot frames among runnable threads. IBM TMDA remains the tool for deeper analysis.jvm_compare_javacores(paths)compares a series of 2–10 javacores from one JVM (a list, or a glob likeruns/javacores/<pod>-<date>*.txt). It reports RUNNABLE threads stuck in the same frame in every dump, threads BLOCKED throughout (with lock owner), idle-I/O and unchanged waiting threads, and per-pool thread creation rates (churn).
All four tools archive their result as evidence and return an
evidence_ref; jvm-analyst is allowed to call them in both harnesses.
This server cannot be deployed
Maintenance
Related MCP Connectors
Provides read access to your GKE and Kubernetes resources.
Read-only MCP access to a documented IT fleet: state, changes, posture. 15 tools.
Read-only MCP access to sessions, funnels, campaigns, errors, live visitors, and anomalies.
Read-only access to a Lumin project's logs, metrics, uptime checks, alerts and infrastructure.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables safe, read-only interaction with Kubernetes clusters, allowing users to list resources and fetch logs without any create/update/delete operations.116Apache 2.0
- FlicenseNot gradedqualityDmaintenanceEnables Kubernetes cluster introspection via MCP tools, such as listing pods, namespaces, nodes, and events.4-
- FlicenseNot gradedqualityBmaintenanceProvides read-only Kubernetes cluster operations via MCP, enabling LLMs to query nodes, pods, logs, events, and watch real-time status for troubleshooting.2-
- FlicenseNot gradedqualityBmaintenanceEnables read-only interaction with a Kubernetes cluster, allowing listing of nodes, pods, deployments, and services, as well as fetching pod logs.-