Dynatrace MCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DYNATRACE_URL | No | Dynatrace SaaS/Grail tenant URL, e.g. https://{env}.apps.dynatrace.com | |
| DT_ENVIRONMENT | No | Alias for DYNATRACE_URL. | |
| DT_PLATFORM_TOKEN | No | Alias for DYNATRACE_API_TOKEN. | |
| DYNATRACE_PROD_URL | No | URL for the named prod connection. | |
| DYNATRACE_API_TOKEN | No | Dynatrace API token only (e.g. dt0s16....). Do not prefix with Api-Token; the server adds that. Pasted Api-Token … / Bearer … values are stripped. | |
| DYNATRACE_STAGE_URL | No | URL for the named stage connection. | |
| NODE_EXTRA_CA_CERTS | No | Corporate CA PEM. | |
| DYNATRACE_PROD_TOKEN | No | Token for the named prod connection. | |
| DYNATRACE_CONNECTIONS | No | JSON map of extra connections: { "name": { "url", "token" } }. | |
| DYNATRACE_STAGE_TOKEN | No | Token for the named stage connection. | |
| DYNATRACE_NO_SSL_VERIFY | No | Skip TLS verify (insecure). | |
| DYNATRACE_DEFAULT_CONNECTION | No | Default connection name. | |
| DYNATRACE_GRAIL_QUERY_BUDGET_GB | No | Session Grail scan cap (default 1000). | 1000 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| prompts | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| dynatrace_healthB | Check Dynatrace connectivity and Grail access for a named connection. |
| dynatrace_list_connectionsB | List configured named connections (default, stage, prod, ...). |
| dynatrace_reset_grail_budgetA | Clear session Grail scanned-bytes budget so queries can continue. |
| dynatrace_execute_dqlB | Run a DQL query against Grail. Use templated tools when possible. Requires valid DQL, not natural language. |
| dynatrace_verify_dqlB | Parse/verify a DQL statement without executing it against Grail storage. |
| dynatrace_query_problemsA | List Davis problems. status: ACTIVE, CLOSED, or ALL. Optional search on description/id/category. Date bounds apply (default last 2h, max 30d). |
| dynatrace_get_problemB | Get a Davis problem by display id (P-123) or event id. |
| dynatrace_list_exceptionsC | Recent ERROR/exception log lines. Use log filters and date bounds. |
| dynatrace_k8s_eventsC | Kubernetes cluster/pod/workload events in a time range. |
| dynatrace_get_entity_idB | Look up Dynatrace entity IDs by type (dt.entity.host, dt.entity.service, ...) and name substring. |
| dynatrace_get_entity_nameB | Resolve a Dynatrace entity ID to its display name. |
| dynatrace_list_clustersA | List Kubernetes clusters known to this Dynatrace tenant (k8s.cluster.name). Call this before filtering pods. |
| dynatrace_list_namespacesC | List namespaces in this tenant, optionally filtered by cluster. |
| dynatrace_list_nodesB | Nodes ranked by container CPU/memory on that node. Pass cluster from dynatrace_list_clusters. |
| dynatrace_list_podsA | Pod table from kube CPU/memory metrics for any cluster/namespace/workload. Quiet pods (no logs) are included. Pass filters from list_clusters / list_namespaces or the user; do not assume names. Unfiltered lists are capped. includeLogs is an optional extra Grail scan. |
| dynatrace_get_podB | Detail for one pod: kube-metrics status (CPU/mem), entity lifetime, ready/restarts if ingested. |
| dynatrace_pod_issuesC | CrashLoop / OOM / not-Ready signals: restarts plus Warning events for a pod. |
| dynatrace_pod_utilizationC | Kubernetes Explorer CPU and Memory quota: usage, throttled, requests, limits, % of request/limit, unused. Filter by cluster/namespace/workload/pod. Same report as dynatrace_pod_quota. |
| dynatrace_pod_quotaB | Kubernetes Explorer quota tables for any namespace: CPU (mcore usage, throttled, requests, limits, %, unused) and Memory (usage, requests, limits, %, unused). Discover cluster/namespace first; do not assume names. |
| dynatrace_pod_eventsC | Kubernetes events for one pod or workload (Explorer Events tab). |
| dynatrace_k8s_issue_eventsC | OOMKilled, CrashLoop/BackOff, Evicted, Failed, Unhealthy, probe failures, Warning events. Scope with cluster/namespace/workload/pod. |
| dynatrace_workload_statusC | Desired vs ready replicas for a Kubernetes workload (covers No pod ready). |
| dynatrace_service_healthC | Request count, failure count, and response time for a service/workload. |
| dynatrace_get_traceC | Fetch spans for a trace_id (from a log line) plus optional linked logs. |
| dynatrace_failed_spansA | Error spans grouped by service/span/pod. Use after service_health shows failures or for latency vs error triage. |
| dynatrace_slow_spansB | Highest-duration spans in the window (latency triage). Follow a row with dynatrace_get_trace. |
| dynatrace_log_summaryA | Count logs by level, namespace, pod, and container. Call this before fetching raw lines. Requires date/time bounds. |
| dynatrace_search_logsB | Exact matching log lines only (not a dump). Filter by cluster/namespace/pod/container, loglevel, contains/regex, trace id, and date bounds. Default limit 50, max 200. |
| dynatrace_logs_surroundingB | Lines around a pivot timestamp for the same pod/container (grep -C in time). before/after default 30s, max 5m. |
| dynatrace_pod_logsB | Last N log lines for a pod (still time-bounded). Prefer log_summary / search_logs for filters. |
| dynatrace_logs_tailB | Polling live tail for one or more pods. Pass since from the previous nextSince. First call defaults to now-30s. Not a websocket; poll from the agent. |
| dynatrace_logs_waitB | Poll until contains/regex appears on the given pods or timeout (default 30s, max 120s). For integration tests. |
| dynatrace_triage_problemA | Investigate a Davis problem: problem record, kube pods/quota, ERROR logs, OOM/CrashLoop events. Pass P-123 or paste alert text. Optional cluster/namespace/workload override extracted hints. |
| dynatrace_triage_scopeA | SRE snapshot without a problem id: active Davis problems, OOM/CrashLoop events, ERROR log counts, CPU/memory quota hotspots, replica mismatch, restarts, service golden signals, failed spans. Pass the user's cluster/namespace/workload (discover first). Default window 30m. |
| dynatrace_security_events_summaryB | Summarize security.events by provider, type, and risk level. Always aggregated. |
| dynatrace_vulnerabilitiesC | Summarized vulnerability findings from Grail security.events (max 100 groups). |
| dynatrace_compliance_findingsB | Summarized compliance findings from Grail (aggregated, max 100). |
| dynatrace_metrics_queryC | Run a DQL timeseries or Environment API v2 metric selector within date bounds. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| dynatrace_triage_alert | Full SRE workflow: discover scope → problems/snapshot → logs → pod quota/events → traces. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 38 tools
Several tools overlap significantly: dynatrace_pod_utilization and dynatrace_pod_quota are explicitly the same report, while dynatrace_k8s_events, dynatrace_pod_events, and dynatrace_k8s_issue_events all surface Kubernetes events with subtly different scopes. Log tools also overlap (list_exceptions vs search_logs, pod_logs vs search_logs), though descriptions provide guidance. Boundaries are mostly documented, but an agent can still misselect among event and log tools.
All tools use the dynatrace_ prefix and snake_case, which is highly predictable. However, many names are noun phrases (pod_issues, workload_status, compliance_findings) rather than the verb_noun pattern, so the convention is consistent in style but not uniformly action-oriented. Minor deviations only.
38 tools is well above the 15-tool threshold for a well-scoped server and falls into the 'too many' range. While the domain is broad, the count increases surface area and contributes to overlap, with several specialized tools that could be consolidated.
The server covers a wide incident-triage surface: DQL, problems, logs, K8s resources, traces, spans, security, and compliance. Gaps remain around metrics discovery (no list_metrics), entity relationships, SLOs, and broader platform resources, but generic execute_dql mitigates many of these. For the apparent observability/troubleshooting purpose, coverage is strong with minor workaroundable gaps.