Skip to main content
Glama

Cloud DevOps MCP Server

CI npm version MCP Registry License: MIT MCP

Cloud DevOps MCP Server is a Model Context Protocol v2 server by Alex C. Godwin. It provides evidence-backed Cloud DevOps analysis across infrastructure, identity, Kubernetes, CI/CD, SRE and software supply-chain controls.

The v0.14 line adds an opt-in OpsChugex change-intelligence and blast-radius gateway. The public MCP forwards bounded planned-change, topology and readiness evidence to a host-configured private OpsChugex service and returns change risk, impacted topology nodes, representative blast paths, blockers, warnings and evidence gaps. Proprietary dependency propagation, risk weighting and blocker rules are not included in this public MIT repository.

The v0.13 line adds an opt-in OpsChugex advanced FinOps gateway. The public MCP forwards bounded cloud-cost, utilization and Kubernetes allocation evidence to a host-configured private OpsChugex service and returns optimization opportunities, cost anomalies, low/high savings ranges, confidence, cost correlations and evidence gaps. Proprietary savings factors, anomaly thresholds, prioritization and deduplication are not included in this public MIT repository.

The v0.12 line adds an opt-in OpsChugex cloud security posture gateway. The public MCP forwards bounded cloud asset, identity, secret and network-reachability evidence to a host-configured private OpsChugex service and returns security score, risk level, ordered findings, attack paths, evidence gaps and recommendations. Proprietary security rules, severity thresholds, scoring and attack-path correlation are not included in this public MIT repository.

The v0.11 line adds an opt-in OpsChugex policy and governance gateway. The public MCP sends bounded resource evidence and approved exception metadata to a host-configured private OpsChugex service and returns pass, review or block decisions with control findings and evidence gaps. Proprietary profiles, rules, weights and exception-processing logic are not included in this public MIT repository.

The v0.10 line adds an opt-in OpsChugex root-cause intelligence gateway. The public MCP sends bounded incident evidence to a host-configured private OpsChugex service and returns evidence-ranked probable causes, contradictions, limitations and next checks. Proprietary ranking and correlation rules are not included in this public MIT repository.

The v0.9 line adds an opt-in distributed-tracing and SLO-intelligence plane for Grafana Tempo and Jaeger v3 trace reads, service dependency mapping, tracing coverage assessment, multi-window SLO burn-rate analysis and trace/SLO incident correlation. The v0.8 production-observability and operations-intelligence plane remains available for Prometheus, Grafana, CloudWatch Logs Insights, Kubernetes health, GitHub Actions diagnosis, cloud health, FinOps and drift. Live AWS, Azure and GCP access remains bounded and read-only. Cloud mutation remains intentionally unavailable.

Table of contents

Related MCP server: safe-runbook-mcp

Why this exists

AI assistants are more useful in engineering work when they can call focused tools with clear inputs and consistent outputs. This server provides a Cloud DevOps tool layer for:

  • Cross-domain release-risk correlation across infrastructure, identity, runtime and delivery.

  • Infrastructure-as-code deployment risk analysis.

  • Production incident runbook generation.

  • CI/CD delivery readiness review.

  • SLO error budget calculations.

  • AWS IAM least-privilege review.

  • AWS, Azure and GCP identity policy packs.

  • Terraform destructive-change, public exposure and encryption security analysis.

  • Kubernetes workload production readiness and security-policy analysis.

  • GitHub Actions workflow security and deployment review.

  • CycloneDX/SPDX SBOM quality and software supply-chain correlation.

  • Optional authenticated Streamable HTTP serving for self-hosted remote access.

  • Optional allowlisted live AWS/Azure/GCP inventory, observability, FinOps and drift signals.

  • Optional Grafana Tempo and Jaeger v3 trace reads, service dependency maps and SLO burn-rate intelligence.

  • Optional OpsChugex private root-cause intelligence across metrics, logs, traces, Kubernetes, cloud, Terraform and CI/CD evidence.

  • Optional OpsChugex private policy and governance intelligence for development, staging, production and regulated profiles.

  • Optional OpsChugex private cloud security posture intelligence for exposure, identity, secrets and attack-path analysis.

  • Optional OpsChugex private advanced FinOps intelligence for rightsizing, anomaly detection, Kubernetes/cloud cost correlation and savings-range analysis.

  • Optional OpsChugex private change intelligence for pre-change dependency propagation, blast-radius mapping, blockers and rollback/readiness analysis.

Tools

Tool

Purpose

assess_cloud_change_bundle

Correlates Terraform, IAM, Kubernetes and GitHub Actions evidence into one deployment-risk assessment with cross-domain change paths.

assess_terraform_change

Scores Terraform/IaC risk and can derive evidence from raw Terraform plan JSON.

build_incident_runbook

Produces a practical incident response runbook for a service, symptom, environment and severity.

review_cicd_pipeline

Reviews CI/CD maturity while separating failed controls from unknown evidence.

estimate_slo_error_budget

Calculates downtime and request-failure budgets with consistency validation.

review_iam_policy

Parses IAM policy JSON and detects wildcard scope and privilege-escalation paths.

review_kubernetes_deployment

Parses Kubernetes YAML for probes, resources, disruption protection, image and exposure risks.

review_github_actions_workflow

Parses workflow YAML for triggers, immutable action pins, permissions, caching and concurrency.

review_cloud_identity_policy

Applies AWS IAM, Azure RBAC or GCP IAM policy packs to raw policy JSON.

review_terraform_security

Reviews Terraform plan JSON for destructive changes, public exposure, encryption, deletion protection and wildcard IAM.

review_kubernetes_security

Reviews privileged mode, host access, service accounts, capabilities, seccomp, root filesystems and NetworkPolicy.

review_software_supply_chain

Correlates CycloneDX/SPDX SBOM quality with CI action pinning, image immutability, signatures and provenance.

Optional live multi-cloud reads

When explicitly enabled, six additional tools provide allowlisted AWS/Azure/GCP identity verification, bounded inventory, managed Kubernetes discovery, observability configuration summaries, FinOps waste signals and drift reporting. No cloud mutation commands are exposed. AWS general inventory is sourced from the Resource Groups Tagging API, so untagged AWS resources may not appear in that inventory or AWS drift comparison.

Optional production observability intelligence

When explicitly enabled, twelve additional tools provide bounded Prometheus queries, Grafana alert summaries, CloudWatch Logs Insights queries, Kubernetes pod-health summaries, GitHub Actions failure diagnosis, cross-signal incident correlation, cloud-health assessment, deployment/incident correlation, observability coverage assessment, FinOps correlation, cross-runtime drift analysis and operations briefs. Endpoints, cloud scopes, log groups, cluster contexts, namespaces and repositories are allowlisted. The layer is read-only and exposes no alert mutation, deployment mutation or arbitrary shell execution.

Optional distributed tracing and SLO intelligence

When explicitly enabled, six additional v0.9 tools provide bounded Tempo/Jaeger trace search and retrieval, service dependency mapping, tracing coverage assessment, multi-window SLO burn-rate analysis and trace/SLO incident correlation. Remote tracing endpoints must be allowlisted and use HTTPS unless loopback. Backend credentials stay in host environment variables. Trace search windows and result sizes are bounded, and the plane exposes no trace ingestion, sampling mutation or telemetry deletion.

Optional OpsChugex root-cause intelligence

When explicitly enabled, v0.10 exposes diagnose_root_cause. The tool accepts bounded evidence from metrics, logs, traces, Kubernetes, cloud, Terraform and CI/CD, then calls a host-configured private OpsChugex service. The public MCP contains no proprietary ranking rules, accepts no service URL or credential as tool input, and performs no remediation. Evidence scores represent evidence strength rather than statistical probability.

Optional OpsChugex policy and governance intelligence

When explicitly enabled, v0.11 exposes assess_governance_policy. The tool accepts bounded factual resource evidence plus optional owner-attributed, time-bounded exceptions and forwards them to the private OpsChugex policy engine. It returns pass, review, or block with control findings and evidence gaps.

The public MCP contains no proprietary governance profiles, policy rules, scoring weights, exception evaluation logic or enforcement capability. The service URL and token remain host-side and cannot be supplied as MCP arguments.

Optional OpsChugex cloud security posture intelligence

When explicitly enabled, v0.12 exposes assess_cloud_security_posture. The tool accepts bounded asset, identity, secret and network-reachability evidence and forwards it to the private OpsChugex security engine. It returns a security score, risk level, ordered findings, correlated attack paths, evidence gaps and recommended next actions.

The public MCP contains no proprietary security detection thresholds, severity rules, scoring logic or attack-path algorithm. It cannot rotate credentials, change IAM, modify network controls, alter encryption settings or remediate infrastructure.

Optional OpsChugex advanced FinOps intelligence

When explicitly enabled, v0.13 exposes analyze_advanced_finops. The tool accepts bounded cloud-cost, utilization and Kubernetes allocation evidence and forwards it to the private OpsChugex FinOps engine. It returns optimization opportunities, anomalies, savings ranges, confidence, cost correlations and evidence gaps.

The public MCP contains no proprietary savings factors, anomaly thresholds, prioritization rules, confidence algorithm or portfolio deduplication logic. It does not resize resources, terminate workloads, purchase commitments, change Kubernetes requests or perform billing actions.

Optional OpsChugex change intelligence and blast-radius analysis

When explicitly enabled, v0.14 exposes analyze_change_blast_radius. The tool accepts bounded planned-change items, topology nodes/edges and readiness evidence, then forwards them to the private OpsChugex change-intelligence engine. It returns risk level/score, impacted nodes, representative blast paths, blockers, warnings, evidence gaps and recommendations.

The public MCP contains no proprietary graph-propagation algorithm, change-risk weights, production-blocker rules or approval logic. The tool is read-only and cannot apply Terraform, mutate Kubernetes, merge pull requests, deploy workloads, approve changes or execute rollback.

Optional infrastructure operations

When explicitly enabled, six additional tools provide Terraform format/validation/plan summaries and Kubernetes read-only runtime inspection. These operations use repository, context, namespace and resource allowlists. Full Terraform plan JSON, Kubernetes Secrets, arbitrary shell execution, Terraform apply and Kubernetes mutation are deliberately excluded.

Architecture

flowchart TD
  LocalClient["Local MCP client"] --> Stdio["stdio"]
  RemoteClient["Remote MCP client"] --> HTTPS["HTTPS reverse proxy / gateway"]
  HTTPS --> AuthHTTP["Bearer-authenticated Streamable HTTP"]
  Stdio --> Server["Cloud DevOps MCP server"]
  AuthHTTP --> Server
  Server --> DomainTools["Domain + policy-pack analyzers"]
  DomainTools --> Correlator["Cross-domain and supply-chain correlation"]
  DomainTools --> Output["Structured guidance"]
  Correlator --> Output

Quickstart

Run the published MCP server directly from npm:

npx -y cloud-devops-mcp-server@0.14.0

On Windows PowerShell systems where script execution policy blocks npx.ps1, use:

npx.cmd -y cloud-devops-mcp-server@0.14.0

Install from npm

Install the CLI globally if you prefer a persistent local command:

npm install -g cloud-devops-mcp-server@0.14.0
cloud-devops-mcp-server

The package is published on npm as cloud-devops-mcp-server and registered in the official MCP Registry as io.github.alexcgodwin/cloud-devops-mcp-server.

MCP clients

Cloud DevOps MCP Server supports local stdio clients and MCP clients capable of connecting to Streamable HTTP endpoints. Common local clients include:

  • Cursor

  • Claude Desktop

  • VS Code with MCP support

  • Claude Code

  • Other clients that follow the Model Context Protocol stdio transport

Use stdio for normal local operation. For self-hosted remote access, start the optional authenticated Streamable HTTP endpoint and place non-local deployments behind an HTTPS reverse proxy or gateway.

Configuration

For MCP clients that support local stdio servers, the recommended public configuration is:

{
  "mcpServers": {
    "cloud-devops": {
      "command": "npx",
      "args": ["-y", "cloud-devops-mcp-server@0.14.0"]
    }
  }
}

Windows clients can use npx.cmd if npx resolves through a blocked PowerShell wrapper:

{
  "mcpServers": {
    "cloud-devops": {
      "command": "npx.cmd",
      "args": ["-y", "cloud-devops-mcp-server@0.14.0"]
    }
  }
}

See docs/configuration.md for npm, global-install, source-development and authenticated Streamable HTTP configuration options.

Authenticated Streamable HTTP

Local loopback example:

$env:CLOUD_DEVOPS_MCP_BEARER_TOKEN="<random secret at least 32 characters>"
npm run start:http

The MCP endpoint is http://127.0.0.1:3000/mcp and requires Authorization: Bearer <token>. A non-local bind additionally requires CLOUD_DEVOPS_MCP_ALLOWED_HOSTS and an HTTPS CLOUD_DEVOPS_MCP_PUBLIC_BASE_URL so remote traffic is expected to terminate TLS at a reverse proxy or gateway.

Public release verification

The v0.14.0 release candidate passes 116 automated tests, with 85.80% statement, 72.18% branch, 85.89% function and 89.11% line coverage. The production dependency audit reports zero vulnerabilities. Public clean-install and MCP Registry acceptance are recorded after publication.

See docs/public-acceptance.md for the verification record.

Example tool input

{
  "changedResources": ["network", "iam", "kubernetes"],
  "includesIamChanges": true,
  "includesPublicIngress": true,
  "modifiesStatefulResources": false,
  "hasRollbackPlan": true,
  "hasPeerReview": true,
  "hasTerraformPlan": true
}

Example output shape:

{
  "riskScore": 78,
  "riskLevel": "critical",
  "changedResources": ["network", "iam", "kubernetes"],
  "recommendedReleasePath": "Change-advisory review, maintenance window and staged execution are recommended."
}

Demo outputs

See docs/demo.md for practical sample inputs and outputs across the toolset.

Docker

Build and run the server in a container:

docker build -t cloud-devops-mcp-server .
docker run --rm -i cloud-devops-mcp-server

Development

npm run dev
npm run build
npm test
npm run check

The core decision logic lives in src/logic.ts and the MCP tool registration lives in src/index.ts.

More project notes are available in DEVELOPMENT.md, RELEASE.md and docs/architecture.md.

Security model

  • Stdio remains the default and requires no secrets.

  • Optional Streamable HTTP requires a bearer token of at least 32 characters.

  • Non-local HTTP binds require an explicit Host allowlist and an HTTPS public base URL for reverse-proxy/gateway termination.

  • Host and Origin validation are enabled through the official MCP Fastify adapter.

  • The default analysis tools do not require cloud credentials or call cloud APIs.

  • Optional Terraform/Kubernetes operations may use locally configured provider or cluster credentials after explicit enablement and allowlisting.

  • Optional live cloud reads use existing AWS CLI, Azure CLI or gcloud authentication and require explicit account/subscription/project allowlists.

  • Optional production-observability reads require explicit endpoint/resource allowlists and host-managed credentials; returned logs and diagnostics are bounded and redacted.

  • Optional v0.10 root-cause intelligence uses a host-configured HTTPS endpoint and host-side token; neither is accepted as a tool argument, and the proprietary ranking engine remains outside the public repository.

  • Optional v0.11 governance intelligence uses a separately gated host-configured HTTPS endpoint and the same host-side OpsChugex token; proprietary policy rules and enforcement remain outside the public repository.

  • Optional v0.12 cloud security posture intelligence uses its own fail-closed gate and host-configured HTTPS endpoint; proprietary security scoring and attack-path correlation remain in the private OpsChugex core.

  • Optional v0.13 advanced FinOps intelligence uses its own fail-closed gate and host-configured HTTPS endpoint; proprietary savings, anomaly, prioritization and deduplication logic remain in the private OpsChugex core.

  • Optional v0.14 change intelligence uses its own fail-closed gate and host-configured HTTPS endpoint; proprietary dependency propagation, risk weighting and blocker logic remain in the private OpsChugex core.

  • No cloud mutation tool is exposed.

  • Analysis remains read-only by default. Controlled execution appears only when explicitly enabled and allowlisted.

  • No generic shell tool or force-push capability is exposed.

  • Direct commit/push on protected branches is blocked, and high-impact GitHub actions require explicit confirmation.

  • Analysis outputs are advisory. Optional operational tools remain bounded by explicit allowlists and fixed command/API surfaces.

Roadmap

  • Current: v0.14.0 Change Intelligence & Blast-Radius Analysis - public gateway to private OpsChugex pre-change dependency propagation, impacted-service mapping, blockers and readiness analysis.

  • v0.15.0 Controlled Remediation Gateway - approval-gated safe fixes for selected cloud, Kubernetes and Terraform operational problems.

  • v0.16.0 Multi-Account / Multi-Organization Operations - AWS Organizations, Azure tenants/subscriptions and GCP organizations/projects topology.

  • v0.17.0 Incident Command & Automated Runbooks - incident timelines, evidence bundles, remediation plans, rollback recommendations and post-incident reports.

  • v0.18.0 Platform Engineering Intelligence - service catalog, ownership, golden paths, environment health and developer-platform checks.

  • v0.19.0 Enterprise Authentication & Authorization - OAuth/OIDC, RBAC, per-user scopes, stronger hosted-MCP access controls and audit trails.

  • v1.0.0 Production Stable Release - stable tool contracts, compatibility guarantees, hardened security model, comprehensive documentation and enterprise-ready release standards.

Author

Built by Alex C. Godwin, Cloud DevOps Engineer.

Available Tools

12 tools
assess_cloud_change_bundleAssess Cloud Change BundleA
Read-onlyIdempotent

Correlate evidence from at least two domains (Terraform, IAM, Kubernetes, GitHub Actions) into one deployment-risk assessment and identify cross-domain change paths. Use domain-specific review tools when only one evidence domain is available. It analyzes caller-supplied artifacts only and does not query providers, clusters, GitHub, or deploy changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
terraformNoOptional Terraform evidence domain. Supply this plus at least one other domain for cross-domain assessment.
changeNameYesName or identifier for the cloud change bundle being assessed.
environmentYesTarget deployment environment; production increases the consequence of correlated risk.
iamPoliciesNoOptional IAM evidence domain with up to 10 policies; combine with at least one other domain.
githubWorkflowsNoOptional GitHub Actions evidence domain with up to 10 workflows; combine with at least one other domain.
kubernetesWorkloadsNoOptional Kubernetes evidence domain with up to 10 workloads; combine with at least one other domain.

Output Schema

ParametersJSON Schema
NameRequiredDescription
changeNameYes
changePathsYes
environmentYes
releaseGateYes
baseRiskScoreYes
domainSummaryYes
uncertaintiesYes
bundleRiskLevelYes
bundleRiskScoreYes
suppliedDomainsYes
correlatedFindingsYes
recommendedActionsYes
assessmentConfidenceYes
correlationAdjustmentYes
correlatedFindingCountYes
environmentRiskAdjustmentYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive operation, so the safety profile is covered. The description adds genuinely new behavioral scope: it analyzes caller-supplied artifacts only and does not query providers, clusters, GitHub, or deploy changes. That closes the open-world question for the caller, though return/report structure is left to the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero waste, front-loaded with the core action and its precondition before the delegation rule and the scope limitation. Each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with nested domains and an output schema, the description supplies the mental model (multi-domain correlation, caller-supplied only) and defers return shape to the output schema. Could be marginally better by noting the two-domain minimum applies to required vs optional inputs, but it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every nested field (hasWildcardActions, usesPinnedActions, hasPodDisruptionBudget, etc.) is already documented. The description names the evidence domains but adds no format, cardinality, or combination semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (correlate) and resource (evidence from Terraform, IAM, Kubernetes, GitHub Actions into one deployment-risk assessment) plus the secondary output (cross-domain change paths). An agent can tell it apart from the single-domain siblings like review_terraform_security or review_iam_policy without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition (at least two evidence domains) and routes the agent to the alternative (domain-specific review tools) when only one domain is available. The when-to-use and when-to-delegate conditions are both spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_terraform_changeAssess Terraform ChangeA
Read-onlyIdempotent

Evaluate Terraform/IaC release risk using plan evidence plus change-governance facts such as rollback, peer review and stateful impact. Use this for overall change/release decisions; for security-only inspection of a raw Terraform plan, use review_terraform_security. It analyzes supplied evidence only and never applies a plan, changes infrastructure, or writes Terraform state.

ParametersJSON Schema
NameRequiredDescriptionDefault
hasPeerReviewNoWhether another qualified reviewer has reviewed the proposed change.
hasRollbackPlanNoWhether a documented rollback or recovery path exists for this change.
changedResourcesNoInfrastructure resource classes changed by the pull request or deployment.
hasTerraformPlanNoWhether a Terraform plan artifact was generated and reviewed for this change.
terraformPlanJsonNoOptional raw Terraform plan JSON. When supplied, the server derives resource classes and evidence.
includesIamChangesNoWhether the change adds, removes or modifies IAM permissions, roles or policies.
includesPublicIngressNoWhether the change introduces or modifies internet-accessible ingress or public network exposure.
modifiesStatefulResourcesNoWhether databases, persistent volumes or other stateful resources are changed or replaced.

Output Schema

ParametersJSON Schema
NameRequiredDescription
evidenceYes
checklistYes
riskLevelYes
riskScoreYes
uncertaintiesYes
changedResourcesYes
assessmentConfidenceYes
recommendedReleasePathYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and a closed world, so the safety profile is covered. The description still adds meaningful nuance beyond them: it analyzes supplied evidence only and never applies a plan, changes infrastructure, or writes Terraform state, which clarifies it is a pure scoring/assessment operation. It does not describe latency or rate-limit behavior, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero padding, with the core purpose front-loaded before the sibling routing and the behavioral boundary. Each sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter assessment tool with a full-coverage schema and an output schema, the description need not explain return values. It covers purpose, sibling disambiguation, and the read-only/no-side-effect boundary, which is everything an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter is self-documented, so the description owes no parameter detail; baseline 3 applies. It names some evidence categories (rollback, peer review, stateful impact) that map to parameters, but adds no syntax or format meaning beyond what the schema already carries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Evaluate) plus resource (Terraform/IaC release risk) and enumerates the evidence classes it consumes (plan evidence, rollback, peer review, stateful impact). It explicitly names the sibling review_terraform_security and the boundary between them, so an agent can differentiate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use condition ('overall change/release decisions') and a named alternative with its own selecting condition ('for security-only inspection of a raw Terraform plan, use review_terraform_security'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_incident_runbookBuild Incident RunbookA
Read-onlyIdempotent

Generate a practical incident-response runbook from a known service symptom, severity, environment and optional signals. Use this to structure response actions and evidence collection; do not use it to fetch or diagnose from live telemetry. It does not execute remediation or make changes to the service.

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceYesService or application name affected by the incident.
signalsNoOptional known alerts, metrics, logs or traces that should guide triage.
symptomYesObserved user-facing or operational symptom to build the runbook around.
severityYesIncident severity used to scale response urgency and communications.
environmentYesEnvironment where the symptom is occurring.

Output Schema

ParametersJSON Schema
NameRequiredDescription
titleYes
contextYes
mitigationYes
rcaEvidenceYes
triageStepsYes
communicationYes
firstFifteenMinutesYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, and the description usefully reinforces and extends this by stating it 'does not execute remediation or make changes to the service' and does not pull live telemetry. That clarifies the generation-only nature of the call beyond the raw safety flags, though it says nothing about cost, latency, or output shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, no filler, with the core action front-loaded and the exclusions following immediately. Every clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a generation tool with a full output schema, 100% parameter coverage, and clear annotations, the description supplies everything needed: what it produces, which inputs matter, and what it explicitly will not do. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with per-field descriptions and two enums, so the schema already carries parameter meaning. The description only echoes the same inputs at a high level and adds no format, constraint, or interaction guidance (e.g. how 'signals' influences triage depth), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and artifact ('Generate a practical incident-response runbook') and scopes the inputs (service symptom, severity, environment, optional signals). It also rules out adjacent behaviors ('fetch or diagnose from live telemetry', 'execute remediation'), so an agent can tell exactly what class of operation this is even though no sibling generates runbooks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('structure response actions and evidence collection') and when-not ('do not use it to fetch or diagnose from live telemetry'), which is more than most definitions. It does not name an alternative sibling tool, so it stops short of a full 5 under the when/when-not/alternatives rubric.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

estimate_slo_error_budgetEstimate SLO Error BudgetA
Read-onlyIdempotent

Calculate remaining SLO downtime budget and, when request counts are supplied, remaining failed-request budget for a fixed period. Use this as a deterministic budget calculator; do not use it to fetch monitoring data or forecast reliability. It performs no external calls and changes no service state.

ParametersJSON Schema
NameRequiredDescriptionDefault
periodDaysYesLength of the SLO measurement period in calendar days.
requestVolumeNoOptional total request count for calculating a request-failure error budget.
failedRequestsNoOptional failed request count; provide together with requestVolume.
sloTargetPercentYesTarget service availability percentage for the measurement period, such as 99.9.
observedDowntimeMinutesYesDowntime already observed during the measurement period, in minutes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
periodDaysYes
budgetStatusYes
failedRequestsNo
sloTargetPercentYes
allowedFailedRequestsNo
allowedDowntimeMinutesYes
observedDowntimeMinutesYes
remainingFailedRequestsNo
remainingDowntimeMinutesYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered structurally. The description adds genuinely useful context beyond them: it is deterministic, makes no external calls, and mutates no service state, which tells the agent the answer is a pure function of the inputs. It does not discuss precision/rounding or failure modes of the calculation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary computation, then scope limits, then purity guarantees. No filler and no restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a pure arithmetic tool with an output schema, complete parameter documentation and full annotation coverage, the description supplies everything an agent needs: what is computed, the optional branch, and the guarantee that nothing external happens.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining the relationship the schema only hints at: the request-failure budget is produced only when request counts are supplied, tying requestVolume and failedRequests together as a pair. It adds no units or format detail beyond what the schema already gives.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (calculate) and resource (remaining SLO downtime budget / failed-request budget) with the exact condition that activates the second computation. Nothing about it is ambiguous against the review_*/assess_* siblings, which all operate on artifacts rather than performing arithmetic on supplied numbers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the tool as a deterministic calculator and names two things it must NOT be used for (fetching monitoring data, forecasting reliability). That is a clear use/do-not-use boundary, which is exactly what an agent needs to route away from this tool when the user wants live telemetry.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_cicd_pipelineReview CI/CD PipelineA
Read-onlyIdempotent

Evaluate generic CI/CD production-readiness controls from structured pipeline facts, separating missing evidence from failed controls. Use this for platform-agnostic delivery process review; for raw GitHub Actions YAML, use review_github_actions_workflow. It is analysis-only and does not trigger builds, deployments, approvals, or pipeline changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
hasRollbackNoWhether the pipeline has a defined rollback or recovery mechanism.
environmentsYesDeployment environments handled by the pipeline, for example dev, staging and production.
pipelineNameYesHuman-readable name of the CI/CD pipeline being reviewed.
hasSecurityScanNoWhether the pipeline performs automated security scanning before deployment.
hasAutomatedTestsNoWhether automated tests run as a release gate.
deploymentStrategyYesPrimary deployment strategy used to release changes.
hasArtifactVersioningNoWhether build artifacts are immutable and versioned for traceability.
hasManualApprovalForProductionNoWhether production deployment requires an explicit human approval gate.

Output Schema

ParametersJSON Schema
NameRequiredDescription
findingsYes
strengthsYes
pipelineNameYes
uncertaintiesYes
readinessLevelYes
readinessScoreYes
recommendedGatesYes
deploymentStrategyYes
assessmentConfidenceYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and idempotentHint=true, so the safety profile is covered; the description reinforces this by enumerating the side effects it will not cause. It also adds a genuine behavioral trait beyond the annotations – that output separates missing evidence from failed controls – but does not cover depth of analysis or evidence requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core purpose front-loaded, then routing guidance, then the analysis-only constraint. Every clause earns its place with no restatement of the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety and idempotency profile. The description supplies scope, sibling routing, and the no-side-effects guarantee, leaving nothing an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 8 parameters, including the deploymentStrategy enum, so the schema carries full semantic load. The description adds no parameter-level guidance beyond the structured fields, making the baseline 3 correct here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Evaluate generic CI/CD production-readiness controls from structured pipeline facts', and adds the differentiator 'separating missing evidence from failed controls'. It explicitly names the sibling review_github_actions_workflow as the alternative for raw YAML, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use ('platform-agnostic delivery process review') and a named alternative with its triggering condition ('for raw GitHub Actions YAML, use review_github_actions_workflow'). It also states the exclusions: analysis-only, no builds, deployments, approvals, or pipeline changes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_cloud_identity_policyReview Cloud Identity PolicyA
Read-onlyIdempotent

Apply deterministic provider-specific identity policy packs to raw AWS IAM, Azure RBAC or GCP IAM policy documents. Use this for cross-cloud identity security analysis; use review_iam_policy when assessing AWS IAM from mixed structured facts or policy JSON. It analyzes supplied policy JSON only and does not call cloud APIs or change permissions.

ParametersJSON Schema
NameRequiredDescriptionDefault
providerYesCloud provider whose identity policy syntax and policy pack should be applied.
policyJsonYesRaw provider policy document in JSON form; AWS IAM, Azure role definition/assignment data, or GCP IAM policy.
policyNameYesName or identifier of the identity policy being reviewed.
environmentNoOptional deployment environment used to contextualize policy risk; defaults are handled by the policy pack.

Output Schema

ParametersJSON Schema
NameRequiredDescription
factsYes
findingsYes
providerYes
riskLevelYes
riskScoreYes
policyNameYes
policyPackYes
environmentYes
findingCountYes
highFindingsYes
mediumFindingsYes
recommendedGateYes
criticalFindingsYes
assessmentConfidenceYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile carries less burden on the description. The description still adds valuable context by stating it analyzes supplied JSON only, does not call cloud APIs, and applies deterministic packs. It could say more about output/return behavior, but the added limits are meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, followed by routing guidance and a scope/limit statement. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the rich annotations cover the safety profile. Combined with the scope, routing, and no-API-call limits in the description, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema does the heavy lifting. The description nonetheless maps the provider options to concrete policy syntaxes (AWS IAM, Azure RBAC, GCP IAM), reinforcing the provider enum's meaning beyond the schema field name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Apply deterministic provider-specific identity policy packs to raw AWS IAM, Azure RBAC or GCP IAM policy documents') and names the sibling it is distinct from. An agent can distinguish this from review_iam_policy without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the routing condition: use this for cross-cloud identity analysis, use review_iam_policy for AWS IAM from mixed structured facts or policy JSON. The alternative and its selecting condition are both named, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_github_actions_workflowReview GitHub Actions WorkflowA
Read-onlyIdempotent

Evaluate GitHub Actions workflow security and deployment readiness from structured facts or raw workflow YAML, including triggers, action pinning, token permissions, caching, environment protection and concurrency. Use review_cicd_pipeline for generic non-GitHub delivery-process review. It examines supplied evidence only and does not call GitHub or dispatch workflows.

ParametersJSON Schema
NameRequiredDescriptionDefault
triggersNoWorkflow trigger events when raw YAML is not supplied, such as push, pull_request or workflow_dispatch.
workflowNameYesHuman-readable name of the GitHub Actions workflow being reviewed.
workflowYamlNoRaw GitHub Actions workflow YAML. The server derives triggers, action pinning, token permissions, caching and concurrency.
hasSecretScanningNoWhether the workflow or surrounding delivery process performs automated secret scanning.
usesPinnedActionsNoWhether third-party actions are pinned to immutable commit SHAs.
deploysToProductionNoWhether the workflow can deploy directly or indirectly to production.
hasDependencyCachingNoWhether dependency caching is configured for repeatable and efficient builds.
hasConcurrencyControlNoWhether concurrency settings prevent overlapping or conflicting workflow runs.
hasEnvironmentProtectionNoWhether protected GitHub environments or equivalent approval controls guard production deployments.
hasLeastPrivilegePermissionsNoWhether GITHUB_TOKEN permissions are explicitly restricted to least privilege.

Output Schema

ParametersJSON Schema
NameRequiredDescription
evidenceYes
findingsYes
triggersYes
strengthsYes
workflowNameYes
uncertaintiesYes
workflowScoreYes
readinessLevelYes
recommendedControlsYes
assessmentConfidenceYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds useful specifics beyond that: 'examines supplied evidence only and does not call GitHub or dispatch workflows', which concretely reinforces the closed-world, non-mutating behavior. It does not discuss limits or assessment trade-offs, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler. The scope is front-loaded in the first sentence and the sibling routing plus closed-world caveat follow immediately after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and annotations are rich, so return values and safety need not be explained. The description covers input modes and routing, leaving only minor ambiguity about whether raw YAML and the boolean fact fields can be combined or which takes precedence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented, and the description's facet list largely mirrors the property names rather than adding syntax, precedence, or format detail. Baseline 3 is appropriate when the schema carries the parameter burden; the one mild addition is the implication that YAML input supersedes the boolean fact fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Evaluate') and resource ('GitHub Actions workflow security and deployment readiness') and enumerates the exact facets assessed (triggers, action pinning, token permissions, caching, environment protection, concurrency). It explicitly distinguishes itself from the sibling review_cicd_pipeline, so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool explicitly ('Use review_cicd_pipeline for generic non-GitHub delivery-process review') and describes the input modes it accepts ('structured facts or raw workflow YAML'). The boundary condition for choosing the sibling versus this tool is stated, not inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_iam_policyReview IAM PolicyA
Read-onlyIdempotent

Evaluate AWS IAM policy risk from structured facts or raw policy JSON, including wildcard scope, privilege-escalation actions and conditions. Use this for AWS IAM operational risk; for provider-specific AWS/Azure/GCP policy-pack checks, use review_cloud_identity_policy. It analyzes supplied policy data only and does not call AWS or modify IAM.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsNoExplicit allowed IAM actions when raw policy JSON is not supplied.
resourcesNoExplicit IAM resource ARNs or patterns when raw policy JSON is not supplied.
policyJsonNoRaw AWS IAM policy JSON. When supplied, the server derives actions, resources, wildcard scope and privilege-escalation evidence.
policyNameYesName of the AWS IAM policy being assessed.
usedByProductionNoWhether the policy is attached to or used by production identities or workloads.
hasConditionBlocksNoWhether policy statements include Condition constraints that narrow access.
hasWildcardActionsNoWhether the policy permits wildcard actions such as * or service:* patterns.
hasWildcardResourcesNoWhether the policy grants permissions against wildcard resources.
allowsPrivilegeEscalationActionsNoWhether the policy contains actions that can enable privilege escalation.

Output Schema

ParametersJSON Schema
NameRequiredDescription
actionsYes
evidenceYes
findingsYes
resourcesYes
riskLevelYes
riskScoreYes
strengthsYes
policyNameYes
uncertaintiesYes
recommendedControlsYes
assessmentConfidenceYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered; the description reinforces it with 'does not call AWS or modify IAM', which is useful confirmation of offline analysis. It adds nothing about output shape or limits, but the output schema covers returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with purpose and risk scope, then routing guidance, then the side-effect boundary. Every sentence carries distinct information with no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given rich annotations, a 100%-covered 9-parameter schema, and an existing output schema, the description supplies exactly what is missing: purpose, sibling routing, and the offline/non-mutating boundary. An agent has everything needed to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description earns extra by explaining the dual input mode ('structured facts or raw policy JSON') and which risk dimensions each path supports, which helps the agent choose between passing policyJson and the discrete boolean/array fields. It stops short of documenting individual field semantics beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (evaluate/review) and resource (AWS IAM policy), and enumerates what is assessed: wildcard scope, privilege-escalation actions, conditions. It explicitly distinguishes itself from the sibling review_cloud_identity_policy, so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (AWS IAM operational risk) and an explicit alternative with its own condition (provider-specific AWS/Azure/GCP policy-pack checks -> review_cloud_identity_policy). It also sets a boundary on inputs ('analyzes supplied policy data only'), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_kubernetes_deploymentReview Kubernetes DeploymentA
Read-onlyIdempotent

Evaluate Kubernetes workload production readiness and reliability from structured facts or multi-document YAML, including probes, resources, replicas, disruption protection and exposure. Use this for deployability/readiness; for workload security hardening, use review_kubernetes_security. It examines supplied evidence only and does not connect to a Kubernetes cluster.

ParametersJSON Schema
NameRequiredDescriptionDefault
replicasNoConfigured replica count used to assess availability and redundancy.
namespaceNoKubernetes namespace containing the workload.
runsAsRootNoWhether the workload is configured to run containers as root.
manifestYamlNoDeployment, StatefulSet or DaemonSet YAML, optionally with Service, Ingress and PodDisruptionBudget documents.
workloadNameNoWorkload name when raw manifest YAML is not the only source of identity.
usesLatestTagNoWhether any container image uses the mutable latest tag.
hasLivenessProbeNoWhether workload containers define liveness probes.
hasReadinessProbeNoWhether workload containers define readiness probes.
hasResourceLimitsNoWhether workload containers define CPU or memory limits.
hasResourceRequestsNoWhether workload containers define CPU or memory requests.
exposesPublicServiceNoWhether the workload is exposed through a public Service or ingress path.
hasPodDisruptionBudgetNoWhether a PodDisruptionBudget protects workload availability during voluntary disruption.

Output Schema

ParametersJSON Schema
NameRequiredDescription
evidenceYes
findingsYes
namespaceYes
strengthsYes
workloadNameYes
uncertaintiesYes
readinessLevelYes
readinessScoreYes
recommendedControlsYes
assessmentConfidenceYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, non-destructive and closed-world behavior, so the bar is lower. The description adds meaningful context beyond them: it operates on 'supplied evidence only and does not connect to a Kubernetes cluster,' which tells the agent not to expect live cluster state. It stops short of describing return shape, though an output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: capability and scope first, sibling routing second, evidence-only constraint third. Zero filler and the scope statement is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 12 optional parameters, an output schema and full annotation coverage, the description supplies exactly the missing pieces: what it assesses, how it routes against its security sibling, and that it is offline/evidence-based. Nothing an agent needs to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% across 12 parameters, so the schema already carries field-level semantics. The description adds framing the schema does not: the two accepted input modes (structured facts or multi-document YAML) and the fact that the checked dimensions map to the boolean flags (probes, resources, replicas, disruption protection, exposure).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Evaluate') plus resource ('Kubernetes workload') and enumerates the readiness dimensions it covers: probes, resources, replicas, disruption protection and exposure. It also names the sibling it is not (review_kubernetes_security), so an agent can differentiate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: 'Use this for deployability/readiness; for workload security hardening, use review_kubernetes_security.' It gives both the when-to-use and the when-to-use-something-else with the alternative named, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_kubernetes_securityReview Kubernetes SecurityA
Read-onlyIdempotent

Apply Kubernetes workload security rules to manifest YAML for privileged mode, host access, Linux capabilities, service accounts, seccomp, filesystem settings and network policy. Use this for security hardening; use review_kubernetes_deployment for reliability and production-readiness checks. It analyzes supplied YAML only and does not connect to a cluster.

ParametersJSON Schema
NameRequiredDescriptionDefault
environmentNoOptional target environment used to contextualize security findings.
manifestYamlYesRaw multi-document Kubernetes YAML containing workloads and related policy objects to evaluate for security hardening.

Output Schema

ParametersJSON Schema
NameRequiredDescription
findingsYes
riskLevelYes
riskScoreYes
policyPackYes
environmentYes
findingCountYes
highFindingsYes
workloadCountYes
mediumFindingsYes
recommendedGateYes
criticalFindingsYes
assessmentConfidenceYes
networkPolicyPresentYes
publicExposureObjectsYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false. The description adds genuine non-redundant context by asserting it 'analyzes supplied YAML only and does not connect to a cluster,' which clarifies that no live cluster state is consulted. It stops short of describing output shape or any limits on manifest size, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: scope first, routing second, boundary condition last. Every sentence carries distinct information and the most decision-relevant content is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations carry the safety profile. The description supplies the remaining gaps an agent needs: the exact analysis scope, the sibling it competes with, and the fact that it is purely static over supplied YAML.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both manifestYaml and environment are already documented in the schema, and the environment enum values are self-explanatory. The description adds no syntax, format, or interpretation guidance for either parameter, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Apply Kubernetes workload security rules to manifest YAML') and enumerates the concrete rule categories covered (privileged mode, host access, capabilities, service accounts, seccomp, filesystem, network policy). It explicitly names the sibling review_kubernetes_deployment as the contrasting tool, so an agent can disambiguate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit routing rule: 'Use this for security hardening; use review_kubernetes_deployment for reliability and production-readiness checks.' Both the when-to-use and the alternative with its selection condition are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_software_supply_chainReview Software Supply ChainA
Read-onlyIdempotent

Correlate CycloneDX/SPDX SBOM quality with CI action pinning, Kubernetes image immutability, artifact signing and build provenance to assess software-supply-chain risk. Use this when an SBOM is available and supply-chain evidence needs to be evaluated together. It analyzes supplied artifacts only and does not call registries, CI systems, or clusters.

ParametersJSON Schema
NameRequiredDescriptionDefault
sbomJsonYesCycloneDX or SPDX SBOM JSON used as the primary software-supply-chain evidence.
environmentNoOptional target environment used to contextualize supply-chain risk.
workflowYamlNoOptional GitHub Actions workflow YAML used to inspect action pinning and build controls.
hasProvenanceNoWhether verifiable build provenance or attestation is produced for released artifacts.
artifactSignedNoWhether released artifacts or images are cryptographically signed.
kubernetesManifestYamlNoOptional Kubernetes manifest YAML used to inspect runtime image immutability.

Output Schema

ParametersJSON Schema
NameRequiredDescription
findingsYes
riskLevelYes
riskScoreYes
policyPackYes
sbomFormatYes
environmentYes
findingCountYes
highFindingsYes
hasProvenanceYes
artifactSignedYes
componentCountYes
mediumFindingsYes
recommendedGateYes
sbomSpecVersionYes
correlationPathsYes
criticalFindingsYes
metadataCoverageYes
assessmentConfidenceYes
mutableRuntimeImagesYes
mutableActionReferencesYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, and the description reinforces this with a concrete operational boundary: 'analyzes supplied artifacts only and does not call registries, CI systems, or clusters.' That closed-world, offline-analysis disclosure is genuinely useful context an agent cannot infer from annotations alone, though it adds little on error behavior or input limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences: the capability, the trigger condition, and the limitation. Every sentence earns its place with no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 100% schema coverage, the description need not explain return values, and annotations cover the safety profile. It adequately covers scope, trigger, and boundaries; only the absence of named alternatives for narrower sibling tools keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by mapping evidence categories onto parameters: action pinning to workflowYaml, image immutability to kubernetesManifestYaml, signing to artifactSigned, provenance to hasProvenance. This explains what role each input plays in the correlation rather than just restating field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (correlate/assess) and resource (software-supply-chain risk) and enumerates the exact evidence dimensions it fuses: SBOM quality, CI action pinning, Kubernetes image immutability, artifact signing, and build provenance. This distinguishes it from single-domain siblings like review_cicd_pipeline, review_github_actions_workflow, and review_kubernetes_security.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit precondition ('Use this when an SBOM is available and supply-chain evidence needs to be evaluated together'), which tells the agent this is the cross-domain aggregator. It stops short of naming which sibling tools to prefer for single-domain checks, so it is clear context rather than a full when/when-not routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_terraform_securityReview Terraform SecurityA
Read-onlyIdempotent

Inspect raw Terraform plan JSON for security-relevant changes such as destructive actions, public exposure, encryption gaps, deletion protection and IAM wildcard risk. Use this for Terraform security posture; use assess_terraform_change for broader release/change risk and governance. It parses the supplied plan only and does not execute Terraform or write state.

ParametersJSON Schema
NameRequiredDescriptionDefault
environmentNoOptional target environment used to contextualize the severity of findings.
terraformPlanJsonYesRaw Terraform plan JSON, typically produced by terraform show -json, to inspect for security-relevant resource changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
findingsYes
riskLevelYes
riskScoreYes
policyPackYes
environmentYes
findingCountYes
highFindingsYes
replacementsYes
mediumFindingsYes
recommendedGateYes
changedResourcesYes
criticalFindingsYes
destructiveChangesYes
assessmentConfidenceYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered structurally. The description still adds real value by clarifying the operational contract: it 'parses the supplied plan only and does not execute Terraform or write state.' It omits any note on size limits or finding/severity output shape, though the output schema covers returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: capability first, routing rule second, execution contract third. Every clause carries information and the scoping constraint (parses only) is placed where it matters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full output schema, complete parameter descriptions, and rich annotations, the description only needs to add scope, sibling routing, and side-effect clarity – all three are present. Nothing an agent needs to select or invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema, including the environment enum's role in contextualizing severity and the plan JSON's expected origin (terraform show -json). The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb ('Inspect') and resource ('raw Terraform plan JSON') and enumerates the concrete security categories it detects: destructive actions, public exposure, encryption gaps, deletion protection, IAM wildcards. It explicitly distinguishes itself from the closest sibling by naming assess_terraform_change, so an agent can differentiate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the selection rule directly: 'Use this for Terraform security posture; use assess_terraform_change for broader release/change risk and governance.' This gives both the when-to-use condition and the named alternative with its own condition, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.14.0
    • Changedassess_cloud_change_bundle45 fields changed
      • addedInput schema / properties / changeName / description
        Added value: +"Name or identifier for the cloud change bundle being assessed."
      • addedInput schema / properties / environment / description
        Added value: +"Target deployment environment; production increases the consequence of correlated risk."
      • addedInput schema / properties / githubWorkflows / description
        Added value: +"Optional GitHub Actions evidence domain with up to 10 workflows; combine with at least one other domain."
      • addedInput schema / properties / githubWorkflows / items / properties / deploysToProduction / description
        Added value: +"Whether the workflow can deploy changes to production."
      • addedInput schema / properties / githubWorkflows / items / properties / hasConcurrencyControl / description
        Added value: +"Whether concurrency settings prevent overlapping or conflicting deployment runs."
      • addedInput schema / properties / githubWorkflows / items / properties / hasDependencyCaching / description
        Added value: +"Whether dependency caching is configured for repeatable efficient builds."
      • addedInput schema / properties / githubWorkflows / items / properties / hasEnvironmentProtection / description
        Added value: +"Whether protected GitHub environments or equivalent approval controls guard deployments."
      • addedInput schema / properties / githubWorkflows / items / properties / hasLeastPrivilegePermissions / description
        Added value: +"Whether GITHUB_TOKEN permissions are explicitly restricted to least privilege."
      • addedInput schema / properties / githubWorkflows / items / properties / hasSecretScanning / description
        Added value: +"Whether the delivery process includes secret-detection controls."
      • addedInput schema / properties / githubWorkflows / items / properties / triggers / description
        Added value: +"Workflow trigger events when raw workflow YAML is not supplied."
      • addedInput schema / properties / githubWorkflows / items / properties / usesPinnedActions / description
        Added value: +"Whether third-party actions are pinned to immutable commit SHAs."
      • addedInput schema / properties / githubWorkflows / items / properties / workflowName / description
        Added value: +"Name of the GitHub Actions workflow represented by this evidence item."
      • addedInput schema / properties / githubWorkflows / items / properties / workflowYaml / description
        Added value: +"Optional raw GitHub Actions workflow YAML used to derive CI/CD evidence."
      • addedInput schema / properties / iamPolicies / description
        Added value: +"Optional IAM evidence domain with up to 10 policies; combine with at least one other domain."
      • addedInput schema / properties / iamPolicies / items / properties / actions / description
        Added value: +"Explicit IAM actions when raw policy JSON is not supplied."
      • addedInput schema / properties / iamPolicies / items / properties / allowsPrivilegeEscalationActions / description
        Added value: +"Whether the policy permits actions commonly associated with privilege escalation."
      • addedInput schema / properties / iamPolicies / items / properties / hasConditionBlocks / description
        Added value: +"Whether policy statements include IAM Condition constraints."
      • addedInput schema / properties / iamPolicies / items / properties / hasWildcardActions / description
        Added value: +"Whether the policy allows wildcard actions such as * or service:* patterns."
      • addedInput schema / properties / iamPolicies / items / properties / hasWildcardResources / description
        Added value: +"Whether the policy grants access to wildcard resources."
      • addedInput schema / properties / iamPolicies / items / properties / policyJson / description
        Added value: +"Optional raw AWS IAM policy JSON used to derive identity-risk evidence."
      • addedInput schema / properties / iamPolicies / items / properties / policyName / description
        Added value: +"Name of the IAM policy represented by this evidence item."
      • addedInput schema / properties / iamPolicies / items / properties / resources / description
        Added value: +"Explicit IAM resource ARNs or resource patterns when raw policy JSON is not supplied."
      • addedInput schema / properties / iamPolicies / items / properties / usedByProduction / description
        Added value: +"Whether the policy is attached to or used by production workloads or identities."
      • addedInput schema / properties / kubernetesWorkloads / description
        Added value: +"Optional Kubernetes evidence domain with up to 10 workloads; combine with at least one other domain."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / exposesPublicService / description
        Added value: +"Whether the workload is exposed through a public Service or ingress path."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / hasLivenessProbe / description
        Added value: +"Whether workload containers define liveness probes."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / hasPodDisruptionBudget / description
        Added value: +"Whether disruption protection is provided by a PodDisruptionBudget."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / hasReadinessProbe / description
        Added value: +"Whether workload containers define readiness probes."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / hasResourceLimits / description
        Added value: +"Whether workload containers define CPU or memory limits."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / hasResourceRequests / description
        Added value: +"Whether workload containers define CPU or memory requests."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / manifestYaml / description
        Added value: +"Optional raw Kubernetes YAML used to derive workload evidence."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / namespace / description
        Added value: +"Kubernetes namespace containing the workload."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / replicas / description
        Added value: +"Configured replica count used for availability assessment."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / runsAsRoot / description
        Added value: +"Whether the workload is configured to run containers as root."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / usesLatestTag / description
        Added value: +"Whether any workload image uses the mutable latest tag."
      • addedInput schema / properties / kubernetesWorkloads / items / properties / workloadName / description
        Added value: +"Kubernetes workload name when raw manifest YAML is not the only identifier."
      • addedInput schema / properties / terraform / description
        Added value: +"Optional Terraform evidence domain. Supply this plus at least one other domain for cross-domain assessment."
      • addedInput schema / properties / terraform / properties / changedResources / description
        Added value: +"Terraform resource classes changed when raw plan JSON is not supplied."
      • addedInput schema / properties / terraform / properties / hasPeerReview / description
        Added value: +"Whether the Terraform change has peer-review evidence."
      • addedInput schema / properties / terraform / properties / hasRollbackPlan / description
        Added value: +"Whether the Terraform change has a documented rollback or recovery plan."
      • addedInput schema / properties / terraform / properties / hasTerraformPlan / description
        Added value: +"Whether a Terraform plan artifact exists and was reviewed."
      • addedInput schema / properties / terraform / properties / includesIamChanges / description
        Added value: +"Whether Terraform changes identity or access-management resources."
      • addedInput schema / properties / terraform / properties / includesPublicIngress / description
        Added value: +"Whether Terraform introduces or changes public ingress."
      • addedInput schema / properties / terraform / properties / modifiesStatefulResources / description
        Added value: +"Whether Terraform changes stateful resources such as databases or persistent storage."
      • addedInput schema / properties / terraform / properties / terraformPlanJson / description
        Added value: +"Optional raw Terraform plan JSON used to derive change evidence."
    • Changedassess_terraform_change6 fields changed
      • addedInput schema / properties / hasPeerReview / description
        Added value: +"Whether another qualified reviewer has reviewed the proposed change."
      • addedInput schema / properties / hasRollbackPlan / description
        Added value: +"Whether a documented rollback or recovery path exists for this change."
      • addedInput schema / properties / hasTerraformPlan / description
        Added value: +"Whether a Terraform plan artifact was generated and reviewed for this change."
      • addedInput schema / properties / includesIamChanges / description
        Added value: +"Whether the change adds, removes or modifies IAM permissions, roles or policies."
      • addedInput schema / properties / includesPublicIngress / description
        Added value: +"Whether the change introduces or modifies internet-accessible ingress or public network exposure."
      • addedInput schema / properties / modifiesStatefulResources / description
        Added value: +"Whether databases, persistent volumes or other stateful resources are changed or replaced."
    • Changedbuild_incident_runbook5 fields changed
      • addedInput schema / properties / environment / description
        Added value: +"Environment where the symptom is occurring."
      • addedInput schema / properties / service / description
        Added value: +"Service or application name affected by the incident."
      • addedInput schema / properties / severity / description
        Added value: +"Incident severity used to scale response urgency and communications."
      • addedInput schema / properties / signals / description
        Added value: +"Optional known alerts, metrics, logs or traces that should guide triage."
      • addedInput schema / properties / symptom / description
        Added value: +"Observed user-facing or operational symptom to build the runbook around."
    • Changedestimate_slo_error_budget5 fields changed
      • addedInput schema / properties / failedRequests / description
        Added value: +"Optional failed request count; provide together with requestVolume."
      • addedInput schema / properties / observedDowntimeMinutes / description
        Added value: +"Downtime already observed during the measurement period, in minutes."
      • addedInput schema / properties / periodDays / description
        Added value: +"Length of the SLO measurement period in calendar days."
      • addedInput schema / properties / requestVolume / description
        Added value: +"Optional total request count for calculating a request-failure error budget."
      • addedInput schema / properties / sloTargetPercent / description
        Added value: +"Target service availability percentage for the measurement period, such as 99.9."
    • Changedreview_cicd_pipeline8 fields changed
      • addedInput schema / properties / deploymentStrategy / description
        Added value: +"Primary deployment strategy used to release changes."
      • addedInput schema / properties / environments / description
        Added value: +"Deployment environments handled by the pipeline, for example dev, staging and production."
      • addedInput schema / properties / hasArtifactVersioning / description
        Added value: +"Whether build artifacts are immutable and versioned for traceability."
      • addedInput schema / properties / hasAutomatedTests / description
        Added value: +"Whether automated tests run as a release gate."
      • addedInput schema / properties / hasManualApprovalForProduction / description
        Added value: +"Whether production deployment requires an explicit human approval gate."
      • addedInput schema / properties / hasRollback / description
        Added value: +"Whether the pipeline has a defined rollback or recovery mechanism."
      • addedInput schema / properties / hasSecurityScan / description
        Added value: +"Whether the pipeline performs automated security scanning before deployment."
      • addedInput schema / properties / pipelineName / description
        Added value: +"Human-readable name of the CI/CD pipeline being reviewed."
    • Changedreview_cloud_identity_policy4 fields changed
      • addedInput schema / properties / environment / description
        Added value: +"Optional deployment environment used to contextualize policy risk; defaults are handled by the policy pack."
      • addedInput schema / properties / policyJson / description
        Added value: +"Raw provider policy document in JSON form; AWS IAM, Azure role definition/assignment data, or GCP IAM policy."
      • addedInput schema / properties / policyName / description
        Added value: +"Name or identifier of the identity policy being reviewed."
      • addedInput schema / properties / provider / description
        Added value: +"Cloud provider whose identity policy syntax and policy pack should be applied."
    • Changedreview_github_actions_workflow9 fields changed
      • addedInput schema / properties / deploysToProduction / description
        Added value: +"Whether the workflow can deploy directly or indirectly to production."
      • addedInput schema / properties / hasConcurrencyControl / description
        Added value: +"Whether concurrency settings prevent overlapping or conflicting workflow runs."
      • addedInput schema / properties / hasDependencyCaching / description
        Added value: +"Whether dependency caching is configured for repeatable and efficient builds."
      • addedInput schema / properties / hasEnvironmentProtection / description
        Added value: +"Whether protected GitHub environments or equivalent approval controls guard production deployments."
      • addedInput schema / properties / hasLeastPrivilegePermissions / description
        Added value: +"Whether GITHUB_TOKEN permissions are explicitly restricted to least privilege."
      • addedInput schema / properties / hasSecretScanning / description
        Added value: +"Whether the workflow or surrounding delivery process performs automated secret scanning."
      • addedInput schema / properties / triggers / description
        Added value: +"Workflow trigger events when raw YAML is not supplied, such as push, pull_request or workflow_dispatch."
      • addedInput schema / properties / usesPinnedActions / description
        Added value: +"Whether third-party actions are pinned to immutable commit SHAs."
      • addedInput schema / properties / workflowName / description
        Added value: +"Human-readable name of the GitHub Actions workflow being reviewed."
    • Changedreview_iam_policy9 fields changed
      • addedInput schema / properties / actions / description
        Added value: +"Explicit allowed IAM actions when raw policy JSON is not supplied."
      • addedInput schema / properties / allowsPrivilegeEscalationActions / description
        Added value: +"Whether the policy contains actions that can enable privilege escalation."
      • addedInput schema / properties / hasConditionBlocks / description
        Added value: +"Whether policy statements include Condition constraints that narrow access."
      • addedInput schema / properties / hasWildcardActions / description
        Added value: +"Whether the policy permits wildcard actions such as * or service:* patterns."
      • addedInput schema / properties / hasWildcardResources / description
        Added value: +"Whether the policy grants permissions against wildcard resources."
      • changedInput schema / properties / policyJson / description
        Previous value: -"Raw IAM policy JSON. The server derives wildcard and privilege-escalation evidence."New value: +"Raw AWS IAM policy JSON. When supplied, the server derives actions, resources, wildcard scope and privilege-escalation evidence."
      • addedInput schema / properties / policyName / description
        Added value: +"Name of the AWS IAM policy being assessed."
      • addedInput schema / properties / resources / description
        Added value: +"Explicit IAM resource ARNs or patterns when raw policy JSON is not supplied."
      • addedInput schema / properties / usedByProduction / description
        Added value: +"Whether the policy is attached to or used by production identities or workloads."
    • Changedreview_kubernetes_deployment12 fields changed
      • addedInput schema / properties / exposesPublicService / description
        Added value: +"Whether the workload is exposed through a public Service or ingress path."
      • addedInput schema / properties / hasLivenessProbe / description
        Added value: +"Whether workload containers define liveness probes."
      • addedInput schema / properties / hasPodDisruptionBudget / description
        Added value: +"Whether a PodDisruptionBudget protects workload availability during voluntary disruption."
      • addedInput schema / properties / hasReadinessProbe / description
        Added value: +"Whether workload containers define readiness probes."
      • addedInput schema / properties / hasResourceLimits / description
        Added value: +"Whether workload containers define CPU or memory limits."
      • addedInput schema / properties / hasResourceRequests / description
        Added value: +"Whether workload containers define CPU or memory requests."
      • changedInput schema / properties / manifestYaml / description
        Previous value: -"Deployment/StatefulSet/DaemonSet YAML, optionally with Service, Ingress and PodDisruptionBudget documents."New value: +"Deployment, StatefulSet or DaemonSet YAML, optionally with Service, Ingress and PodDisruptionBudget documents."
      • addedInput schema / properties / namespace / description
        Added value: +"Kubernetes namespace containing the workload."
      • addedInput schema / properties / replicas / description
        Added value: +"Configured replica count used to assess availability and redundancy."
      • addedInput schema / properties / runsAsRoot / description
        Added value: +"Whether the workload is configured to run containers as root."
      • addedInput schema / properties / usesLatestTag / description
        Added value: +"Whether any container image uses the mutable latest tag."
      • addedInput schema / properties / workloadName / description
        Added value: +"Workload name when raw manifest YAML is not the only source of identity."
    • Changedreview_kubernetes_security2 fields changed
      • addedInput schema / properties / environment / description
        Added value: +"Optional target environment used to contextualize security findings."
      • addedInput schema / properties / manifestYaml / description
        Added value: +"Raw multi-document Kubernetes YAML containing workloads and related policy objects to evaluate for security hardening."
    • Changedreview_software_supply_chain6 fields changed
      • addedInput schema / properties / artifactSigned / description
        Added value: +"Whether released artifacts or images are cryptographically signed."
      • addedInput schema / properties / environment / description
        Added value: +"Optional target environment used to contextualize supply-chain risk."
      • addedInput schema / properties / hasProvenance / description
        Added value: +"Whether verifiable build provenance or attestation is produced for released artifacts."
      • addedInput schema / properties / kubernetesManifestYaml / description
        Added value: +"Optional Kubernetes manifest YAML used to inspect runtime image immutability."
      • addedInput schema / properties / sbomJson / description
        Added value: +"CycloneDX or SPDX SBOM JSON used as the primary software-supply-chain evidence."
      • addedInput schema / properties / workflowYaml / description
        Added value: +"Optional GitHub Actions workflow YAML used to inspect action pinning and build controls."
    • Changedreview_terraform_security2 fields changed
      • addedInput schema / properties / environment / description
        Added value: +"Optional target environment used to contextualize the severity of findings."
      • addedInput schema / properties / terraformPlanJson / description
        Added value: +"Raw Terraform plan JSON, typically produced by terraform show -json, to inspect for security-relevant resource changes."
  2. 12 tool updatesv0.4.0
    • First observedassess_cloud_change_bundle
    • First observedassess_terraform_change
    • First observedbuild_incident_runbook
    • First observedestimate_slo_error_budget
    • First observedreview_cicd_pipeline
    • First observedreview_cloud_identity_policy
    • First observedreview_github_actions_workflow
    • First observedreview_iam_policy
    • First observedreview_kubernetes_deployment
    • First observedreview_kubernetes_security
    • First observedreview_software_supply_chain
    • First observedreview_terraform_security

TDQS

A4.4/5.0

Scored across 12 tools

Disambiguation4/5

Descriptions explicitly cross-reference paired tools (e.g., review_iam_policy vs review_cloud_identity_policy, review_terraform_security vs assess_terraform_change), helping agents distinguish them. However, the set includes several adjacent 'review_' tools (Kubernetes deployment vs security, GitHub Actions vs generic CI/CD) that could still cause hesitation without careful reading.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern using assess_, build_, estimate_, and review_ prefixes. No naming inconsistencies or mixed conventions are present.

Tool Count5/5

With 12 tools, the set sits comfortably within the ideal 3–15 range. Each tool targets a distinct domain or analysis type, and no tool appears redundant.

Completeness4/5

The surface covers major DevOps analysis domains: Terraform change and security, cloud change bundles, CI/CD and GitHub Actions, IAM and cloud identity, Kubernetes deployment and security, SLO budgets, incident runbooks, and supply chain. Some gaps exist (e.g., cost optimization, compliance, or actual remediation), but the server explicitly scopes itself as analysis-only, so coverage is strong for that purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides Kubernetes deployment intelligence with typed query tools and AI-synthesized risk briefs, enabling users to list workloads, get detailed snapshots, query Prometheus metrics, record deployment history, and generate risk assessments before promoting to production.
    Apache 2.0