Cloud DevOps MCP Server
The server provides read-only, evidence-backed Cloud DevOps analysis and risk assessment for infrastructure, identity, Kubernetes, CI/CD, SRE, and supply chain.
Assess Terraform changes and security posture from plan JSON.
Correlate Terraform, IAM, Kubernetes, and GitHub Actions evidence into a deployment-risk bundle.
Build incident runbooks and calculate SLO error budgets.
Review CI/CD pipelines and GitHub Actions workflows.
Review AWS IAM and AWS/Azure/GCP identity policies.
Review Kubernetes deployment readiness and security hardening.
Review software supply chain using SBOMs, CI pinning, image immutability, signing, and provenance.
Optional: enable allowlisted live multi-cloud reads, observability, tracing/SLO, and OpsChugex private intelligence gateways (root-cause, governance, security posture, FinOps, change blast radius).
Provides a tool to assess Terraform or IaC deployment risk based on changed resource classes and release controls.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Cloud DevOps MCP Serverassess deployment risk for terraform changes with IAM and public ingress"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Cloud DevOps MCP Server
Cloud DevOps MCP Server is a Model Context Protocol v2 server by Alex C. Godwin. It provides evidence-backed Cloud DevOps analysis across infrastructure, identity, Kubernetes, CI/CD, SRE and software supply-chain controls.
The v0.14 line adds an opt-in OpsChugex change-intelligence and blast-radius gateway. The public MCP forwards bounded planned-change, topology and readiness evidence to a host-configured private OpsChugex service and returns change risk, impacted topology nodes, representative blast paths, blockers, warnings and evidence gaps. Proprietary dependency propagation, risk weighting and blocker rules are not included in this public MIT repository.
The v0.13 line adds an opt-in OpsChugex advanced FinOps gateway. The public MCP forwards bounded cloud-cost, utilization and Kubernetes allocation evidence to a host-configured private OpsChugex service and returns optimization opportunities, cost anomalies, low/high savings ranges, confidence, cost correlations and evidence gaps. Proprietary savings factors, anomaly thresholds, prioritization and deduplication are not included in this public MIT repository.
The v0.12 line adds an opt-in OpsChugex cloud security posture gateway. The public MCP forwards bounded cloud asset, identity, secret and network-reachability evidence to a host-configured private OpsChugex service and returns security score, risk level, ordered findings, attack paths, evidence gaps and recommendations. Proprietary security rules, severity thresholds, scoring and attack-path correlation are not included in this public MIT repository.
The v0.11 line adds an opt-in OpsChugex policy and governance gateway. The public MCP sends bounded resource evidence and approved exception metadata to a host-configured private OpsChugex service and returns pass, review or block decisions with control findings and evidence gaps. Proprietary profiles, rules, weights and exception-processing logic are not included in this public MIT repository.
The v0.10 line adds an opt-in OpsChugex root-cause intelligence gateway. The public MCP sends bounded incident evidence to a host-configured private OpsChugex service and returns evidence-ranked probable causes, contradictions, limitations and next checks. Proprietary ranking and correlation rules are not included in this public MIT repository.
The v0.9 line adds an opt-in distributed-tracing and SLO-intelligence plane for Grafana Tempo and Jaeger v3 trace reads, service dependency mapping, tracing coverage assessment, multi-window SLO burn-rate analysis and trace/SLO incident correlation. The v0.8 production-observability and operations-intelligence plane remains available for Prometheus, Grafana, CloudWatch Logs Insights, Kubernetes health, GitHub Actions diagnosis, cloud health, FinOps and drift. Live AWS, Azure and GCP access remains bounded and read-only. Cloud mutation remains intentionally unavailable.
Table of contents
Related MCP server: safe-runbook-mcp
Why this exists
AI assistants are more useful in engineering work when they can call focused tools with clear inputs and consistent outputs. This server provides a Cloud DevOps tool layer for:
Cross-domain release-risk correlation across infrastructure, identity, runtime and delivery.
Infrastructure-as-code deployment risk analysis.
Production incident runbook generation.
CI/CD delivery readiness review.
SLO error budget calculations.
AWS IAM least-privilege review.
AWS, Azure and GCP identity policy packs.
Terraform destructive-change, public exposure and encryption security analysis.
Kubernetes workload production readiness and security-policy analysis.
GitHub Actions workflow security and deployment review.
CycloneDX/SPDX SBOM quality and software supply-chain correlation.
Optional authenticated Streamable HTTP serving for self-hosted remote access.
Optional allowlisted live AWS/Azure/GCP inventory, observability, FinOps and drift signals.
Optional Grafana Tempo and Jaeger v3 trace reads, service dependency maps and SLO burn-rate intelligence.
Optional OpsChugex private root-cause intelligence across metrics, logs, traces, Kubernetes, cloud, Terraform and CI/CD evidence.
Optional OpsChugex private policy and governance intelligence for development, staging, production and regulated profiles.
Optional OpsChugex private cloud security posture intelligence for exposure, identity, secrets and attack-path analysis.
Optional OpsChugex private advanced FinOps intelligence for rightsizing, anomaly detection, Kubernetes/cloud cost correlation and savings-range analysis.
Optional OpsChugex private change intelligence for pre-change dependency propagation, blast-radius mapping, blockers and rollback/readiness analysis.
Tools
Tool | Purpose |
| Correlates Terraform, IAM, Kubernetes and GitHub Actions evidence into one deployment-risk assessment with cross-domain change paths. |
| Scores Terraform/IaC risk and can derive evidence from raw Terraform plan JSON. |
| Produces a practical incident response runbook for a service, symptom, environment and severity. |
| Reviews CI/CD maturity while separating failed controls from unknown evidence. |
| Calculates downtime and request-failure budgets with consistency validation. |
| Parses IAM policy JSON and detects wildcard scope and privilege-escalation paths. |
| Parses Kubernetes YAML for probes, resources, disruption protection, image and exposure risks. |
| Parses workflow YAML for triggers, immutable action pins, permissions, caching and concurrency. |
| Applies AWS IAM, Azure RBAC or GCP IAM policy packs to raw policy JSON. |
| Reviews Terraform plan JSON for destructive changes, public exposure, encryption, deletion protection and wildcard IAM. |
| Reviews privileged mode, host access, service accounts, capabilities, seccomp, root filesystems and NetworkPolicy. |
| Correlates CycloneDX/SPDX SBOM quality with CI action pinning, image immutability, signatures and provenance. |
Optional live multi-cloud reads
When explicitly enabled, six additional tools provide allowlisted AWS/Azure/GCP identity verification, bounded inventory, managed Kubernetes discovery, observability configuration summaries, FinOps waste signals and drift reporting. No cloud mutation commands are exposed. AWS general inventory is sourced from the Resource Groups Tagging API, so untagged AWS resources may not appear in that inventory or AWS drift comparison.
Optional production observability intelligence
When explicitly enabled, twelve additional tools provide bounded Prometheus queries, Grafana alert summaries, CloudWatch Logs Insights queries, Kubernetes pod-health summaries, GitHub Actions failure diagnosis, cross-signal incident correlation, cloud-health assessment, deployment/incident correlation, observability coverage assessment, FinOps correlation, cross-runtime drift analysis and operations briefs. Endpoints, cloud scopes, log groups, cluster contexts, namespaces and repositories are allowlisted. The layer is read-only and exposes no alert mutation, deployment mutation or arbitrary shell execution.
Optional distributed tracing and SLO intelligence
When explicitly enabled, six additional v0.9 tools provide bounded Tempo/Jaeger trace search and retrieval, service dependency mapping, tracing coverage assessment, multi-window SLO burn-rate analysis and trace/SLO incident correlation. Remote tracing endpoints must be allowlisted and use HTTPS unless loopback. Backend credentials stay in host environment variables. Trace search windows and result sizes are bounded, and the plane exposes no trace ingestion, sampling mutation or telemetry deletion.
Optional OpsChugex root-cause intelligence
When explicitly enabled, v0.10 exposes diagnose_root_cause. The tool accepts bounded evidence from metrics, logs, traces, Kubernetes, cloud, Terraform and CI/CD, then calls a host-configured private OpsChugex service. The public MCP contains no proprietary ranking rules, accepts no service URL or credential as tool input, and performs no remediation. Evidence scores represent evidence strength rather than statistical probability.
Optional OpsChugex policy and governance intelligence
When explicitly enabled, v0.11 exposes assess_governance_policy. The tool accepts bounded factual resource evidence plus optional owner-attributed, time-bounded exceptions and forwards them to the private OpsChugex policy engine. It returns pass, review, or block with control findings and evidence gaps.
The public MCP contains no proprietary governance profiles, policy rules, scoring weights, exception evaluation logic or enforcement capability. The service URL and token remain host-side and cannot be supplied as MCP arguments.
Optional OpsChugex cloud security posture intelligence
When explicitly enabled, v0.12 exposes assess_cloud_security_posture. The tool accepts bounded asset, identity, secret and network-reachability evidence and forwards it to the private OpsChugex security engine. It returns a security score, risk level, ordered findings, correlated attack paths, evidence gaps and recommended next actions.
The public MCP contains no proprietary security detection thresholds, severity rules, scoring logic or attack-path algorithm. It cannot rotate credentials, change IAM, modify network controls, alter encryption settings or remediate infrastructure.
Optional OpsChugex advanced FinOps intelligence
When explicitly enabled, v0.13 exposes analyze_advanced_finops. The tool accepts bounded cloud-cost, utilization and Kubernetes allocation evidence and forwards it to the private OpsChugex FinOps engine. It returns optimization opportunities, anomalies, savings ranges, confidence, cost correlations and evidence gaps.
The public MCP contains no proprietary savings factors, anomaly thresholds, prioritization rules, confidence algorithm or portfolio deduplication logic. It does not resize resources, terminate workloads, purchase commitments, change Kubernetes requests or perform billing actions.
Optional OpsChugex change intelligence and blast-radius analysis
When explicitly enabled, v0.14 exposes analyze_change_blast_radius. The tool accepts bounded planned-change items, topology nodes/edges and readiness evidence, then forwards them to the private OpsChugex change-intelligence engine. It returns risk level/score, impacted nodes, representative blast paths, blockers, warnings, evidence gaps and recommendations.
The public MCP contains no proprietary graph-propagation algorithm, change-risk weights, production-blocker rules or approval logic. The tool is read-only and cannot apply Terraform, mutate Kubernetes, merge pull requests, deploy workloads, approve changes or execute rollback.
Optional infrastructure operations
When explicitly enabled, six additional tools provide Terraform format/validation/plan summaries and Kubernetes read-only runtime inspection. These operations use repository, context, namespace and resource allowlists. Full Terraform plan JSON, Kubernetes Secrets, arbitrary shell execution, Terraform apply and Kubernetes mutation are deliberately excluded.
Architecture
flowchart TD
LocalClient["Local MCP client"] --> Stdio["stdio"]
RemoteClient["Remote MCP client"] --> HTTPS["HTTPS reverse proxy / gateway"]
HTTPS --> AuthHTTP["Bearer-authenticated Streamable HTTP"]
Stdio --> Server["Cloud DevOps MCP server"]
AuthHTTP --> Server
Server --> DomainTools["Domain + policy-pack analyzers"]
DomainTools --> Correlator["Cross-domain and supply-chain correlation"]
DomainTools --> Output["Structured guidance"]
Correlator --> OutputQuickstart
Run the published MCP server directly from npm:
npx -y cloud-devops-mcp-server@0.14.0On Windows PowerShell systems where script execution policy blocks npx.ps1, use:
npx.cmd -y cloud-devops-mcp-server@0.14.0Install from npm
Install the CLI globally if you prefer a persistent local command:
npm install -g cloud-devops-mcp-server@0.14.0
cloud-devops-mcp-serverThe package is published on npm as cloud-devops-mcp-server and registered in the official MCP Registry as io.github.alexcgodwin/cloud-devops-mcp-server.
MCP clients
Cloud DevOps MCP Server supports local stdio clients and MCP clients capable of connecting to Streamable HTTP endpoints. Common local clients include:
Cursor
Claude Desktop
VS Code with MCP support
Claude Code
Other clients that follow the Model Context Protocol stdio transport
Use stdio for normal local operation. For self-hosted remote access, start the optional authenticated Streamable HTTP endpoint and place non-local deployments behind an HTTPS reverse proxy or gateway.
Configuration
For MCP clients that support local stdio servers, the recommended public configuration is:
{
"mcpServers": {
"cloud-devops": {
"command": "npx",
"args": ["-y", "cloud-devops-mcp-server@0.14.0"]
}
}
}Windows clients can use npx.cmd if npx resolves through a blocked PowerShell wrapper:
{
"mcpServers": {
"cloud-devops": {
"command": "npx.cmd",
"args": ["-y", "cloud-devops-mcp-server@0.14.0"]
}
}
}See docs/configuration.md for npm, global-install, source-development and authenticated Streamable HTTP configuration options.
Authenticated Streamable HTTP
Local loopback example:
$env:CLOUD_DEVOPS_MCP_BEARER_TOKEN="<random secret at least 32 characters>"
npm run start:httpThe MCP endpoint is http://127.0.0.1:3000/mcp and requires Authorization: Bearer <token>. A non-local bind additionally requires CLOUD_DEVOPS_MCP_ALLOWED_HOSTS and an HTTPS CLOUD_DEVOPS_MCP_PUBLIC_BASE_URL so remote traffic is expected to terminate TLS at a reverse proxy or gateway.
Public release verification
The v0.14.0 release candidate passes 116 automated tests, with 85.80% statement, 72.18% branch, 85.89% function and 89.11% line coverage. The production dependency audit reports zero vulnerabilities. Public clean-install and MCP Registry acceptance are recorded after publication.
See docs/public-acceptance.md for the verification record.
Example tool input
{
"changedResources": ["network", "iam", "kubernetes"],
"includesIamChanges": true,
"includesPublicIngress": true,
"modifiesStatefulResources": false,
"hasRollbackPlan": true,
"hasPeerReview": true,
"hasTerraformPlan": true
}Example output shape:
{
"riskScore": 78,
"riskLevel": "critical",
"changedResources": ["network", "iam", "kubernetes"],
"recommendedReleasePath": "Change-advisory review, maintenance window and staged execution are recommended."
}Demo outputs
See docs/demo.md for practical sample inputs and outputs across the toolset.
Docker
Build and run the server in a container:
docker build -t cloud-devops-mcp-server .
docker run --rm -i cloud-devops-mcp-serverDevelopment
npm run dev
npm run build
npm test
npm run checkThe core decision logic lives in src/logic.ts and the MCP tool registration lives in src/index.ts.
More project notes are available in DEVELOPMENT.md, RELEASE.md and docs/architecture.md.
Security model
Stdio remains the default and requires no secrets.
Optional Streamable HTTP requires a bearer token of at least 32 characters.
Non-local HTTP binds require an explicit Host allowlist and an HTTPS public base URL for reverse-proxy/gateway termination.
Host and Origin validation are enabled through the official MCP Fastify adapter.
The default analysis tools do not require cloud credentials or call cloud APIs.
Optional Terraform/Kubernetes operations may use locally configured provider or cluster credentials after explicit enablement and allowlisting.
Optional live cloud reads use existing AWS CLI, Azure CLI or gcloud authentication and require explicit account/subscription/project allowlists.
Optional production-observability reads require explicit endpoint/resource allowlists and host-managed credentials; returned logs and diagnostics are bounded and redacted.
Optional v0.10 root-cause intelligence uses a host-configured HTTPS endpoint and host-side token; neither is accepted as a tool argument, and the proprietary ranking engine remains outside the public repository.
Optional v0.11 governance intelligence uses a separately gated host-configured HTTPS endpoint and the same host-side OpsChugex token; proprietary policy rules and enforcement remain outside the public repository.
Optional v0.12 cloud security posture intelligence uses its own fail-closed gate and host-configured HTTPS endpoint; proprietary security scoring and attack-path correlation remain in the private OpsChugex core.
Optional v0.13 advanced FinOps intelligence uses its own fail-closed gate and host-configured HTTPS endpoint; proprietary savings, anomaly, prioritization and deduplication logic remain in the private OpsChugex core.
Optional v0.14 change intelligence uses its own fail-closed gate and host-configured HTTPS endpoint; proprietary dependency propagation, risk weighting and blocker logic remain in the private OpsChugex core.
No cloud mutation tool is exposed.
Analysis remains read-only by default. Controlled execution appears only when explicitly enabled and allowlisted.
No generic shell tool or force-push capability is exposed.
Direct commit/push on protected branches is blocked, and high-impact GitHub actions require explicit confirmation.
Analysis outputs are advisory. Optional operational tools remain bounded by explicit allowlists and fixed command/API surfaces.
Roadmap
Current: v0.14.0 Change Intelligence & Blast-Radius Analysis - public gateway to private OpsChugex pre-change dependency propagation, impacted-service mapping, blockers and readiness analysis.
v0.15.0 Controlled Remediation Gateway - approval-gated safe fixes for selected cloud, Kubernetes and Terraform operational problems.
v0.16.0 Multi-Account / Multi-Organization Operations - AWS Organizations, Azure tenants/subscriptions and GCP organizations/projects topology.
v0.17.0 Incident Command & Automated Runbooks - incident timelines, evidence bundles, remediation plans, rollback recommendations and post-incident reports.
v0.18.0 Platform Engineering Intelligence - service catalog, ownership, golden paths, environment health and developer-platform checks.
v0.19.0 Enterprise Authentication & Authorization - OAuth/OIDC, RBAC, per-user scopes, stronger hosted-MCP access controls and audit trails.
v1.0.0 Production Stable Release - stable tool contracts, compatibility guarantees, hardened security model, comprehensive documentation and enterprise-ready release standards.
Author
Built by Alex C. Godwin, Cloud DevOps Engineer.
Available Tools
12 toolsassess_cloud_change_bundleAssess Cloud Change BundleARead-onlyIdempotent
Correlate evidence from at least two domains (Terraform, IAM, Kubernetes, GitHub Actions) into one deployment-risk assessment and identify cross-domain change paths. Use domain-specific review tools when only one evidence domain is available. It analyzes caller-supplied artifacts only and does not query providers, clusters, GitHub, or deploy changes.
| Name | Required | Description | Default |
|---|---|---|---|
| terraform | No | Optional Terraform evidence domain. Supply this plus at least one other domain for cross-domain assessment. | |
| changeName | Yes | Name or identifier for the cloud change bundle being assessed. | |
| environment | Yes | Target deployment environment; production increases the consequence of correlated risk. | |
| iamPolicies | No | Optional IAM evidence domain with up to 10 policies; combine with at least one other domain. | |
| githubWorkflows | No | Optional GitHub Actions evidence domain with up to 10 workflows; combine with at least one other domain. | |
| kubernetesWorkloads | No | Optional Kubernetes evidence domain with up to 10 workloads; combine with at least one other domain. |
Output Schema
| Name | Required | Description |
|---|---|---|
| changeName | Yes | |
| changePaths | Yes | |
| environment | Yes | |
| releaseGate | Yes | |
| baseRiskScore | Yes | |
| domainSummary | Yes | |
| uncertainties | Yes | |
| bundleRiskLevel | Yes | |
| bundleRiskScore | Yes | |
| suppliedDomains | Yes | |
| correlatedFindings | Yes | |
| recommendedActions | Yes | |
| assessmentConfidence | Yes | |
| correlationAdjustment | Yes | |
| correlatedFindingCount | Yes | |
| environmentRiskAdjustment | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive operation, so the safety profile is covered. The description adds genuinely new behavioral scope: it analyzes caller-supplied artifacts only and does not query providers, clusters, GitHub, or deploy changes. That closes the open-world question for the caller, though return/report structure is left to the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero waste, front-loaded with the core action and its precondition before the delegation rule and the scope limitation. Each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with nested domains and an output schema, the description supplies the mental model (multi-domain correlation, caller-supplied only) and defers return shape to the output schema. Could be marginally better by noting the two-domain minimum applies to required vs optional inputs, but it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every nested field (hasWildcardActions, usesPinnedActions, hasPodDisruptionBudget, etc.) is already documented. The description names the evidence domains but adds no format, cardinality, or combination semantics beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (correlate) and resource (evidence from Terraform, IAM, Kubernetes, GitHub Actions into one deployment-risk assessment) plus the secondary output (cross-domain change paths). An agent can tell it apart from the single-domain siblings like review_terraform_security or review_iam_policy without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition (at least two evidence domains) and routes the agent to the alternative (domain-specific review tools) when only one domain is available. The when-to-use and when-to-delegate conditions are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_terraform_changeAssess Terraform ChangeARead-onlyIdempotent
Evaluate Terraform/IaC release risk using plan evidence plus change-governance facts such as rollback, peer review and stateful impact. Use this for overall change/release decisions; for security-only inspection of a raw Terraform plan, use review_terraform_security. It analyzes supplied evidence only and never applies a plan, changes infrastructure, or writes Terraform state.
| Name | Required | Description | Default |
|---|---|---|---|
| hasPeerReview | No | Whether another qualified reviewer has reviewed the proposed change. | |
| hasRollbackPlan | No | Whether a documented rollback or recovery path exists for this change. | |
| changedResources | No | Infrastructure resource classes changed by the pull request or deployment. | |
| hasTerraformPlan | No | Whether a Terraform plan artifact was generated and reviewed for this change. | |
| terraformPlanJson | No | Optional raw Terraform plan JSON. When supplied, the server derives resource classes and evidence. | |
| includesIamChanges | No | Whether the change adds, removes or modifies IAM permissions, roles or policies. | |
| includesPublicIngress | No | Whether the change introduces or modifies internet-accessible ingress or public network exposure. | |
| modifiesStatefulResources | No | Whether databases, persistent volumes or other stateful resources are changed or replaced. |
Output Schema
| Name | Required | Description |
|---|---|---|
| evidence | Yes | |
| checklist | Yes | |
| riskLevel | Yes | |
| riskScore | Yes | |
| uncertainties | Yes | |
| changedResources | Yes | |
| assessmentConfidence | Yes | |
| recommendedReleasePath | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and a closed world, so the safety profile is covered. The description still adds meaningful nuance beyond them: it analyzes supplied evidence only and never applies a plan, changes infrastructure, or writes Terraform state, which clarifies it is a pure scoring/assessment operation. It does not describe latency or rate-limit behavior, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero padding, with the core purpose front-loaded before the sibling routing and the behavioral boundary. Each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter assessment tool with a full-coverage schema and an output schema, the description need not explain return values. It covers purpose, sibling disambiguation, and the read-only/no-side-effect boundary, which is everything an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter is self-documented, so the description owes no parameter detail; baseline 3 applies. It names some evidence categories (rollback, peer review, stateful impact) that map to parameters, but adds no syntax or format meaning beyond what the schema already carries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Evaluate) plus resource (Terraform/IaC release risk) and enumerates the evidence classes it consumes (plan evidence, rollback, peer review, stateful impact). It explicitly names the sibling review_terraform_security and the boundary between them, so an agent can differentiate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('overall change/release decisions') and a named alternative with its own selecting condition ('for security-only inspection of a raw Terraform plan, use review_terraform_security'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_incident_runbookBuild Incident RunbookARead-onlyIdempotent
Generate a practical incident-response runbook from a known service symptom, severity, environment and optional signals. Use this to structure response actions and evidence collection; do not use it to fetch or diagnose from live telemetry. It does not execute remediation or make changes to the service.
| Name | Required | Description | Default |
|---|---|---|---|
| service | Yes | Service or application name affected by the incident. | |
| signals | No | Optional known alerts, metrics, logs or traces that should guide triage. | |
| symptom | Yes | Observed user-facing or operational symptom to build the runbook around. | |
| severity | Yes | Incident severity used to scale response urgency and communications. | |
| environment | Yes | Environment where the symptom is occurring. |
Output Schema
| Name | Required | Description |
|---|---|---|
| title | Yes | |
| context | Yes | |
| mitigation | Yes | |
| rcaEvidence | Yes | |
| triageSteps | Yes | |
| communication | Yes | |
| firstFifteenMinutes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, and the description usefully reinforces and extends this by stating it 'does not execute remediation or make changes to the service' and does not pull live telemetry. That clarifies the generation-only nature of the call beyond the raw safety flags, though it says nothing about cost, latency, or output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, no filler, with the core action front-loaded and the exclusions following immediately. Every clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generation tool with a full output schema, 100% parameter coverage, and clear annotations, the description supplies everything needed: what it produces, which inputs matter, and what it explicitly will not do. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with per-field descriptions and two enums, so the schema already carries parameter meaning. The description only echoes the same inputs at a high level and adds no format, constraint, or interaction guidance (e.g. how 'signals' influences triage depth), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and artifact ('Generate a practical incident-response runbook') and scopes the inputs (service symptom, severity, environment, optional signals). It also rules out adjacent behaviors ('fetch or diagnose from live telemetry', 'execute remediation'), so an agent can tell exactly what class of operation this is even though no sibling generates runbooks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('structure response actions and evidence collection') and when-not ('do not use it to fetch or diagnose from live telemetry'), which is more than most definitions. It does not name an alternative sibling tool, so it stops short of a full 5 under the when/when-not/alternatives rubric.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_slo_error_budgetEstimate SLO Error BudgetARead-onlyIdempotent
Calculate remaining SLO downtime budget and, when request counts are supplied, remaining failed-request budget for a fixed period. Use this as a deterministic budget calculator; do not use it to fetch monitoring data or forecast reliability. It performs no external calls and changes no service state.
| Name | Required | Description | Default |
|---|---|---|---|
| periodDays | Yes | Length of the SLO measurement period in calendar days. | |
| requestVolume | No | Optional total request count for calculating a request-failure error budget. | |
| failedRequests | No | Optional failed request count; provide together with requestVolume. | |
| sloTargetPercent | Yes | Target service availability percentage for the measurement period, such as 99.9. | |
| observedDowntimeMinutes | Yes | Downtime already observed during the measurement period, in minutes. |
Output Schema
| Name | Required | Description |
|---|---|---|
| periodDays | Yes | |
| budgetStatus | Yes | |
| failedRequests | No | |
| sloTargetPercent | Yes | |
| allowedFailedRequests | No | |
| allowedDowntimeMinutes | Yes | |
| observedDowntimeMinutes | Yes | |
| remainingFailedRequests | No | |
| remainingDowntimeMinutes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered structurally. The description adds genuinely useful context beyond them: it is deterministic, makes no external calls, and mutates no service state, which tells the agent the answer is a pure function of the inputs. It does not discuss precision/rounding or failure modes of the calculation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the primary computation, then scope limits, then purity guarantees. No filler and no restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a pure arithmetic tool with an output schema, complete parameter documentation and full annotation coverage, the description supplies everything an agent needs: what is computed, the optional branch, and the guarantee that nothing external happens.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining the relationship the schema only hints at: the request-failure budget is produced only when request counts are supplied, tying requestVolume and failedRequests together as a pair. It adds no units or format detail beyond what the schema already gives.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (calculate) and resource (remaining SLO downtime budget / failed-request budget) with the exact condition that activates the second computation. Nothing about it is ambiguous against the review_*/assess_* siblings, which all operate on artifacts rather than performing arithmetic on supplied numbers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the tool as a deterministic calculator and names two things it must NOT be used for (fetching monitoring data, forecasting reliability). That is a clear use/do-not-use boundary, which is exactly what an agent needs to route away from this tool when the user wants live telemetry.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_cicd_pipelineReview CI/CD PipelineARead-onlyIdempotent
Evaluate generic CI/CD production-readiness controls from structured pipeline facts, separating missing evidence from failed controls. Use this for platform-agnostic delivery process review; for raw GitHub Actions YAML, use review_github_actions_workflow. It is analysis-only and does not trigger builds, deployments, approvals, or pipeline changes.
| Name | Required | Description | Default |
|---|---|---|---|
| hasRollback | No | Whether the pipeline has a defined rollback or recovery mechanism. | |
| environments | Yes | Deployment environments handled by the pipeline, for example dev, staging and production. | |
| pipelineName | Yes | Human-readable name of the CI/CD pipeline being reviewed. | |
| hasSecurityScan | No | Whether the pipeline performs automated security scanning before deployment. | |
| hasAutomatedTests | No | Whether automated tests run as a release gate. | |
| deploymentStrategy | Yes | Primary deployment strategy used to release changes. | |
| hasArtifactVersioning | No | Whether build artifacts are immutable and versioned for traceability. | |
| hasManualApprovalForProduction | No | Whether production deployment requires an explicit human approval gate. |
Output Schema
| Name | Required | Description |
|---|---|---|
| findings | Yes | |
| strengths | Yes | |
| pipelineName | Yes | |
| uncertainties | Yes | |
| readinessLevel | Yes | |
| readinessScore | Yes | |
| recommendedGates | Yes | |
| deploymentStrategy | Yes | |
| assessmentConfidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and idempotentHint=true, so the safety profile is covered; the description reinforces this by enumerating the side effects it will not cause. It also adds a genuine behavioral trait beyond the annotations – that output separates missing evidence from failed controls – but does not cover depth of analysis or evidence requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core purpose front-loaded, then routing guidance, then the analysis-only constraint. Every clause earns its place with no restatement of the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations cover the safety and idempotency profile. The description supplies scope, sibling routing, and the no-side-effects guarantee, leaving nothing an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 8 parameters, including the deploymentStrategy enum, so the schema carries full semantic load. The description adds no parameter-level guidance beyond the structured fields, making the baseline 3 correct here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Evaluate generic CI/CD production-readiness controls from structured pipeline facts', and adds the differentiator 'separating missing evidence from failed controls'. It explicitly names the sibling review_github_actions_workflow as the alternative for raw YAML, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('platform-agnostic delivery process review') and a named alternative with its triggering condition ('for raw GitHub Actions YAML, use review_github_actions_workflow'). It also states the exclusions: analysis-only, no builds, deployments, approvals, or pipeline changes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_cloud_identity_policyReview Cloud Identity PolicyARead-onlyIdempotent
Apply deterministic provider-specific identity policy packs to raw AWS IAM, Azure RBAC or GCP IAM policy documents. Use this for cross-cloud identity security analysis; use review_iam_policy when assessing AWS IAM from mixed structured facts or policy JSON. It analyzes supplied policy JSON only and does not call cloud APIs or change permissions.
| Name | Required | Description | Default |
|---|---|---|---|
| provider | Yes | Cloud provider whose identity policy syntax and policy pack should be applied. | |
| policyJson | Yes | Raw provider policy document in JSON form; AWS IAM, Azure role definition/assignment data, or GCP IAM policy. | |
| policyName | Yes | Name or identifier of the identity policy being reviewed. | |
| environment | No | Optional deployment environment used to contextualize policy risk; defaults are handled by the policy pack. |
Output Schema
| Name | Required | Description |
|---|---|---|
| facts | Yes | |
| findings | Yes | |
| provider | Yes | |
| riskLevel | Yes | |
| riskScore | Yes | |
| policyName | Yes | |
| policyPack | Yes | |
| environment | Yes | |
| findingCount | Yes | |
| highFindings | Yes | |
| mediumFindings | Yes | |
| recommendedGate | Yes | |
| criticalFindings | Yes | |
| assessmentConfidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile carries less burden on the description. The description still adds valuable context by stating it analyzes supplied JSON only, does not call cloud APIs, and applies deterministic packs. It could say more about output/return behavior, but the added limits are meaningful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, followed by routing guidance and a scope/limit statement. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the rich annotations cover the safety profile. Combined with the scope, routing, and no-API-call limits in the description, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema does the heavy lifting. The description nonetheless maps the provider options to concrete policy syntaxes (AWS IAM, Azure RBAC, GCP IAM), reinforcing the provider enum's meaning beyond the schema field name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Apply deterministic provider-specific identity policy packs to raw AWS IAM, Azure RBAC or GCP IAM policy documents') and names the sibling it is distinct from. An agent can distinguish this from review_iam_policy without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the routing condition: use this for cross-cloud identity analysis, use review_iam_policy for AWS IAM from mixed structured facts or policy JSON. The alternative and its selecting condition are both named, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_github_actions_workflowReview GitHub Actions WorkflowARead-onlyIdempotent
Evaluate GitHub Actions workflow security and deployment readiness from structured facts or raw workflow YAML, including triggers, action pinning, token permissions, caching, environment protection and concurrency. Use review_cicd_pipeline for generic non-GitHub delivery-process review. It examines supplied evidence only and does not call GitHub or dispatch workflows.
| Name | Required | Description | Default |
|---|---|---|---|
| triggers | No | Workflow trigger events when raw YAML is not supplied, such as push, pull_request or workflow_dispatch. | |
| workflowName | Yes | Human-readable name of the GitHub Actions workflow being reviewed. | |
| workflowYaml | No | Raw GitHub Actions workflow YAML. The server derives triggers, action pinning, token permissions, caching and concurrency. | |
| hasSecretScanning | No | Whether the workflow or surrounding delivery process performs automated secret scanning. | |
| usesPinnedActions | No | Whether third-party actions are pinned to immutable commit SHAs. | |
| deploysToProduction | No | Whether the workflow can deploy directly or indirectly to production. | |
| hasDependencyCaching | No | Whether dependency caching is configured for repeatable and efficient builds. | |
| hasConcurrencyControl | No | Whether concurrency settings prevent overlapping or conflicting workflow runs. | |
| hasEnvironmentProtection | No | Whether protected GitHub environments or equivalent approval controls guard production deployments. | |
| hasLeastPrivilegePermissions | No | Whether GITHUB_TOKEN permissions are explicitly restricted to least privilege. |
Output Schema
| Name | Required | Description |
|---|---|---|
| evidence | Yes | |
| findings | Yes | |
| triggers | Yes | |
| strengths | Yes | |
| workflowName | Yes | |
| uncertainties | Yes | |
| workflowScore | Yes | |
| readinessLevel | Yes | |
| recommendedControls | Yes | |
| assessmentConfidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds useful specifics beyond that: 'examines supplied evidence only and does not call GitHub or dispatch workflows', which concretely reinforces the closed-world, non-mutating behavior. It does not discuss limits or assessment trade-offs, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The scope is front-loaded in the first sentence and the sibling routing plus closed-world caveat follow immediately after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and annotations are rich, so return values and safety need not be explained. The description covers input modes and routing, leaving only minor ambiguity about whether raw YAML and the boolean fact fields can be combined or which takes precedence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented, and the description's facet list largely mirrors the property names rather than adding syntax, precedence, or format detail. Baseline 3 is appropriate when the schema carries the parameter burden; the one mild addition is the implication that YAML input supersedes the boolean fact fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Evaluate') and resource ('GitHub Actions workflow security and deployment readiness') and enumerates the exact facets assessed (triggers, action pinning, token permissions, caching, environment protection, concurrency). It explicitly distinguishes itself from the sibling review_cicd_pipeline, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative tool explicitly ('Use review_cicd_pipeline for generic non-GitHub delivery-process review') and describes the input modes it accepts ('structured facts or raw workflow YAML'). The boundary condition for choosing the sibling versus this tool is stated, not inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_iam_policyReview IAM PolicyARead-onlyIdempotent
Evaluate AWS IAM policy risk from structured facts or raw policy JSON, including wildcard scope, privilege-escalation actions and conditions. Use this for AWS IAM operational risk; for provider-specific AWS/Azure/GCP policy-pack checks, use review_cloud_identity_policy. It analyzes supplied policy data only and does not call AWS or modify IAM.
| Name | Required | Description | Default |
|---|---|---|---|
| actions | No | Explicit allowed IAM actions when raw policy JSON is not supplied. | |
| resources | No | Explicit IAM resource ARNs or patterns when raw policy JSON is not supplied. | |
| policyJson | No | Raw AWS IAM policy JSON. When supplied, the server derives actions, resources, wildcard scope and privilege-escalation evidence. | |
| policyName | Yes | Name of the AWS IAM policy being assessed. | |
| usedByProduction | No | Whether the policy is attached to or used by production identities or workloads. | |
| hasConditionBlocks | No | Whether policy statements include Condition constraints that narrow access. | |
| hasWildcardActions | No | Whether the policy permits wildcard actions such as * or service:* patterns. | |
| hasWildcardResources | No | Whether the policy grants permissions against wildcard resources. | |
| allowsPrivilegeEscalationActions | No | Whether the policy contains actions that can enable privilege escalation. |
Output Schema
| Name | Required | Description |
|---|---|---|
| actions | Yes | |
| evidence | Yes | |
| findings | Yes | |
| resources | Yes | |
| riskLevel | Yes | |
| riskScore | Yes | |
| strengths | Yes | |
| policyName | Yes | |
| uncertainties | Yes | |
| recommendedControls | Yes | |
| assessmentConfidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered; the description reinforces it with 'does not call AWS or modify IAM', which is useful confirmation of offline analysis. It adds nothing about output shape or limits, but the output schema covers returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose and risk scope, then routing guidance, then the side-effect boundary. Every sentence carries distinct information with no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given rich annotations, a 100%-covered 9-parameter schema, and an existing output schema, the description supplies exactly what is missing: purpose, sibling routing, and the offline/non-mutating boundary. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description earns extra by explaining the dual input mode ('structured facts or raw policy JSON') and which risk dimensions each path supports, which helps the agent choose between passing policyJson and the discrete boolean/array fields. It stops short of documenting individual field semantics beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (evaluate/review) and resource (AWS IAM policy), and enumerates what is assessed: wildcard scope, privilege-escalation actions, conditions. It explicitly distinguishes itself from the sibling review_cloud_identity_policy, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use (AWS IAM operational risk) and an explicit alternative with its own condition (provider-specific AWS/Azure/GCP policy-pack checks -> review_cloud_identity_policy). It also sets a boundary on inputs ('analyzes supplied policy data only'), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_kubernetes_deploymentReview Kubernetes DeploymentARead-onlyIdempotent
Evaluate Kubernetes workload production readiness and reliability from structured facts or multi-document YAML, including probes, resources, replicas, disruption protection and exposure. Use this for deployability/readiness; for workload security hardening, use review_kubernetes_security. It examines supplied evidence only and does not connect to a Kubernetes cluster.
| Name | Required | Description | Default |
|---|---|---|---|
| replicas | No | Configured replica count used to assess availability and redundancy. | |
| namespace | No | Kubernetes namespace containing the workload. | |
| runsAsRoot | No | Whether the workload is configured to run containers as root. | |
| manifestYaml | No | Deployment, StatefulSet or DaemonSet YAML, optionally with Service, Ingress and PodDisruptionBudget documents. | |
| workloadName | No | Workload name when raw manifest YAML is not the only source of identity. | |
| usesLatestTag | No | Whether any container image uses the mutable latest tag. | |
| hasLivenessProbe | No | Whether workload containers define liveness probes. | |
| hasReadinessProbe | No | Whether workload containers define readiness probes. | |
| hasResourceLimits | No | Whether workload containers define CPU or memory limits. | |
| hasResourceRequests | No | Whether workload containers define CPU or memory requests. | |
| exposesPublicService | No | Whether the workload is exposed through a public Service or ingress path. | |
| hasPodDisruptionBudget | No | Whether a PodDisruptionBudget protects workload availability during voluntary disruption. |
Output Schema
| Name | Required | Description |
|---|---|---|
| evidence | Yes | |
| findings | Yes | |
| namespace | Yes | |
| strengths | Yes | |
| workloadName | Yes | |
| uncertainties | Yes | |
| readinessLevel | Yes | |
| readinessScore | Yes | |
| recommendedControls | Yes | |
| assessmentConfidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, non-destructive and closed-world behavior, so the bar is lower. The description adds meaningful context beyond them: it operates on 'supplied evidence only and does not connect to a Kubernetes cluster,' which tells the agent not to expect live cluster state. It stops short of describing return shape, though an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: capability and scope first, sibling routing second, evidence-only constraint third. Zero filler and the scope statement is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 optional parameters, an output schema and full annotation coverage, the description supplies exactly the missing pieces: what it assesses, how it routes against its security sibling, and that it is offline/evidence-based. Nothing an agent needs to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% across 12 parameters, so the schema already carries field-level semantics. The description adds framing the schema does not: the two accepted input modes (structured facts or multi-document YAML) and the fact that the checked dimensions map to the boolean flags (probes, resources, replicas, disruption protection, exposure).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Evaluate') plus resource ('Kubernetes workload') and enumerates the readiness dimensions it covers: probes, resources, replicas, disruption protection and exposure. It also names the sibling it is not (review_kubernetes_security), so an agent can differentiate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'Use this for deployability/readiness; for workload security hardening, use review_kubernetes_security.' It gives both the when-to-use and the when-to-use-something-else with the alternative named, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_kubernetes_securityReview Kubernetes SecurityARead-onlyIdempotent
Apply Kubernetes workload security rules to manifest YAML for privileged mode, host access, Linux capabilities, service accounts, seccomp, filesystem settings and network policy. Use this for security hardening; use review_kubernetes_deployment for reliability and production-readiness checks. It analyzes supplied YAML only and does not connect to a cluster.
| Name | Required | Description | Default |
|---|---|---|---|
| environment | No | Optional target environment used to contextualize security findings. | |
| manifestYaml | Yes | Raw multi-document Kubernetes YAML containing workloads and related policy objects to evaluate for security hardening. |
Output Schema
| Name | Required | Description |
|---|---|---|
| findings | Yes | |
| riskLevel | Yes | |
| riskScore | Yes | |
| policyPack | Yes | |
| environment | Yes | |
| findingCount | Yes | |
| highFindings | Yes | |
| workloadCount | Yes | |
| mediumFindings | Yes | |
| recommendedGate | Yes | |
| criticalFindings | Yes | |
| assessmentConfidence | Yes | |
| networkPolicyPresent | Yes | |
| publicExposureObjects | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false. The description adds genuine non-redundant context by asserting it 'analyzes supplied YAML only and does not connect to a cluster,' which clarifies that no live cluster state is consulted. It stops short of describing output shape or any limits on manifest size, so not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: scope first, routing second, boundary condition last. Every sentence carries distinct information and the most decision-relevant content is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations carry the safety profile. The description supplies the remaining gaps an agent needs: the exact analysis scope, the sibling it competes with, and the fact that it is purely static over supplied YAML.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both manifestYaml and environment are already documented in the schema, and the environment enum values are self-explanatory. The description adds no syntax, format, or interpretation guidance for either parameter, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Apply Kubernetes workload security rules to manifest YAML') and enumerates the concrete rule categories covered (privileged mode, host access, capabilities, service accounts, seccomp, filesystem, network policy). It explicitly names the sibling review_kubernetes_deployment as the contrasting tool, so an agent can disambiguate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit routing rule: 'Use this for security hardening; use review_kubernetes_deployment for reliability and production-readiness checks.' Both the when-to-use and the alternative with its selection condition are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_software_supply_chainReview Software Supply ChainARead-onlyIdempotent
Correlate CycloneDX/SPDX SBOM quality with CI action pinning, Kubernetes image immutability, artifact signing and build provenance to assess software-supply-chain risk. Use this when an SBOM is available and supply-chain evidence needs to be evaluated together. It analyzes supplied artifacts only and does not call registries, CI systems, or clusters.
| Name | Required | Description | Default |
|---|---|---|---|
| sbomJson | Yes | CycloneDX or SPDX SBOM JSON used as the primary software-supply-chain evidence. | |
| environment | No | Optional target environment used to contextualize supply-chain risk. | |
| workflowYaml | No | Optional GitHub Actions workflow YAML used to inspect action pinning and build controls. | |
| hasProvenance | No | Whether verifiable build provenance or attestation is produced for released artifacts. | |
| artifactSigned | No | Whether released artifacts or images are cryptographically signed. | |
| kubernetesManifestYaml | No | Optional Kubernetes manifest YAML used to inspect runtime image immutability. |
Output Schema
| Name | Required | Description |
|---|---|---|
| findings | Yes | |
| riskLevel | Yes | |
| riskScore | Yes | |
| policyPack | Yes | |
| sbomFormat | Yes | |
| environment | Yes | |
| findingCount | Yes | |
| highFindings | Yes | |
| hasProvenance | Yes | |
| artifactSigned | Yes | |
| componentCount | Yes | |
| mediumFindings | Yes | |
| recommendedGate | Yes | |
| sbomSpecVersion | Yes | |
| correlationPaths | Yes | |
| criticalFindings | Yes | |
| metadataCoverage | Yes | |
| assessmentConfidence | Yes | |
| mutableRuntimeImages | Yes | |
| mutableActionReferences | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, and the description reinforces this with a concrete operational boundary: 'analyzes supplied artifacts only and does not call registries, CI systems, or clusters.' That closed-world, offline-analysis disclosure is genuinely useful context an agent cannot infer from annotations alone, though it adds little on error behavior or input limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences: the capability, the trigger condition, and the limitation. Every sentence earns its place with no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and 100% schema coverage, the description need not explain return values, and annotations cover the safety profile. It adequately covers scope, trigger, and boundaries; only the absence of named alternatives for narrower sibling tools keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by mapping evidence categories onto parameters: action pinning to workflowYaml, image immutability to kubernetesManifestYaml, signing to artifactSigned, provenance to hasProvenance. This explains what role each input plays in the correlation rather than just restating field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (correlate/assess) and resource (software-supply-chain risk) and enumerates the exact evidence dimensions it fuses: SBOM quality, CI action pinning, Kubernetes image immutability, artifact signing, and build provenance. This distinguishes it from single-domain siblings like review_cicd_pipeline, review_github_actions_workflow, and review_kubernetes_security.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit precondition ('Use this when an SBOM is available and supply-chain evidence needs to be evaluated together'), which tells the agent this is the cross-domain aggregator. It stops short of naming which sibling tools to prefer for single-domain checks, so it is clear context rather than a full when/when-not routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_terraform_securityReview Terraform SecurityARead-onlyIdempotent
Inspect raw Terraform plan JSON for security-relevant changes such as destructive actions, public exposure, encryption gaps, deletion protection and IAM wildcard risk. Use this for Terraform security posture; use assess_terraform_change for broader release/change risk and governance. It parses the supplied plan only and does not execute Terraform or write state.
| Name | Required | Description | Default |
|---|---|---|---|
| environment | No | Optional target environment used to contextualize the severity of findings. | |
| terraformPlanJson | Yes | Raw Terraform plan JSON, typically produced by terraform show -json, to inspect for security-relevant resource changes. |
Output Schema
| Name | Required | Description |
|---|---|---|
| findings | Yes | |
| riskLevel | Yes | |
| riskScore | Yes | |
| policyPack | Yes | |
| environment | Yes | |
| findingCount | Yes | |
| highFindings | Yes | |
| replacements | Yes | |
| mediumFindings | Yes | |
| recommendedGate | Yes | |
| changedResources | Yes | |
| criticalFindings | Yes | |
| destructiveChanges | Yes | |
| assessmentConfidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered structurally. The description still adds real value by clarifying the operational contract: it 'parses the supplied plan only and does not execute Terraform or write state.' It omits any note on size limits or finding/severity output shape, though the output schema covers returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: capability first, routing rule second, execution contract third. Every clause carries information and the scoping constraint (parses only) is placed where it matters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema, complete parameter descriptions, and rich annotations, the description only needs to add scope, sibling routing, and side-effect clarity – all three are present. Nothing an agent needs to select or invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema, including the environment enum's role in contextualizing severity and the plan JSON's expected origin (terraform show -json). The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb ('Inspect') and resource ('raw Terraform plan JSON') and enumerates the concrete security categories it detects: destructive actions, public exposure, encryption gaps, deletion protection, IAM wildcards. It explicitly distinguishes itself from the closest sibling by naming assess_terraform_change, so an agent can differentiate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the selection rule directly: 'Use this for Terraform security posture; use assess_terraform_change for broader release/change risk and governance.' This gives both the when-to-use condition and the named alternative with its own condition, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.14.0- Changed
assess_cloud_change_bundle45 fields changed- added
Input schema / properties / changeName / descriptionAdded value: +"Name or identifier for the cloud change bundle being assessed." - added
Input schema / properties / environment / descriptionAdded value: +"Target deployment environment; production increases the consequence of correlated risk." - added
Input schema / properties / githubWorkflows / descriptionAdded value: +"Optional GitHub Actions evidence domain with up to 10 workflows; combine with at least one other domain." - added
Input schema / properties / githubWorkflows / items / properties / deploysToProduction / descriptionAdded value: +"Whether the workflow can deploy changes to production." - added
Input schema / properties / githubWorkflows / items / properties / hasConcurrencyControl / descriptionAdded value: +"Whether concurrency settings prevent overlapping or conflicting deployment runs." - added
Input schema / properties / githubWorkflows / items / properties / hasDependencyCaching / descriptionAdded value: +"Whether dependency caching is configured for repeatable efficient builds." - added
Input schema / properties / githubWorkflows / items / properties / hasEnvironmentProtection / descriptionAdded value: +"Whether protected GitHub environments or equivalent approval controls guard deployments." - added
Input schema / properties / githubWorkflows / items / properties / hasLeastPrivilegePermissions / descriptionAdded value: +"Whether GITHUB_TOKEN permissions are explicitly restricted to least privilege." - added
Input schema / properties / githubWorkflows / items / properties / hasSecretScanning / descriptionAdded value: +"Whether the delivery process includes secret-detection controls." - added
Input schema / properties / githubWorkflows / items / properties / triggers / descriptionAdded value: +"Workflow trigger events when raw workflow YAML is not supplied." - added
Input schema / properties / githubWorkflows / items / properties / usesPinnedActions / descriptionAdded value: +"Whether third-party actions are pinned to immutable commit SHAs." - added
Input schema / properties / githubWorkflows / items / properties / workflowName / descriptionAdded value: +"Name of the GitHub Actions workflow represented by this evidence item." - added
Input schema / properties / githubWorkflows / items / properties / workflowYaml / descriptionAdded value: +"Optional raw GitHub Actions workflow YAML used to derive CI/CD evidence." - added
Input schema / properties / iamPolicies / descriptionAdded value: +"Optional IAM evidence domain with up to 10 policies; combine with at least one other domain." - added
Input schema / properties / iamPolicies / items / properties / actions / descriptionAdded value: +"Explicit IAM actions when raw policy JSON is not supplied." - added
Input schema / properties / iamPolicies / items / properties / allowsPrivilegeEscalationActions / descriptionAdded value: +"Whether the policy permits actions commonly associated with privilege escalation." - added
Input schema / properties / iamPolicies / items / properties / hasConditionBlocks / descriptionAdded value: +"Whether policy statements include IAM Condition constraints." - added
Input schema / properties / iamPolicies / items / properties / hasWildcardActions / descriptionAdded value: +"Whether the policy allows wildcard actions such as * or service:* patterns." - added
Input schema / properties / iamPolicies / items / properties / hasWildcardResources / descriptionAdded value: +"Whether the policy grants access to wildcard resources." - added
Input schema / properties / iamPolicies / items / properties / policyJson / descriptionAdded value: +"Optional raw AWS IAM policy JSON used to derive identity-risk evidence." - added
Input schema / properties / iamPolicies / items / properties / policyName / descriptionAdded value: +"Name of the IAM policy represented by this evidence item." - added
Input schema / properties / iamPolicies / items / properties / resources / descriptionAdded value: +"Explicit IAM resource ARNs or resource patterns when raw policy JSON is not supplied." - added
Input schema / properties / iamPolicies / items / properties / usedByProduction / descriptionAdded value: +"Whether the policy is attached to or used by production workloads or identities." - added
Input schema / properties / kubernetesWorkloads / descriptionAdded value: +"Optional Kubernetes evidence domain with up to 10 workloads; combine with at least one other domain." - added
Input schema / properties / kubernetesWorkloads / items / properties / exposesPublicService / descriptionAdded value: +"Whether the workload is exposed through a public Service or ingress path." - added
Input schema / properties / kubernetesWorkloads / items / properties / hasLivenessProbe / descriptionAdded value: +"Whether workload containers define liveness probes." - added
Input schema / properties / kubernetesWorkloads / items / properties / hasPodDisruptionBudget / descriptionAdded value: +"Whether disruption protection is provided by a PodDisruptionBudget." - added
Input schema / properties / kubernetesWorkloads / items / properties / hasReadinessProbe / descriptionAdded value: +"Whether workload containers define readiness probes." - added
Input schema / properties / kubernetesWorkloads / items / properties / hasResourceLimits / descriptionAdded value: +"Whether workload containers define CPU or memory limits." - added
Input schema / properties / kubernetesWorkloads / items / properties / hasResourceRequests / descriptionAdded value: +"Whether workload containers define CPU or memory requests." - added
Input schema / properties / kubernetesWorkloads / items / properties / manifestYaml / descriptionAdded value: +"Optional raw Kubernetes YAML used to derive workload evidence." - added
Input schema / properties / kubernetesWorkloads / items / properties / namespace / descriptionAdded value: +"Kubernetes namespace containing the workload." - added
Input schema / properties / kubernetesWorkloads / items / properties / replicas / descriptionAdded value: +"Configured replica count used for availability assessment." - added
Input schema / properties / kubernetesWorkloads / items / properties / runsAsRoot / descriptionAdded value: +"Whether the workload is configured to run containers as root." - added
Input schema / properties / kubernetesWorkloads / items / properties / usesLatestTag / descriptionAdded value: +"Whether any workload image uses the mutable latest tag." - added
Input schema / properties / kubernetesWorkloads / items / properties / workloadName / descriptionAdded value: +"Kubernetes workload name when raw manifest YAML is not the only identifier." - added
Input schema / properties / terraform / descriptionAdded value: +"Optional Terraform evidence domain. Supply this plus at least one other domain for cross-domain assessment." - added
Input schema / properties / terraform / properties / changedResources / descriptionAdded value: +"Terraform resource classes changed when raw plan JSON is not supplied." - added
Input schema / properties / terraform / properties / hasPeerReview / descriptionAdded value: +"Whether the Terraform change has peer-review evidence." - added
Input schema / properties / terraform / properties / hasRollbackPlan / descriptionAdded value: +"Whether the Terraform change has a documented rollback or recovery plan." - added
Input schema / properties / terraform / properties / hasTerraformPlan / descriptionAdded value: +"Whether a Terraform plan artifact exists and was reviewed." - added
Input schema / properties / terraform / properties / includesIamChanges / descriptionAdded value: +"Whether Terraform changes identity or access-management resources." - added
Input schema / properties / terraform / properties / includesPublicIngress / descriptionAdded value: +"Whether Terraform introduces or changes public ingress." - added
Input schema / properties / terraform / properties / modifiesStatefulResources / descriptionAdded value: +"Whether Terraform changes stateful resources such as databases or persistent storage." - added
Input schema / properties / terraform / properties / terraformPlanJson / descriptionAdded value: +"Optional raw Terraform plan JSON used to derive change evidence."
- Changed
assess_terraform_change6 fields changed- added
Input schema / properties / hasPeerReview / descriptionAdded value: +"Whether another qualified reviewer has reviewed the proposed change." - added
Input schema / properties / hasRollbackPlan / descriptionAdded value: +"Whether a documented rollback or recovery path exists for this change." - added
Input schema / properties / hasTerraformPlan / descriptionAdded value: +"Whether a Terraform plan artifact was generated and reviewed for this change." - added
Input schema / properties / includesIamChanges / descriptionAdded value: +"Whether the change adds, removes or modifies IAM permissions, roles or policies." - added
Input schema / properties / includesPublicIngress / descriptionAdded value: +"Whether the change introduces or modifies internet-accessible ingress or public network exposure." - added
Input schema / properties / modifiesStatefulResources / descriptionAdded value: +"Whether databases, persistent volumes or other stateful resources are changed or replaced."
- Changed
build_incident_runbook5 fields changed- added
Input schema / properties / environment / descriptionAdded value: +"Environment where the symptom is occurring." - added
Input schema / properties / service / descriptionAdded value: +"Service or application name affected by the incident." - added
Input schema / properties / severity / descriptionAdded value: +"Incident severity used to scale response urgency and communications." - added
Input schema / properties / signals / descriptionAdded value: +"Optional known alerts, metrics, logs or traces that should guide triage." - added
Input schema / properties / symptom / descriptionAdded value: +"Observed user-facing or operational symptom to build the runbook around."
- Changed
estimate_slo_error_budget5 fields changed- added
Input schema / properties / failedRequests / descriptionAdded value: +"Optional failed request count; provide together with requestVolume." - added
Input schema / properties / observedDowntimeMinutes / descriptionAdded value: +"Downtime already observed during the measurement period, in minutes." - added
Input schema / properties / periodDays / descriptionAdded value: +"Length of the SLO measurement period in calendar days." - added
Input schema / properties / requestVolume / descriptionAdded value: +"Optional total request count for calculating a request-failure error budget." - added
Input schema / properties / sloTargetPercent / descriptionAdded value: +"Target service availability percentage for the measurement period, such as 99.9."
- Changed
review_cicd_pipeline8 fields changed- added
Input schema / properties / deploymentStrategy / descriptionAdded value: +"Primary deployment strategy used to release changes." - added
Input schema / properties / environments / descriptionAdded value: +"Deployment environments handled by the pipeline, for example dev, staging and production." - added
Input schema / properties / hasArtifactVersioning / descriptionAdded value: +"Whether build artifacts are immutable and versioned for traceability." - added
Input schema / properties / hasAutomatedTests / descriptionAdded value: +"Whether automated tests run as a release gate." - added
Input schema / properties / hasManualApprovalForProduction / descriptionAdded value: +"Whether production deployment requires an explicit human approval gate." - added
Input schema / properties / hasRollback / descriptionAdded value: +"Whether the pipeline has a defined rollback or recovery mechanism." - added
Input schema / properties / hasSecurityScan / descriptionAdded value: +"Whether the pipeline performs automated security scanning before deployment." - added
Input schema / properties / pipelineName / descriptionAdded value: +"Human-readable name of the CI/CD pipeline being reviewed."
- Changed
review_cloud_identity_policy4 fields changed- added
Input schema / properties / environment / descriptionAdded value: +"Optional deployment environment used to contextualize policy risk; defaults are handled by the policy pack." - added
Input schema / properties / policyJson / descriptionAdded value: +"Raw provider policy document in JSON form; AWS IAM, Azure role definition/assignment data, or GCP IAM policy." - added
Input schema / properties / policyName / descriptionAdded value: +"Name or identifier of the identity policy being reviewed." - added
Input schema / properties / provider / descriptionAdded value: +"Cloud provider whose identity policy syntax and policy pack should be applied."
- Changed
review_github_actions_workflow9 fields changed- added
Input schema / properties / deploysToProduction / descriptionAdded value: +"Whether the workflow can deploy directly or indirectly to production." - added
Input schema / properties / hasConcurrencyControl / descriptionAdded value: +"Whether concurrency settings prevent overlapping or conflicting workflow runs." - added
Input schema / properties / hasDependencyCaching / descriptionAdded value: +"Whether dependency caching is configured for repeatable and efficient builds." - added
Input schema / properties / hasEnvironmentProtection / descriptionAdded value: +"Whether protected GitHub environments or equivalent approval controls guard production deployments." - added
Input schema / properties / hasLeastPrivilegePermissions / descriptionAdded value: +"Whether GITHUB_TOKEN permissions are explicitly restricted to least privilege." - added
Input schema / properties / hasSecretScanning / descriptionAdded value: +"Whether the workflow or surrounding delivery process performs automated secret scanning." - added
Input schema / properties / triggers / descriptionAdded value: +"Workflow trigger events when raw YAML is not supplied, such as push, pull_request or workflow_dispatch." - added
Input schema / properties / usesPinnedActions / descriptionAdded value: +"Whether third-party actions are pinned to immutable commit SHAs." - added
Input schema / properties / workflowName / descriptionAdded value: +"Human-readable name of the GitHub Actions workflow being reviewed."
- Changed
review_iam_policy9 fields changed- added
Input schema / properties / actions / descriptionAdded value: +"Explicit allowed IAM actions when raw policy JSON is not supplied." - added
Input schema / properties / allowsPrivilegeEscalationActions / descriptionAdded value: +"Whether the policy contains actions that can enable privilege escalation." - added
Input schema / properties / hasConditionBlocks / descriptionAdded value: +"Whether policy statements include Condition constraints that narrow access." - added
Input schema / properties / hasWildcardActions / descriptionAdded value: +"Whether the policy permits wildcard actions such as * or service:* patterns." - added
Input schema / properties / hasWildcardResources / descriptionAdded value: +"Whether the policy grants permissions against wildcard resources." - changed
Input schema / properties / policyJson / descriptionPrevious value: -"Raw IAM policy JSON. The server derives wildcard and privilege-escalation evidence."New value: +"Raw AWS IAM policy JSON. When supplied, the server derives actions, resources, wildcard scope and privilege-escalation evidence." - added
Input schema / properties / policyName / descriptionAdded value: +"Name of the AWS IAM policy being assessed." - added
Input schema / properties / resources / descriptionAdded value: +"Explicit IAM resource ARNs or patterns when raw policy JSON is not supplied." - added
Input schema / properties / usedByProduction / descriptionAdded value: +"Whether the policy is attached to or used by production identities or workloads."
- Changed
review_kubernetes_deployment12 fields changed- added
Input schema / properties / exposesPublicService / descriptionAdded value: +"Whether the workload is exposed through a public Service or ingress path." - added
Input schema / properties / hasLivenessProbe / descriptionAdded value: +"Whether workload containers define liveness probes." - added
Input schema / properties / hasPodDisruptionBudget / descriptionAdded value: +"Whether a PodDisruptionBudget protects workload availability during voluntary disruption." - added
Input schema / properties / hasReadinessProbe / descriptionAdded value: +"Whether workload containers define readiness probes." - added
Input schema / properties / hasResourceLimits / descriptionAdded value: +"Whether workload containers define CPU or memory limits." - added
Input schema / properties / hasResourceRequests / descriptionAdded value: +"Whether workload containers define CPU or memory requests." - changed
Input schema / properties / manifestYaml / descriptionPrevious value: -"Deployment/StatefulSet/DaemonSet YAML, optionally with Service, Ingress and PodDisruptionBudget documents."New value: +"Deployment, StatefulSet or DaemonSet YAML, optionally with Service, Ingress and PodDisruptionBudget documents." - added
Input schema / properties / namespace / descriptionAdded value: +"Kubernetes namespace containing the workload." - added
Input schema / properties / replicas / descriptionAdded value: +"Configured replica count used to assess availability and redundancy." - added
Input schema / properties / runsAsRoot / descriptionAdded value: +"Whether the workload is configured to run containers as root." - added
Input schema / properties / usesLatestTag / descriptionAdded value: +"Whether any container image uses the mutable latest tag." - added
Input schema / properties / workloadName / descriptionAdded value: +"Workload name when raw manifest YAML is not the only source of identity."
- Changed
review_kubernetes_security2 fields changed- added
Input schema / properties / environment / descriptionAdded value: +"Optional target environment used to contextualize security findings." - added
Input schema / properties / manifestYaml / descriptionAdded value: +"Raw multi-document Kubernetes YAML containing workloads and related policy objects to evaluate for security hardening."
- Changed
review_software_supply_chain6 fields changed- added
Input schema / properties / artifactSigned / descriptionAdded value: +"Whether released artifacts or images are cryptographically signed." - added
Input schema / properties / environment / descriptionAdded value: +"Optional target environment used to contextualize supply-chain risk." - added
Input schema / properties / hasProvenance / descriptionAdded value: +"Whether verifiable build provenance or attestation is produced for released artifacts." - added
Input schema / properties / kubernetesManifestYaml / descriptionAdded value: +"Optional Kubernetes manifest YAML used to inspect runtime image immutability." - added
Input schema / properties / sbomJson / descriptionAdded value: +"CycloneDX or SPDX SBOM JSON used as the primary software-supply-chain evidence." - added
Input schema / properties / workflowYaml / descriptionAdded value: +"Optional GitHub Actions workflow YAML used to inspect action pinning and build controls."
- Changed
review_terraform_security2 fields changed- added
Input schema / properties / environment / descriptionAdded value: +"Optional target environment used to contextualize the severity of findings." - added
Input schema / properties / terraformPlanJson / descriptionAdded value: +"Raw Terraform plan JSON, typically produced by terraform show -json, to inspect for security-relevant resource changes."
12 tool updates
v0.4.0- First observed
assess_cloud_change_bundle - First observed
assess_terraform_change - First observed
build_incident_runbook - First observed
estimate_slo_error_budget - First observed
review_cicd_pipeline - First observed
review_cloud_identity_policy - First observed
review_github_actions_workflow - First observed
review_iam_policy - First observed
review_kubernetes_deployment - First observed
review_kubernetes_security - First observed
review_software_supply_chain - First observed
review_terraform_security
TDQS
Scored across 12 tools
Descriptions explicitly cross-reference paired tools (e.g., review_iam_policy vs review_cloud_identity_policy, review_terraform_security vs assess_terraform_change), helping agents distinguish them. However, the set includes several adjacent 'review_' tools (Kubernetes deployment vs security, GitHub Actions vs generic CI/CD) that could still cause hesitation without careful reading.
All tool names follow a consistent snake_case verb_noun pattern using assess_, build_, estimate_, and review_ prefixes. No naming inconsistencies or mixed conventions are present.
With 12 tools, the set sits comfortably within the ideal 3–15 range. Each tool targets a distinct domain or analysis type, and no tool appears redundant.
The surface covers major DevOps analysis domains: Terraform change and security, cloud change bundles, CI/CD and GitHub Actions, IAM and cloud identity, Kubernetes deployment and security, SLO budgets, incident runbooks, and supply chain. Some gaps exist (e.g., cost optimization, compliance, or actual remediation), but the server explicitly scopes itself as analysis-only, so coverage is strong for that purpose.
Maintenance
Related MCP Connectors
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
Read-only, deterministic AI triage and readiness tools implementing Sophon's published rubrics.
- mcpOAuthcom.vibgrate
Query your team's drift, vulnerability, and upgrade data from any AI assistant. OAuth 2.1, 51 tools.
Fail-closed policy guardrails for AI agents running kubectl, terraform, helm, and argocd.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides Kubernetes deployment intelligence with typed query tools and AI-synthesized risk briefs, enabling users to list workloads, get detailed snapshots, query Prometheus metrics, record deployment history, and generate risk assessments before promoting to production.Apache 2.0
- AlicenseAqualityBmaintenanceEnables AI agents to safely inspect and execute version-controlled operational runbooks with policy checks, dry-run planning, and out-of-band approvals.3MIT
- AlicenseAqualityBmaintenanceEnables analyzing Salesforce deployment logs, validating metadata manifests, assessing permission risks, and generating remediation plans through deterministic, local-only rules.4MIT
- AlicenseAqualityAmaintenanceEnables engineers to estimate LLM costs, check GPU VRAM fit, audit MCP configs for security risks, and analyze network device config changes and compliance—all locally, with no account or telemetry.7MIT