AWS Infrastructure Operations MCP
Provides tools to check nginx service status (active state, sub-state, boot-enabled) and retrieve recent error-related events from nginx log groups.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AWS Infrastructure Operations MCPCheck web01's EC2 health and recent nginx errors."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AWS Infrastructure Operations MCP
This project runs a small, local server that helps Codex investigate one AWS lab host using live evidence. It is useful when you want an assistant to check whether an EC2 instance is healthy, inspect a fixed set of performance metrics, look for recent system or nginx errors, inspect nginx service state, and read a bounded nginx systemd journal without giving the assistant general AWS or shell access.
MCP (Model Context Protocol) is a standard way for an AI client to call well-defined tools provided by another process. Here, Codex starts the server locally and exchanges messages with it over standard input and output. The server validates each request, calls a narrow set of AWS APIs, and returns structured evidence with its data source.
The server is read-only. It has no generic shell tool and no restart, stop,
deploy, configuration, or other remediation tools. It supports only the
allowlisted instance web01 and the service nginx.
What the five tools check
Tool | AWS source | What it returns |
| EC2 | Instance state plus AWS system and instance status checks. |
| CloudWatch metrics | CPU average/maximum, status-check maximums, and total network bytes over a bounded lookback. |
| CloudWatch Logs Insights | A bounded set of recent error-related events from the approved system and nginx log groups. |
| Systems Manager | nginx active state, sub-state, and boot-enabled state from a fixed, restricted SSM document. |
| Systems Manager | A bounded nginx systemd journal from a separate fixed, restricted SSM document. |
Every response identifies its source, including aws-cloudwatch-metrics for
instance metrics. Tests inject fake AWS clients, so the automated test suite
makes no AWS calls.
get_instance_metrics accepts only an approved instance name and a lookback of
5–1440 minutes (60 minutes by default). It uses five-minute periods and the
fixed AWS/EC2 allowlist: CPUUtilization Average and Maximum;
StatusCheckFailed, StatusCheckFailed_Instance, and
StatusCheckFailed_System Maximum; and NetworkIn and NetworkOut Sum. The
caller cannot supply an instance ID, namespace, metric, dimension, statistic,
period, or CloudWatch query. Missing datapoints are returned as null, not
invented zero values.
get_recent_errors accepts a lookback of 5–1440 minutes and a result limit of
1–50. The caller cannot provide log-group names or Logs Insights query text.
The server always queries:
/aws/mcp-lab/web01/system
/aws/mcp-lab/web01/nginxget_service_status invokes only the Terraform-managed
mcp-lab-get-nginx-status document. That document has no parameters and runs a
fixed set of read-only systemctl checks. The caller cannot provide a command,
document name, instance ID, path, or shell argument.
get_service_journal invokes only the Terraform-managed
mcp-lab-get-nginx-journal document. It accepts lookbacks of only 5, 10, 15,
30, 60, or 120 minutes and result limits of only 10, 25, 50, or 100. The
document is Linux-only, has a short timeout, fixes the unit to nginx, and runs a
fixed journalctl operation with no pager and bounded output. Its only
parameters are the enumerated lookback and result limit. The caller cannot
provide a unit, command, journal argument, path, filter, instance ID, or
document name. Returned lines, entry count, and total response size are bounded;
an empty journal is returned as an empty entries array.
Related MCP server: ssh-remote-mcp
Architecture
Codex
│ MCP over local stdio
▼
server.py
│
├─ get_instance_health ── EC2 status APIs
├─ get_instance_metrics ─ CloudWatch GetMetricData
├─ get_recent_errors ──── CloudWatch Logs Insights
├─ get_service_status ─── restricted SSM document ── web01
└─ get_service_journal ── restricted SSM document ── web01Terraform creates the lab instance, log groups, restricted SSM documents,
least-privilege runtime policy and role, and supporting network resources.
Application code lives in aws_infra_ops_mcp/; Terraform lives in
infrastructure/.
Security boundaries
The MCP runtime role can perform only the diagnostic API calls required by the four tools.
Instance and service names are allowlisted in code.
CloudWatch metric and log queries are fixed. Metrics are limited to the documented
AWS/EC2allowlist, while logs are limited to two log groups. Thelogs:StartQueryIAM resources use the required log-group ARN form ending in:*; Terraform normalizes inputs before adding that suffix.The status SSM document accepts no user parameters and cannot be replaced with an arbitrary Run Command document.
The journal SSM document accepts only enumerated numeric bounds. Its nginx unit and
journalctlcommand are fixed, and IAM permitsssm:SendCommandonly against the exact custom status or journal document and the exact lab instance.The instance security group has no inbound rules. Administration uses Systems Manager rather than SSH.
The MCP server contains no generic shell or remediation capability.
AWS credentials come from the normal AWS credential chain. Never store static AWS credentials,
.envfiles, AWS config directories, private keys, Terraform state, plans, or local variable files in this repository.
Some AWS read-only APIs require Resource: "*", including EC2 Describe calls,
cloudwatch:GetMetricData, and the query-level logs:GetQueryResults and
logs:StopQuery calls. The runtime policy grants only that single CloudWatch
metrics action—no write, alarm, or broad CloudWatch read permissions. The
resource-scoped actions remain limited to the two approved log groups, the
exact lab instance, and the fixed SSM document.
Prerequisites
Python 3.11 or newer
Terraform 1.6 or newer
Codex with MCP server support
AWS CLI and the Session Manager plugin for administrative checks
An AWS Identity Center or other source profile with permission to deploy the Terraform configuration
An AWS Region configured; the example defaults to
ap-southeast-1
The target instance must have these tags:
Tag | Value |
|
|
|
|
Local setup
On Linux, macOS, or WSL:
cd <PROJECT_DIR>
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"On Windows PowerShell:
Set-Location <PROJECT_DIR>
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"If PowerShell blocks activation, call .venv\Scripts\python.exe directly.
Running tests
With the virtual environment active:
python -m pytest -vUseful Terraform checks that do not change AWS are:
terraform -chdir=infrastructure fmt -check -recursive
terraform -chdir=infrastructure validateTwo AWS profiles
Use separate profiles for deployment and diagnostics:
defaultis the source/administrator profile. Terraform uses it to create, update, and destroy the lab.mcp-lab-runtimeassumes the restricted runtime role. Only the MCP server should use it for diagnostics.
Do not run Terraform with mcp-lab-runtime; its limited permissions are
intentional and are not enough to refresh or manage Terraform resources.
Configure the assumed-role profile in your user-level AWS config, outside this repository:
[profile mcp-lab-runtime]
role_arn = arn:aws:iam::<AWS_ACCOUNT_ID>:role/aws-infra-ops-mcp-lab-runtime
source_profile = default
role_session_name = aws-infra-ops-mcp
duration_seconds = 3600
region = ap-southeast-1For an Identity Center deployment, set Terraform's
trusted_sso_role_arn_pattern to a placeholder-based value such as:
arn:aws:iam::<AWS_ACCOUNT_ID>:role/aws-reserved/sso.amazonaws.com/<AWS_REGION>/AWSReservedSSO_<PERMISSION_SET>_<IDENTITY_CENTER_ROLE_SUFFIX>Verify each profile outside the test suite:
aws sts get-caller-identity --profile default
aws sts get-caller-identity --profile mcp-lab-runtimeThe runtime result should contain
assumed-role/aws-infra-ops-mcp-lab-runtime/.
Terraform deployment
Terraform uses local state. Keep terraform.tfstate secure and backed up; it is
ignored by Git. Run deployment commands with the default profile:
export AWS_PROFILE=default
terraform -chdir=infrastructure init
terraform -chdir=infrastructure fmt -check -recursive
terraform -chdir=infrastructure validate
terraform -chdir=infrastructure plan -out=tfplan
terraform -chdir=infrastructure apply tfplanCopy infrastructure/terraform.tfvars.example to
infrastructure/terraform.tfvars only when overrides are needed. The local
file is ignored and must not be committed.
Applying the configuration creates billable AWS resources, including EC2, EBS,
a public IPv4 address, and CloudWatch Logs. Review the plan and current AWS
pricing before applying. For more infrastructure detail, see
infrastructure/README.md.
Connecting the server to Codex
Configure Codex to start the virtual-environment Python executable with the
repository's server.py. Use absolute paths and set the restricted AWS profile
for the child process:
[mcp_servers.aws-infra-ops-lab]
command = "<PROJECT_DIR>/.venv/bin/python"
args = ["<PROJECT_DIR>/server.py"]
env = { AWS_PROFILE = "mcp-lab-runtime", AWS_REGION = "ap-southeast-1" }On Windows, use paths such as
<PROJECT_DIR>\\.venv\\Scripts\\python.exe and
<PROJECT_DIR>\\server.py.
Restart Codex after changing its MCP configuration. The server waits silently for MCP messages on standard input; stdout is reserved for the protocol.
Three optional, non-secret environment variables are available:
AWS_LOG_GROUP_PREFIXchanges the default/aws/mcp-labprefix.AWS_SSM_DOCUMENT_NAMEchanges the deployed document name frommcp-lab-get-nginx-status.AWS_SSM_JOURNAL_DOCUMENT_NAMEchanges the deployed journal document name frommcp-lab-get-nginx-journal.
Changing these values also requires matching infrastructure and IAM policy configuration.
Example troubleshooting prompts
Run all four diagnostics:
Investigate the current health of web01 using all four approved diagnostic
tools. Check EC2 health, CPU, status-check and network metrics from the last 15
minutes, recent errors from the same period, and nginx service status. Separate
confirmed evidence from conclusions and show each data source.Check only nginx:
Call only get_service_status for web01 and nginx. Show the structured evidence
and its data source.Distinguish a clean stop from a startup or configuration failure:
Investigate why nginx is unavailable on web01. Call get_service_status and
get_service_journal with a 60-minute lookback and at most 50 entries. Separate
confirmed journal evidence from inference and identify every data source.get_service_journal is not a generic shell, Run Command, or remediation tool.
It cannot accept arbitrary commands and cannot start, stop, restart, or modify
nginx. The custom SSM document internally uses the aws:runShellScript
document plugin to execute its fixed, reviewed journalctl command. This is
distinct from granting access to the AWS-managed AWS-RunShellScript document,
which the runtime role is not permitted to invoke.
Correlate an outage:
The web service on web01 appears unavailable. Correlate EC2 health, errors from
the last 15 minutes, and nginx status. Identify the most likely cause using
only confirmed evidence.Controlled nginx incident test
This test deliberately stops nginx. Run it only in the disposable lab and use
the default administrator profile—not the MCP runtime profile.
Start a Session Manager session:
export AWS_PROFILE=default
aws ssm start-session --target "$(terraform -chdir=infrastructure output -raw instance_id)"Inside the session, stop nginx and write a recognizable system-log event:
sudo systemctl stop nginx
logger "MCP-LAB ERROR: nginx intentionally stopped for diagnostic validation"Ask Codex to run all four diagnostics. Confirm that EC2 remains healthy, the
metrics remain diagnostic-only, the log event appears, and nginx reports
inactive/dead.
Always restore the service before ending the test:
sudo systemctl start nginx
systemctl is-active nginx
curl --fail http://127.0.0.1/The expected final results are active and a successful local HTTP response.
The MCP server cannot perform this stop or restoration.
Common troubleshooting
Terraform returns AccessDenied while using mcp-lab-runtime
Switch to AWS_PROFILE=default. The runtime role is intentionally unable to
administer or fully refresh Terraform-managed resources.
CloudWatch StartQuery returns AccessDenied
Apply the current Terraform policy with the administrator profile. The two
approved log-group resource ARNs must end in :*, not the raw log-group ARN.
The MCP server cannot assume its runtime role
Check the source_profile, account ID, Region, runtime role ARN, and
trusted_sso_role_arn_pattern. Identity Center role suffixes can change when a
permission-set assignment is recreated.
get_service_status times out or returns an SSM error
Confirm that web01 is online in Systems Manager, the fixed document exists,
and the runtime policy references the exact instance and document ARNs.
Recent errors are empty
Check that the CloudWatch Agent is running, both approved log groups have a current stream, and the requested lookback covers the event. An empty result means no matching events were returned; it does not test HTTP reachability.
Codex does not show the tools
Use absolute paths in the MCP configuration, confirm the virtual environment contains the package, and restart Codex after configuration changes.
Current limitations
Only
web01andnginxare supported.There is no HTTP reachability or end-to-end request tool.
The bounded journal can provide service-local evidence but does not expose arbitrary units, full journal history, or general host inspection.
Error search uses a fixed keyword query rather than application-specific parsing.
CloudWatch Logs Insights queries may incur scan charges.
Terraform state is local; there is no remote backend or state locking.
The lab uses a public subnet and public IPv4 address for outbound access, although its security group allows no inbound traffic.
Diagnostics are intentionally read-only; recovery remains a separate, operator-controlled action.
Sensible next steps
Add a read-only HTTP or load-balancer health check with a fixed target.
Add CloudWatch metrics and selected CloudTrail events as bounded tools.
Move Terraform state to an encrypted remote backend with locking.
Add CI checks for tests, formatting, Terraform validation, and secret scanning.
Generalize allowlists through reviewed configuration while keeping tool schemas narrow and avoiding arbitrary queries or commands.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseBqualityBmaintenanceAn MCP server for read-only Linux system administration and diagnostics on RHEL-based systems via SSH. It enables users to troubleshoot remote hosts by accessing system information, services, logs, and network configurations through natural language.Last updated19272Apache 2.0
- Alicense-qualityCmaintenanceA secure SSH-based MCP server for diagnosing remote servers. It allows AI agents to execute read-only commands and read files automatically, while requiring user confirmation for write operations.Last updatedMIT
- Alicense-qualityCmaintenanceA read-only MCP server that lets an LLM inspect an AWS account — list EC2 instances, S3 buckets, IAM users, and cost — with a structural guarantee against any mutations.Last updatedMIT

masaro-infra-mcpofficial
Flicense-qualityCmaintenanceA secure MCP server providing read-only tools to interact with Cloudflare, Coolify, and other infrastructure services, enabling AI clients to safely diagnose and validate environments.Last updated
Related MCP Connectors
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Hosted Amazon Seller and Vendor MCP server for Claude, ChatGPT, Cursor, Codex, Gemini, Copilot.
An MCP server for deep research or task groups
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nhkm95/aws-infra-ops-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server