mcp-aws-observability-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-aws-observability-serverAre there any active alarms on the platform right now?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-aws-observability-server
A small, dependency-free MCP (Model Context Protocol) server that
exposes AWS CloudWatch Logs and Alarms as tools an LLM client (Claude
Desktop, an agent runtime, a custom MCP client) can call — list_log_groups,
search_logs, and get_active_alarms.
This is a reference implementation modelled on the kind of MCP server I've built and deployed to AWS at Equal Experts for centralized logging/monitoring/observability across a shared GenAI platform: the same tool surface, but standing on a mock backend here instead of a real account, so anyone can clone and run it in under a minute.
Why no SDK dependency
The official mcp SDK is great, but for a reference/demo repo I wanted
zero install friction and the transport mechanics to be visible rather
than hidden behind a library. src/mcp_observability/protocol.py
implements the newline-delimited JSON-RPC 2.0 stdio transport and the
initialize / tools/list / tools/call lifecycle directly from the
MCP specification. It's
~200 lines and fully tested — a good place to actually read how MCP
works under the hood.
Related MCP server: mcp-cloudwatch-explorer
Quickstart
git clone https://github.com/swapnilbabladkar/mcp-aws-observability-server.git
cd mcp-aws-observability-server
# run the test suite (stdlib unittest, no install required)
PYTHONPATH=src python3 -m unittest discover -s tests -v
# run the full stdio flow against a real subprocess
python3 examples/demo_client.py
# or run the server directly (reads JSON-RPC from stdin, writes to stdout)
PYTHONPATH=src python3 -m mcp_observabilityUsing it from Claude Desktop
Add to your MCP client config (e.g. Claude Desktop's
claude_desktop_config.json):
{
"mcpServers": {
"aws-observability": {
"command": "python3",
"args": ["-m", "mcp_observability"],
"env": { "PYTHONPATH": "/absolute/path/to/mcp-aws-observability-server/src" }
}
}
}Restart Claude Desktop and ask it something like "any active alarms on the platform right now?" or "search the mcp-server log group for errors in the last two hours."
Switching to real AWS data
By default the server runs on MockObservabilityBackend, which returns
realistic canned data (log groups/events/alarms shaped like a real
EKS-hosted MCP server + RAG pipeline platform) so the whole tool-call
flow works with zero AWS setup.
To point it at a real account, install the optional AWS extra and swap
the backend in server.py:
pip install -e ".[aws]"# server.py
from .backends import AWSObservabilityBackend
def default_server() -> MCPServer:
return build_server(AWSObservabilityBackend(region_name="eu-west-1"))AWSObservabilityBackend (in backends.py) implements the same
interface via boto3's logs and cloudwatch clients — real
describe_log_groups / filter_log_events / describe_alarms calls,
paginated. It needs a role/profile with logs:Describe*,
logs:FilterLogEvents, and cloudwatch:DescribeAlarms.
Project layout
src/mcp_observability/
protocol.py # MCP JSON-RPC/stdio transport — the actual protocol implementation
backends.py # ObservabilityBackend interface + Mock and AWS implementations
server.py # registers the 3 tools against a backend
__main__.py # `python -m mcp_observability` entrypoint
tests/ # unittest coverage for protocol + mock backend
examples/
demo_client.py # spawns the server as a subprocess and drives it end-to-endRunning the tests
PYTHONPATH=src python3 -m unittest discover -s tests -v16 tests, covering the JSON-RPC error cases (parse errors, unknown methods, unknown tools, tool-level failures vs. protocol failures) as well as the mock backend's filtering logic.
License
MIT — see LICENSE.
Available Tools
3 toolsget_active_alarmsA
List CloudWatch alarms currently in ALARM state, optionally filtered by namespace (e.g. 'AWS/EKS', 'Platform/RAG').
| Name | Required | Description | Default |
|---|---|---|---|
| namespace | No | Optional CloudWatch namespace filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses a behavioral constraint (only ALARM-state alarms, not OK/INSUFFICIENT_DATA), but says nothing about permissions, pagination, result-size limits, or ordering for what is implicitly a read-only list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence that front-loads the resource and the state constraint before the optional filter. Every clause earns its place; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional parameter and no output schema, the description supplies what an agent needs to call it correctly. Minor gaps around pagination/limits keep it from a 5, but given the tool's low complexity this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single optional parameter is already documented as a 'CloudWatch namespace filter'. The description reinforces this with concrete example values ('AWS/EKS', 'Platform/RAG'), adding modest practical value, which lands at the baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List), resource (CloudWatch alarms), and a precise scope (currently in ALARM state, optionally namespace-filtered). An agent can immediately tell this returns only alarms in the ALARM state, distinct from the log-oriented siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description's scoping ('currently in ALARM state'), and the sibling tools (list_log_groups, search_logs) operate on a different domain so no explicit alternative is needed. However, there is no explicit when-to-use statement or mention of when the namespace filter should be applied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_log_groupsB
List CloudWatch log groups on the platform, optionally filtered by name prefix (e.g. '/eks/platform-cluster/').
| Name | Required | Description | Default |
|---|---|---|---|
| prefix | No | Optional log group name prefix filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose pagination limits, result ordering, max results, or required permissions for listing log groups — significant omissions for a list tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the core action and the optional modifier are both stated economically with nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-optional-param list tool with no output schema, the description covers the action and filter adequately but omits any hint of return shape, result volume, or pagination. Completeness is minimal-viable rather than thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one optional parameter, so the baseline is 3. The description's example prefix ('/eks/platform-cluster/') adds a small amount of practical value beyond the schema's 'Optional log group name prefix filter', but no format or matching semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (CloudWatch log groups) plus scope (on the platform) and the optional filter. It does not explicitly distinguish itself from search_logs or get_active_alarms, but the read-and-enumerate purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'optionally filtered by name prefix' implies when the prefix parameter applies, but there is no guidance on when to use this tool versus search_logs (log events) or get_active_alarms. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_logsA
Search a CloudWatch log group for a text pattern within a recent time window. Use this for quick root-cause triage across the platform's services.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max events to return (default 20) | |
| pattern | Yes | Substring to search for | |
| log_group | Yes | Log group name | |
| since_minutes | No | How far back to search, in minutes (default 60) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Search' implies a read-only, non-destructive operation and the 'recent time window' phrase hints at the default lookback, which is useful. However, it omits permissions/auth requirements, rate limits, and result/pagination behavior (e.g., how limit interacts with results).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and followed by a compact routing hint. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, 2-required read tool with no output schema and no annotations, the description covers the what and a rough when, but leaves behavioral gaps: no return shape, no pagination/limit interaction, no auth caveats. Adequate but not fully complete given the disclosure burden.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented with type, default, and meaning in the schema. The description adds only the loose notion of a 'recent time window,' which does not extend the since_minutes or limit semantics. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Search a CloudWatch log group for a text pattern within a recent time window.' It is clearly distinguishable from list_log_groups and get_active_alarms by naming the CloudWatch log group and pattern-search action. No sibling is named explicitly, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Adds a usage context: 'quick root-cause triage across the platform's services.' This implies when to reach for it, but gives no when-not guidance, prerequisites, or explicit routing against list_log_groups or get_active_alarms. Usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_active_alarms - First observed
list_log_groups - First observed
search_logs
TDQS
Scored across 3 tools
Each tool targets a distinct CloudWatch resource and action: listing log groups, searching within a log group, and listing active alarms. There is no meaningful overlap between them, so an agent can select correctly without hesitation.
All three tools follow a strict verb_noun snake_case pattern (list_log_groups, search_logs, get_active_alarms). The convention is uniform and predictable.
Three tools is on the thin side for an observability server, but each one is well-scoped and earns its place for log/alarm triage. It sits slightly under what the domain could support rather than being excessive.
Logs and alarms are covered, but the observability surface omits other core CloudWatch pillars such as metrics and traces, and there is no way to inspect alarm history or log stream metadata. Agents can triage common issues but will hit dead ends for metric-based investigations.
Maintenance
Related MCP Connectors
MCP-Native LLM Orchestration Agent
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
- gatewayOAuthai.sealgate
MCP gateway with runtime security policy, tool-call-level control, and audit of agent actions.
Query application logs, traces, and metrics from your AI coding assistant via Foam's MCP server.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables LLMs to autonomously query AWS CloudWatch Logs and perform structured root-cause analysis via natural language prompts, using MCP tools for log group listing and Insights queries.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to query AWS CloudWatch metrics, alarms, and logs read-only via MCP, providing rapid health snapshots and triage without console navigation.1MIT
- FlicenseNot gradedqualityCmaintenanceEnables natural language queries for AWS CloudWatch logs, metrics, and alarms via an LLM agent with MCP tools.-
- AlicenseNot gradedqualityBmaintenanceA read-only MCP server that exposes tools to query Datadog monitors and logs, and AWS CloudWatch Logs, and to group recurring errors by fingerprint. Designed for use with Claude Code to diagnose issues and propose fixes.MIT