Skip to main content
Glama
swapnilbabladkar

mcp-aws-observability-server

mcp-aws-observability-server

A small, dependency-free MCP (Model Context Protocol) server that exposes AWS CloudWatch Logs and Alarms as tools an LLM client (Claude Desktop, an agent runtime, a custom MCP client) can call — list_log_groups, search_logs, and get_active_alarms.

This is a reference implementation modelled on the kind of MCP server I've built and deployed to AWS at Equal Experts for centralized logging/monitoring/observability across a shared GenAI platform: the same tool surface, but standing on a mock backend here instead of a real account, so anyone can clone and run it in under a minute.

Why no SDK dependency

The official mcp SDK is great, but for a reference/demo repo I wanted zero install friction and the transport mechanics to be visible rather than hidden behind a library. src/mcp_observability/protocol.py implements the newline-delimited JSON-RPC 2.0 stdio transport and the initialize / tools/list / tools/call lifecycle directly from the MCP specification. It's ~200 lines and fully tested — a good place to actually read how MCP works under the hood.

Related MCP server: mcp-cloudwatch-explorer

Quickstart

git clone https://github.com/swapnilbabladkar/mcp-aws-observability-server.git
cd mcp-aws-observability-server

# run the test suite (stdlib unittest, no install required)
PYTHONPATH=src python3 -m unittest discover -s tests -v

# run the full stdio flow against a real subprocess
python3 examples/demo_client.py

# or run the server directly (reads JSON-RPC from stdin, writes to stdout)
PYTHONPATH=src python3 -m mcp_observability

Using it from Claude Desktop

Add to your MCP client config (e.g. Claude Desktop's claude_desktop_config.json):

{
  "mcpServers": {
    "aws-observability": {
      "command": "python3",
      "args": ["-m", "mcp_observability"],
      "env": { "PYTHONPATH": "/absolute/path/to/mcp-aws-observability-server/src" }
    }
  }
}

Restart Claude Desktop and ask it something like "any active alarms on the platform right now?" or "search the mcp-server log group for errors in the last two hours."

Switching to real AWS data

By default the server runs on MockObservabilityBackend, which returns realistic canned data (log groups/events/alarms shaped like a real EKS-hosted MCP server + RAG pipeline platform) so the whole tool-call flow works with zero AWS setup.

To point it at a real account, install the optional AWS extra and swap the backend in server.py:

pip install -e ".[aws]"
# server.py
from .backends import AWSObservabilityBackend

def default_server() -> MCPServer:
    return build_server(AWSObservabilityBackend(region_name="eu-west-1"))

AWSObservabilityBackend (in backends.py) implements the same interface via boto3's logs and cloudwatch clients — real describe_log_groups / filter_log_events / describe_alarms calls, paginated. It needs a role/profile with logs:Describe*, logs:FilterLogEvents, and cloudwatch:DescribeAlarms.

Project layout

src/mcp_observability/
  protocol.py   # MCP JSON-RPC/stdio transport — the actual protocol implementation
  backends.py   # ObservabilityBackend interface + Mock and AWS implementations
  server.py     # registers the 3 tools against a backend
  __main__.py   # `python -m mcp_observability` entrypoint
tests/          # unittest coverage for protocol + mock backend
examples/
  demo_client.py  # spawns the server as a subprocess and drives it end-to-end

Running the tests

PYTHONPATH=src python3 -m unittest discover -s tests -v

16 tests, covering the JSON-RPC error cases (parse errors, unknown methods, unknown tools, tool-level failures vs. protocol failures) as well as the mock backend's filtering logic.

License

MIT — see LICENSE.

Available Tools

3 tools
get_active_alarmsA

List CloudWatch alarms currently in ALARM state, optionally filtered by namespace (e.g. 'AWS/EKS', 'Platform/RAG').

ParametersJSON Schema
NameRequiredDescriptionDefault
namespaceNoOptional CloudWatch namespace filter

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses a behavioral constraint (only ALARM-state alarms, not OK/INSUFFICIENT_DATA), but says nothing about permissions, pagination, result-size limits, or ordering for what is implicitly a read-only list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence that front-loads the resource and the state constraint before the optional filter. Every clause earns its place; nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional parameter and no output schema, the description supplies what an agent needs to call it correctly. Minor gaps around pagination/limits keep it from a 5, but given the tool's low complexity this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single optional parameter is already documented as a 'CloudWatch namespace filter'. The description reinforces this with concrete example values ('AWS/EKS', 'Platform/RAG'), adding modest practical value, which lands at the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List), resource (CloudWatch alarms), and a precise scope (currently in ALARM state, optionally namespace-filtered). An agent can immediately tell this returns only alarms in the ALARM state, distinct from the log-oriented siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the description's scoping ('currently in ALARM state'), and the sibling tools (list_log_groups, search_logs) operate on a different domain so no explicit alternative is needed. However, there is no explicit when-to-use statement or mention of when the namespace filter should be applied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_log_groupsB

List CloudWatch log groups on the platform, optionally filtered by name prefix (e.g. '/eks/platform-cluster/').

ParametersJSON Schema
NameRequiredDescriptionDefault
prefixNoOptional log group name prefix filter

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose pagination limits, result ordering, max results, or required permissions for listing log groups — significant omissions for a list tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the core action and the optional modifier are both stated economically with nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-optional-param list tool with no output schema, the description covers the action and filter adequately but omits any hint of return shape, result volume, or pagination. Completeness is minimal-viable rather than thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is only one optional parameter, so the baseline is 3. The description's example prefix ('/eks/platform-cluster/') adds a small amount of practical value beyond the schema's 'Optional log group name prefix filter', but no format or matching semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (CloudWatch log groups) plus scope (on the platform) and the optional filter. It does not explicitly distinguish itself from search_logs or get_active_alarms, but the read-and-enumerate purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'optionally filtered by name prefix' implies when the prefix parameter applies, but there is no guidance on when to use this tool versus search_logs (log events) or get_active_alarms. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_logsA

Search a CloudWatch log group for a text pattern within a recent time window. Use this for quick root-cause triage across the platform's services.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax events to return (default 20)
patternYesSubstring to search for
log_groupYesLog group name
since_minutesNoHow far back to search, in minutes (default 60)

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Search' implies a read-only, non-destructive operation and the 'recent time window' phrase hints at the default lookback, which is useful. However, it omits permissions/auth requirements, rate limits, and result/pagination behavior (e.g., how limit interacts with results).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and followed by a compact routing hint. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, 2-required read tool with no output schema and no annotations, the description covers the what and a rough when, but leaves behavioral gaps: no return shape, no pagination/limit interaction, no auth caveats. Adequate but not fully complete given the disclosure burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented with type, default, and meaning in the schema. The description adds only the loose notion of a 'recent time window,' which does not extend the since_minutes or limit semantics. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Search a CloudWatch log group for a text pattern within a recent time window.' It is clearly distinguishable from list_log_groups and get_active_alarms by naming the CloudWatch log group and pattern-search action. No sibling is named explicitly, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Adds a usage context: 'quick root-cause triage across the platform's services.' This implies when to reach for it, but gives no when-not guidance, prerequisites, or explicit routing against list_log_groups or get_active_alarms. Usage is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedget_active_alarms
    • First observedlist_log_groups
    • First observedsearch_logs

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a distinct CloudWatch resource and action: listing log groups, searching within a log group, and listing active alarms. There is no meaningful overlap between them, so an agent can select correctly without hesitation.

Naming Consistency5/5

All three tools follow a strict verb_noun snake_case pattern (list_log_groups, search_logs, get_active_alarms). The convention is uniform and predictable.

Tool Count4/5

Three tools is on the thin side for an observability server, but each one is well-scoped and earns its place for log/alarm triage. It sits slightly under what the domain could support rather than being excessive.

Completeness3/5

Logs and alarms are covered, but the observability surface omits other core CloudWatch pillars such as metrics and traces, and there is no way to inspect alarm history or log stream metadata. Agents can triage common issues but will hit dead ends for metric-based investigations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to query AWS CloudWatch metrics, alarms, and logs read-only via MCP, providing rapid health snapshots and triage without console navigation.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A read-only MCP server that exposes tools to query Datadog monitors and logs, and AWS CloudWatch Logs, and to group recurring errors by fingerprint. Designed for use with Claude Code to diagnose issues and propose fixes.
    MIT