Skip to main content
Glama
vivekanjana76

azure-compliance-mcp

README.md
# azure-compliance-mcp

[![CI](https://github.com/vivekanjana76/azure-compliance-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/vivekanjana76/azure-compliance-mcp/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](./LICENSE)
[![Python 3.12+](https://img.shields.io/badge/python-3.12%2B-blue.svg)](https://www.python.org/)
[![FastMCP 3.x](https://img.shields.io/badge/FastMCP-3.x-purple.svg)](https://gofastmcp.com)

> **Honest Azure compliance for LLM agents** — ask infra-health questions in plain English and get answers that distinguish a real `fail` from a `not_evaluable`, and **never fake a `pass`**.

An [MCP](https://modelcontextprotocol.io) server, built with **[FastMCP](https://gofastmcp.com) 3.x**, that exposes read-only Azure resource-compliance data so an agent can answer questions like *"which storage accounts allow TLS 1.0?"* or *"which VMs have no disk encryption?"* — each finding tagged with **where the verdict came from** and **whether it could be evaluated at all**.

![Demo](docs/demo.gif)

## Why honest compliance?

Most tooling collapses "compliant", "non-compliant", and "couldn't tell" into a single pass/fail bit — so a control it *can't actually measure* silently shows green. This server refuses to guess:

- **`fail`** — positive evidence a protection is absent (e.g. `encryptionAtHost = false`). Actionable; shows up in the default view.
- **`not_evaluable`** — the signal genuinely isn't in the data (e.g. no guest-config policy assigned). Surfaced explicitly, never as a pass.
- **`source`** — every finding records the signal that decided its verdict: `arg` | `azure_policy` | `graph` | `defender`.

## Quickstart

Requires [uv](https://docs.astral.sh/uv/) and Python 3.12+. **Runs out of the box** — mock mode needs zero Azure setup.

```bash
uv sync                          # install dependencies
uv run python scripts/demo.py    # fastest look — runs the tools against mock data
uv run server.py                 # mock data, stdio transport — works immediately
```

Point it at your own tenant when you're ready:

```bash
uv run server.py --mode live     # real Azure Resource Graph (e.g. after `az login`)
uv run server.py --transport http  # remote: Streamable HTTP
uv run fastmcp dev inspector server.py  # explore the tools interactively
```

Live mode authenticates with `DefaultAzureCredential` against **your own** tenant and is strictly read-only.

## Architecture

Tools depend only on a `Provider` **protocol**, never on a concrete data source — so the *same* tool code runs against synthetic mock data or live Azure Resource Graph (ARG), selected at startup with `--mode`.

```mermaid
flowchart LR
    Agent["LLM agent / MCP client"] -->|stdio · HTTP| Server["FastMCP server\n(5 read-only tools)"]
    Server --> Protocol["Provider protocol\n(read-only)"]
    Protocol -->|--mode mock| Mock["MockProvider\nseeded ARG-shaped data"]
    Protocol -->|--mode live| Live["LiveProvider\nDefaultAzureCredential"]
    Live -->|injection-safe KQL| ARG[("Azure Resource Graph\nresources · policyresources ·\npatchassessmentresources · authorizationresources")]
    Live -->|getByIds| Graph[("Microsoft Graph\nprincipal existence")]
```

**Controls** (`check_compliance` evaluates five):

| control | source | signal |
|---|---|---|
| `required_tags` | `arg` | `resources.tags` (env / owner / costCenter) |
| `tls_min_1_2` | `arg` | `resources.properties.minimumTlsVersion` |
| `public_network_access` | `arg` | `resources.properties.publicNetworkAccess` |
| `disk_encryption` | `arg` | `securityProfile.encryptionAtHost` — host-level only; OFF ⇒ `fail`, ADE not assessed |
| `guest_config_extension` | `azure_policy` | `policyresources` guest-config states; no data ⇒ `not_evaluable` |

**Tools** (5 of a 5–6 budget):

| tool | source | honest behavior |
|---|---|---|
| `check_compliance` | `arg` / `azure_policy` | the five controls above; unmeasurable signal ⇒ `not_evaluable`, never a fake pass |
| `query_resources` | `arg` | read-only projection with filters pushed into KQL |
| `get_patch_status` | `arg` | Update Manager assessments from `patchassessmentresources`; deallocated VM or no assessment data ⇒ `not_evaluable` with `pendingUpdateCount: null` (unknown ≠ 0), never a fake "current" |
| `find_orphaned_rbac` | `graph` | assignments from ARG `authorizationresources`, principal existence resolved via Microsoft Graph; `orphaned` only on a positive directory not-found — unresolvable (no `Directory.Read.All`) ⇒ `not_evaluable`, **never** orphaned on a guess |

**Transports:** `stdio` (default, local) and Streamable HTTP (remote, intended behind OAuth 2.1). In stdio mode stdout is reserved for the protocol — logging goes to stderr only.

See [`SPEC.md`](./SPEC.md) for the full tool contracts.

## Design decisions & tradeoffs

- **Mock provider is the default.** The repo runs with zero Azure account, credentials, or network — so reviewers, CI, and new contributors get meaningful output immediately. The mock dataset is shaped exactly like ARG rows, so one mapping serves both modes.
- **Read-only by construction.** No tool can create, update, or delete a resource. The live path only ever issues ARG *queries*; there is no mutation surface to misuse.
- **Control-based, not Azure-Policy-based.** Compliance is evaluated from resource *configuration* (the ARG row), so it works even when no Azure Policy is assigned — except where the signal genuinely lives elsewhere (guest-config), which routes to `policyresources` and is honest about it.
- **The `not_evaluable` + `source` honesty model.** A third status plus a provenance field means the agent can tell "this is broken" from "I can't see this from here" — the core reason to trust the output.
- **KQL pushdown for scale.** Live filters (`query_resources`) are pushed *into* the ARG query (`where` + `take`) rather than fetched-then-filtered, so large tenants stay cheap. An opt-in contract test asserts the pushdown matches the reference Python filter, so both modes stay provably consistent.
- **Injection-safe escaping.** ARG has no bind parameters, so every user-supplied value is encoded as an escaped KQL string literal — never concatenated raw. (ARG is read-only regardless, but the discipline is enforced and unit-tested.)

## Status & roadmap

**v0.3.0** — `check_compliance`, `query_resources`, `get_patch_status`, and `find_orphaned_rbac` shipped (mock + live), honest `not_evaluable`/`source` model across all of them, KQL pushdown, CI on every PR.

Next:

- [ ] `summarize_health` — rolled-up infra-health summary
- [ ] evals — task-level evaluation of agent answers over the mock dataset

## Development

```bash
uv run pytest               # tests (live tests are opt-in: RUN_LIVE_TESTS=1)
uv run ruff check .         # lint
uv run ruff format --check . # format
```

### Evals (opt-in)

Tests prove the tools behave; the eval suite ([SPEC §5](./SPEC.md)) checks the
property the server exists for: *given a natural-language question, does an
agent call the right tool, with the right arguments, and answer correctly* —
including saying "that can't be determined" when the honest answer is
`not_evaluable`.

```bash
# needs ANTHROPIC_API_KEY (env or the gitignored .env); calls the API — costs money
uv run --group evals python evals/run.py                     # full suite (~20 cases)
uv run --group evals python evals/run.py --cases rbac        # filter by case id
uv run --group evals python evals/run.py --agent-model <id>  # evaluate another agent model

# alternate provider for cost-free smoke runs (needs OPENROUTER_API_KEY);
# a non-Anthropic judge is OFF-PIN vs SPEC §5.4 — the report flags it
uv run --group evals python evals/run.py --provider openrouter \
    --agent-model <openrouter-id> --judge-model <openrouter-id>
```

Runs against **mock mode only** (deterministic, no Azure credentials). Results
land in `evals/results/<timestamp>-<model>.json` plus a rendered summary;
they're gitignored except deliberately committed `baseline-*` files. Structured
assertions are exact-match; open-ended prose is scored by a **pinned** LLM
judge whose ~85–92% human-agreement caveat is printed in every report — judge
scores are signal, not ground truth. Deliberately **not** in CI (SPEC §5.6).

## Security

- All Azure tools are **read-only** — nothing in this server modifies Azure resources.
- Secrets, tenant IDs, and `.env` are gitignored and must never be committed; live mode uses only your local `DefaultAzureCredential`.
- In stdio mode, logging goes to **stderr only** (stdout is reserved for the protocol).

## License

[MIT](./LICENSE)

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a distinct purpose: check_compliance evaluates specific security controls, get_patch_status reports VM patch posture, and query_resources performs generic read-only queries. No overlap in functionality.

Naming Consistency5/5

All tool names use lowercase with underscores and follow a clear verb_noun pattern (check_compliance, get_patch_status, query_resources), providing a predictable and consistent naming scheme.

Tool Count4/5

With only 3 tools, the server is tightly focused on Azure compliance and resource queries. While the count is somewhat minimal, each tool is well-justified and covers the core use cases without being overloaded.

Completeness4/5

The tool set covers compliance evaluation, patch status, and generic resource queries, which together address a broad range of compliance workflows. The generic query tool compensates for potential gaps, though some specialized compliance checks might be missing.

Maintenance

ActivityStale
ResponsivenessUnresponsive