Skip to main content
Glama
psprakhar020

Azure Image Pipeline MCP

by psprakhar020

Azure Image Pipeline MCP

A focused Model Context Protocol server for an Azure VM image-baking pipeline. It exposes three read-only tools for pipeline runs, build status and release evidence, plus one strongly guarded tool for queueing a pipeline run.

Why a focused server?

Microsoft provides an official Azure DevOps MCP server. Use that server for broad Azure DevOps access. This sample is intentionally narrower: it demonstrates how to present image-release concepts as domain tools, enforce a parameter allow-list and keep pipeline queueing disabled by default.

Related MCP server: Buildkite MCP Server

Tools

  • list_recent_image_runs

  • get_image_run_status

  • get_image_run_evidence

  • queue_image_pipeline (disabled by default)

Prerequisites

  • Python 3.10+

  • uv or pip

  • Azure DevOps project and pipeline access

  • For local development, an Azure DevOps PAT with the minimum required scope

Run locally

cp .env.example .env
# Edit .env with non-secret settings; inject AZDO_PAT through your shell or secret store.
uv sync --extra dev
uv run mcp dev src/azure_image_mcp/server.py

To use stdio from VS Code, retain .vscode/mcp.json, restart VS Code, and start the server from the MCP server controls.

Safe write behaviour

Queueing remains blocked unless both conditions are met:

  1. MCP_ALLOW_QUEUE_RUN=true

  2. The caller supplies the exact confirmation value QUEUE_IMAGE_PIPELINE

The server also rejects template parameters outside MCP_ALLOWED_PARAMETERS. These controls do not replace Azure DevOps checks, approvals, branch policies, environments or service-connection permissions.

Test

uv sync --extra dev
uv run pytest -q

The test suite uses mocked HTTP responses and does not contact Azure DevOps.

Example prompts

  • "Show the latest 10 image pipeline runs."

  • "Get status and source commit for image run 1234."

  • "Which expected evidence artifacts are missing from run 1234?"

  • "Queue a Java 21 CIS image build in UK South." The host should first display parameters and request explicit confirmation.

Authentication note

This repository uses PAT authentication only as a compact local-development example. For enterprise usage, prefer Microsoft Entra-backed authentication and the hosted Azure DevOps MCP server where the client supports it.

Available Tools

4 tools
get_image_run_evidenceC

List evidence artifacts and report which expected image-release artifacts are missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that missing artifacts are reported, but says nothing about permissions, whether it reads remote storage, how 'expected' artifacts are determined, or pagination/size limits for a potentially large artifact list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence with no filler, front-loading the listing behavior before the diagnostic report behavior. Appropriate length for the scope described.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description needn't enumerate return fields, and the summary of both outputs (artifacts list + missing report) covers the shape. However, with no annotations and an undocumented required parameter, the definition leaves the agent guessing on safety profile and precondition details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'run_id' is undocumented; the description only implicitly ties it to a run. It adds no format, range, or sourcing detail (e.g., where to obtain a valid run_id) beyond the property name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('evidence artifacts') plus a second function (report missing expected image-release artifacts). This is clearly distinct from siblings like get_image_run_status or list_recent_image_runs, though it never explicitly names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this tool versus get_image_run_status or list_recent_image_runs, and no prerequisites. Usage is only inferable from the verb 'evidence' implying a diagnostic/verification role.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_image_run_statusB

Get pipeline and build status for one image-pipeline run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Get' implies a read-only operation and 'status' hints at the data returned, but since an output schema exists that already defines the response, the description adds no behavioral context such as freshness, whether status is cached, or what states are possible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and resource come first. It is appropriately terse, though the terseness borders on under-specification for a tool with an undocumented parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values and the read-only nature is inferable from 'Get', so the definition is minimally workable. But with no annotations and an undocumented required parameter, it leaves gaps an agent would need to guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter run_id has 0% schema description coverage, so the description must compensate and does not — it only implies that the id selects 'one image-pipeline run' without stating format, origin, or how to obtain a valid value. This falls well short of what a 0%-coverage parameter needs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('pipeline and build status for one image-pipeline run'), which clearly separates it from get_image_run_evidence (evidence, not status) and list_recent_image_runs (list, not single-run). It does not explicitly name those siblings, but the resource noun is distinct enough that an agent can tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for one image-pipeline run' implies the single-run scope and that a run_id is needed, which distinguishes it from the list tool by inference. However, there is no explicit when-to-use guidance and no mention of the alternative tools (evidence vs status vs list) despite three relevant siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recent_image_runsC

List recent runs for the configured Azure DevOps image pipeline.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It indicates a read-style listing operation but says nothing about authentication, ordering, pagination, rate limits, or what 'recent' means, leaving important behavior undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single well-formed sentence with no wasted words and the core action is front-loaded. It is concise, though its brevity contributes to the gaps in guidance and parameter detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return value details need not be in the description. For a low-complexity list tool, the description is minimally adequate, but it omits the meaning of 'recent,' the 'top' limit parameter, and any routing guidance relative to sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the single optional parameter 'top,' and the description does not mention it at all. The word 'recent' hints at a limit but does not explain the parameter's meaning, default, or effect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource ('List recent runs') and scopes it to the 'configured Azure DevOps image pipeline.' It is distinguishable from siblings like get_image_run_status and queue_image_pipeline, but it does not explicitly contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The implied use case is listing recent runs, but the agent is left to infer it without help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queue_image_pipelineB

Queue an image pipeline only when writes are enabled and the exact confirmation token is supplied.

ParametersJSON Schema
NameRequiredDescriptionDefault
branchNorefs/heads/main
locationYes
confirmationYes
image_flavourYes
image_versionNo
runtime_versionYes
enable_security_scanNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the write-enablement gate and exact confirmation-token requirement, but it omits side effects of queuing, error behavior for invalid tokens, idempotency, and any authorization details beyond 'writes are enabled.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is structurally clean, though extremely terse for a mutation tool with seven parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a seven-parameter mutation tool with no annotations and zero schema description coverage, the description is too thin. The output schema can cover return values, but the description still needs to clarify input semantics, preconditions, and side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across seven parameters, so the description must compensate and largely does not. It hints that confirmation must be an exact token, but runtime_version, image_flavour, location, branch, image_version, and enable_security_scan remain semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Queue an image pipeline.' The action is clearly distinct from the read-oriented siblings list_recent_image_runs, get_image_run_status, and get_image_run_evidence, but the description never explicitly differentiates or names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear precondition for use: only when writes are enabled and the exact confirmation token is supplied. This implies when not to use it, but it does not name alternative tools or explain what to do if the condition is not met.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedget_image_run_evidence
    • First observedget_image_run_status
    • First observedlist_recent_image_runs
    • First observedqueue_image_pipeline

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct resource+action: listing recent runs, fetching status of one run, gathering evidence artifacts, and queueing a pipeline. No two tools overlap in purpose, so misselection is unlikely.

Naming Consistency5/5

All four names follow a consistent verb_noun pattern (get_image_run_*, list_recent_image_runs, queue_image_pipeline) with a shared 'image' domain prefix. Predictable and readable throughout.

Tool Count4/5

Four tools is on the lean side but fits the narrow monitoring-and-queueing scope of an image pipeline. Each tool earns its place, though one or two more (e.g. logs/cancel) could round it out.

Completeness4/5

The surface covers the core lifecycle: discover runs, inspect status, gather release evidence, and trigger a run under a guarded confirmation. Minor gaps like canceling/retrying a run or fetching raw logs are workable around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers