Azure Image Pipeline MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Azure Image Pipeline MCPShow the latest 10 image pipeline runs."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Azure Image Pipeline MCP
A focused Model Context Protocol server for an Azure VM image-baking pipeline. It exposes three read-only tools for pipeline runs, build status and release evidence, plus one strongly guarded tool for queueing a pipeline run.
Why a focused server?
Microsoft provides an official Azure DevOps MCP server. Use that server for broad Azure DevOps access. This sample is intentionally narrower: it demonstrates how to present image-release concepts as domain tools, enforce a parameter allow-list and keep pipeline queueing disabled by default.
Related MCP server: Buildkite MCP Server
Tools
list_recent_image_runsget_image_run_statusget_image_run_evidencequeue_image_pipeline(disabled by default)
Prerequisites
Python 3.10+
uvorpipAzure DevOps project and pipeline access
For local development, an Azure DevOps PAT with the minimum required scope
Run locally
cp .env.example .env
# Edit .env with non-secret settings; inject AZDO_PAT through your shell or secret store.
uv sync --extra dev
uv run mcp dev src/azure_image_mcp/server.pyTo use stdio from VS Code, retain .vscode/mcp.json, restart VS Code, and start the server from the MCP server controls.
Safe write behaviour
Queueing remains blocked unless both conditions are met:
MCP_ALLOW_QUEUE_RUN=trueThe caller supplies the exact confirmation value
QUEUE_IMAGE_PIPELINE
The server also rejects template parameters outside MCP_ALLOWED_PARAMETERS. These controls do not replace Azure DevOps checks, approvals, branch policies, environments or service-connection permissions.
Test
uv sync --extra dev
uv run pytest -qThe test suite uses mocked HTTP responses and does not contact Azure DevOps.
Example prompts
"Show the latest 10 image pipeline runs."
"Get status and source commit for image run 1234."
"Which expected evidence artifacts are missing from run 1234?"
"Queue a Java 21 CIS image build in UK South." The host should first display parameters and request explicit confirmation.
Authentication note
This repository uses PAT authentication only as a compact local-development example. For enterprise usage, prefer Microsoft Entra-backed authentication and the hosted Azure DevOps MCP server where the client supports it.
Available Tools
4 toolsget_image_run_evidenceC
List evidence artifacts and report which expected image-release artifacts are missing.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that missing artifacts are reported, but says nothing about permissions, whether it reads remote storage, how 'expected' artifacts are determined, or pagination/size limits for a potentially large artifact list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with no filler, front-loading the listing behavior before the diagnostic report behavior. Appropriate length for the scope described.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't enumerate return fields, and the summary of both outputs (artifacts list + missing report) covers the shape. However, with no annotations and an undocumented required parameter, the definition leaves the agent guessing on safety profile and precondition details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter 'run_id' is undocumented; the description only implicitly ties it to a run. It adds no format, range, or sourcing detail (e.g., where to obtain a valid run_id) beyond the property name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('evidence artifacts') plus a second function (report missing expected image-release artifacts). This is clearly distinct from siblings like get_image_run_status or list_recent_image_runs, though it never explicitly names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this tool versus get_image_run_status or list_recent_image_runs, and no prerequisites. Usage is only inferable from the verb 'evidence' implying a diagnostic/verification role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_image_run_statusB
Get pipeline and build status for one image-pipeline run.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Get' implies a read-only operation and 'status' hints at the data returned, but since an output schema exists that already defines the response, the description adds no behavioral context such as freshness, whether status is cached, or what states are possible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the verb and resource come first. It is appropriately terse, though the terseness borders on under-specification for a tool with an undocumented parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and the read-only nature is inferable from 'Get', so the definition is minimally workable. But with no annotations and an undocumented required parameter, it leaves gaps an agent would need to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter run_id has 0% schema description coverage, so the description must compensate and does not — it only implies that the id selects 'one image-pipeline run' without stating format, origin, or how to obtain a valid value. This falls well short of what a 0%-coverage parameter needs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('pipeline and build status for one image-pipeline run'), which clearly separates it from get_image_run_evidence (evidence, not status) and list_recent_image_runs (list, not single-run). It does not explicitly name those siblings, but the resource noun is distinct enough that an agent can tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one image-pipeline run' implies the single-run scope and that a run_id is needed, which distinguishes it from the list tool by inference. However, there is no explicit when-to-use guidance and no mention of the alternative tools (evidence vs status vs list) despite three relevant siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recent_image_runsC
List recent runs for the configured Azure DevOps image pipeline.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It indicates a read-style listing operation but says nothing about authentication, ordering, pagination, rate limits, or what 'recent' means, leaving important behavior undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-formed sentence with no wasted words and the core action is front-loaded. It is concise, though its brevity contributes to the gaps in guidance and parameter detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return value details need not be in the description. For a low-complexity list tool, the description is minimally adequate, but it omits the meaning of 'recent,' the 'top' limit parameter, and any routing guidance relative to sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the single optional parameter 'top,' and the description does not mention it at all. The word 'recent' hints at a limit but does not explain the parameter's meaning, default, or effect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource ('List recent runs') and scopes it to the 'configured Azure DevOps image pipeline.' It is distinguishable from siblings like get_image_run_status and queue_image_pipeline, but it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The implied use case is listing recent runs, but the agent is left to infer it without help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queue_image_pipelineB
Queue an image pipeline only when writes are enabled and the exact confirmation token is supplied.
| Name | Required | Description | Default |
|---|---|---|---|
| branch | No | refs/heads/main | |
| location | Yes | ||
| confirmation | Yes | ||
| image_flavour | Yes | ||
| image_version | No | ||
| runtime_version | Yes | ||
| enable_security_scan | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the write-enablement gate and exact confirmation-token requirement, but it omits side effects of queuing, error behavior for invalid tokens, idempotency, and any authorization details beyond 'writes are enabled.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It is structurally clean, though extremely terse for a mutation tool with seven parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a seven-parameter mutation tool with no annotations and zero schema description coverage, the description is too thin. The output schema can cover return values, but the description still needs to clarify input semantics, preconditions, and side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across seven parameters, so the description must compensate and largely does not. It hints that confirmation must be an exact token, but runtime_version, image_flavour, location, branch, image_version, and enable_security_scan remain semantically undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Queue an image pipeline.' The action is clearly distinct from the read-oriented siblings list_recent_image_runs, get_image_run_status, and get_image_run_evidence, but the description never explicitly differentiates or names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition for use: only when writes are enabled and the exact confirmation token is supplied. This implies when not to use it, but it does not name alternative tools or explain what to do if the condition is not met.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
get_image_run_evidence - First observed
get_image_run_status - First observed
list_recent_image_runs - First observed
queue_image_pipeline
TDQS
Scored across 4 tools
Each tool targets a distinct resource+action: listing recent runs, fetching status of one run, gathering evidence artifacts, and queueing a pipeline. No two tools overlap in purpose, so misselection is unlikely.
All four names follow a consistent verb_noun pattern (get_image_run_*, list_recent_image_runs, queue_image_pipeline) with a shared 'image' domain prefix. Predictable and readable throughout.
Four tools is on the lean side but fits the narrow monitoring-and-queueing scope of an image pipeline. Each tool earns its place, though one or two more (e.g. logs/cancel) could round it out.
The surface covers the core lifecycle: discover runs, inspect status, gather release evidence, and trigger a run under a guarded confirmation. Minor gaps like canceling/retrying a run or fetching raw logs are workable around.
Maintenance
Related MCP Connectors
Inspect Depot builds, CI runs, job logs and Actions runners; retry or cancel CI runs.
Trigger and inspect Codemagic CI builds, apps and artifacts.
Remote MCP for Kiro release readiness, evidence binders, signoff, and CI approval receipts.
Provides capabilities that let LLM agents perform a range of infrastructure management tasks.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceExposes Power Platform Pipeline operations as MCP tools for Copilot Studio agents, enabling pipeline discovery, deployments, approvals, and configuration management.-
- FlicenseNot gradedqualityDmaintenanceEnables interaction with Buildkite CI/CD to list organizations, pipelines, builds, jobs, and logs, as well as retry jobs.4-
- FlicenseNot gradedqualityDmaintenanceEnables interaction with Azure DevOps, including listing pipelines and runs, viewing logs and artifacts, collecting and normalizing security reports, and triggering pipelines with safety controls.-
- FlicenseNot gradedqualityBmaintenanceExposes Jenkins CI/CD capabilities as MCP tools, including job management, builds, pipelines, and Terraform/GCP log inspection.-