Skip to main content
Glama

image-gen-mcp

A local Model Context Protocol (MCP) server for generating and editing images through your own Azure Foundry deployment, with GitHub Copilot and editable presentation workflows.

Status: @juanmicrosoft/image-gen-mcp@0.1.0 is published. Anonymous archive integrity, fresh registry installation and MCP discovery are verified. The implemented tools have real Azure evidence, actual Copilot CLI evidence and a rendered presentation example. Verified paths include local macOS / Node 22.22.2 / Copilot CLI 1.0.79 and the CLI-backed VS Code 1.139.0 session with CLI-token generation/editing from an installed tarball. Native Local Copilot Chat has a demonstrated preview-off recovery/edit workflow; inline previews encountered a host image-transport failure. Do not equate the two session backends. See the qualification matrix and boundaries. A separate data-only Azure principal also generated and edited while an authenticated management read was denied. Live expired-session behavior remains unverified; ambiguous diagnostics do not claim a unique cause. Publishing-token revocation remains an open cleanup gate.

The configured model profile is gpt-image-2.5-sunburst. V1 deliberately enables only the live-verified 1536x864, high-quality PNG combination: one image, one concurrent submission, no automatic retries, prompt rewrites or model fallback. No exact ChatGPT backend or output parity is claimed.

Tool

Purpose

get_capabilities

Nonbillable diagnostics; keeps configured, observed and unverified facts separate

generate_image

One new immutable image, identified by a caller-retained operation UUID

edit_image

One explicit artifact or approved local PNG reference; immutable parent/child lineage

get_operation

Recover a saved result after interruption without submitting again

Get started

Let your coding agent do the setup. Use Copilot CLI, VS Code Copilot in a local agent session, or Claude Code on the machine where the MCP will run. Paste the prompts below in order. No Azure portal configuration is needed: the agent uses Azure CLI and this repository's provisioning script. You still complete interactive sign-in and approve subscription, resource costs and permissions. A subscription with model access/quota is required; an agent cannot bypass an administrator or capacity restriction.

Here, Claude means Claude Code, not Claude web or Claude Desktop. The Copilot workflows have live qualification; Claude Code setup is based on its documented stdio interface, not a live-tested image workflow. These instructions are for local agents, not a hosted coding agent that cannot access your local Azure login or files.

1. Ask the agent to prepare Azure

Set up https://github.com/juanmicrosoft/image-gen-mcp for me using Azure CLI,
without Azure portal configuration. Read AGENTS.md, docs/azure-setup.md,
docs/setup.md and docs/clients.md from a trusted checkout first.

Check Node (tested 22.22.2), npm, Azure CLI and Bicep. If something is missing,
propose the official installation command for my OS before installing it.
Check my az login and available subscriptions. If sign-in is needed, start
az login and let me complete it; never ask me to paste credentials into chat.
Ask me to choose/confirm the tenant and subscription, and whether to reuse
an existing compatible image deployment or create a dedicated new one.

For reuse, ask for the inference endpoint and deployment alias; verify the
model/version when authorized, otherwise get confirmation from its owner.
Do not modify an existing resource or treat its alias as proof of its model.

For a new deployment, use scripts/azure.mjs and infra/main.bicep from the
checkout, not an improvised deployment. Explain the proposed region, model,
public-network/key-enabled development topology, expected costs and required
resource-creation plus role-assignment permissions. Get my approval before
resource creation, provider registration or role changes. Check actual
model/version/SKU/quota availability; do not silently switch model or region.
Use --auth azure-cli, an explicitly confirmed subscription, unique account
and dedicated resource-group names, and an absolute private --state path.
Do not grant subscription-wide Owner or fall back to API keys on failure.
If access/quota is blocked, report the exact blocker and required admin action.

Keep ownership state for recovery/cleanup, including after partial failure.
Never delete/adopt unrelated resources or reset state to bypass ownership.
Read back deployment success and save the endpoint, deployment and output
directory in private local configuration, not source control. Do not expose
keys/tokens or generate any images. Finish with the configuration values
needed for the next prompt and the state path; do not claim inference is tested.

The agent can run az login --use-device-code when normal interactive login is unavailable. Sign-in may open a browser; that is authentication, not manual Azure resource configuration. New resources use the existing ownership-safe provision/cleanup flow. Provisioning is in the source checkout, not the npm runtime package.

2. Ask the agent to connect your client

Replace the bracketed choice with Copilot CLI, VS Code Copilot, or Claude Code:

Connect image-gen-mcp to [Copilot CLI / VS Code Copilot / Claude Code] on this
machine, using the Azure configuration from the previous step.

Install @juanmicrosoft/image-gen-mcp@0.1.0 from https://registry.npmjs.org/
with --ignore-scripts in a persistent local prefix. Use an absolute installed
entrypoint and the correct client-specific format in docs/clients.md.
Use IMAGE_GEN_AUTH=azure-cli, IMAGE_GEN_PREVIEW=false and an absolute private
output directory; do not pass AZURE_OPENAI_API_KEY in CLI-auth mode.
Pass the endpoint, deployment and any explicit tenant.
If we used an isolated Azure CLI profile, explicitly pass AZURE_CONFIG_DIR
to the MCP process. Ensure that process can find node and az.

Inspect existing client configuration first. Add only the image-gen entry
without overwriting other servers; stop on a name conflict. Keep personal
configuration outside source control and preserve host approval policies.
For Copilot CLI, prefer the installed configure-client.mjs helper and a
private session-local --additional-mcp-config file.
For VS Code, merge into my user MCP configuration; do not commit personal
settings in .vscode/mcp.json.
For Claude Code, use claude mcp add --scope local --transport stdio with
explicit environment variables and the installed executable; do not copy
Copilot-only tools/timeout fields into Claude configuration.

Reload/restart the host as needed; show me the exact launch or reload step.
Verify that this client discovers get_capabilities, generate_image,
edit_image and get_operation. Call only get_capabilities and explain any
unverified inference/permission checks. If you cannot inspect the client
session, tell me the exact check to run rather than claiming it is connected.
Do not request an image or automatically approve billable tools.

The client guide includes concrete configuration commands. There is no MCP OAuth login here: the local server uses the Azure CLI identity you authorized. Valid configuration or a working az login is not proof of inference permission.

3. Try one image, deliberately

Only after you are ready for an Azure image charge, paste:

I approve at most one billable image-generation request through image-gen.
Create a fresh operation UUID and retain it before submission. Generate a
1536x864 high-quality PNG of a lighthouse on a quiet rocky coast at dawn,
with open sky on the left for a title. Use the configured deployment only.
Do not retry, edit, switch models or submit another request automatically.
If the outcome or client transport is uncertain, call get_operation with
the same UUID instead of generating again. On success, inspect the saved
full-resolution PNG using a local image-reading tool if available and show
me its path. If you cannot inspect it, say so; do not claim visual quality.

To remove a new, task-owned deployment later, ask the agent to read the saved state, show the exact resource group and deletion impact, obtain your explicit approval, then run the documented ownership-checked cleanup command. Do not apply that cleanup to a reused deployment.

Manual installation alternative

For an existing compatible deployment:

Install the pinned release in a persistent directory:

npm install --prefix "$HOME/.local/share/image-gen-mcp" \
  --registry=https://registry.npmjs.org/ --ignore-scripts --no-audit --no-fund \
  @juanmicrosoft/image-gen-mcp@0.1.0

Or build from source:

git clone https://github.com/juanmicrosoft/image-gen-mcp.git
cd image-gen-mcp
npm ci
npm run build

Then follow existing-deployment setup to select credentials, set the inference endpoint/deployment/output directory, and create a private client configuration. Pin the package version and never paste a key into shell history. The server is a stdio protocol process, not an interactive image-generation command.

Starting from scratch? Use the separate owner-tagged Bicep/Azure CLI setup. The normal MCP never creates resources or retrieves management keys. The authorization record explains the required resource-scoped grant and remaining limits; a token or management access alone is not inference permission.

Ask Copilot to generate a hero with negative space, inspect it, then explicitly edit that artifact using a new operation UUID. Results include the immutable full-resolution path, dimensions, hash, lineage, available usage and an optional bounded image preview. Image requests are billable; diagnostics are not.

Related MCP server: platform-eng-copilot

Presentations, safety and evidence

The PptxGenJS example consumes full-resolution artifacts into three editable slides. Its dependencies stay outside the runtime. Reference continuity is probabilistic; generated art does not replace editable text.

Read configuration, client setup/limits, generation, editing, privacy/cost/recovery and bounded live evaluation. Previews can be disabled. Cancellation does not prove that a request stopped or was free; retain the operation ID and inspect it before explicitly submitting again.

For contributors, npm test is offline and requires no Azure credentials. See testing, CONTRIBUTING.md and AGENTS.md for the individual-PR, independent-review and evidence rules. The changelog and release evidence/gates separate registry verification from bounded live-client evidence and remaining cleanup.

The software and documentation are MIT licensed. See SECURITY.md. This license does not replace Azure/OpenAI service terms or guarantee rights in generated images or third-party source assets.

Available Tools

4 tools
edit_imageA
Idempotent

Create one new billable image from exactly one explicit source_artifact_id or approved absolute PNG source_path. Provide a self-contained prompt and new operation UUID. The source remains unchanged; masks, URLs and multiple references are unsupported.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNo1536x864
promptYes
qualityNohigh
source_pathNoAbsolute .png path inside an explicitly configured input root. Choose exactly one source; omit source_artifact_id when using this path.
operation_idYes
source_artifact_idNoChoose exactly one source. Omit source_path when using this artifact ID.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=false, destructiveHint=false and idempotentHint=true, but the description adds genuinely non-structured context: the call is billable, the source image is left unchanged, and several input shapes (masks, URLs, multiple references) are rejected. It doesn't describe the async/polling behavior or what the call returns, which matters for a create-and-bill operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler; the core action and its source constraint are front-loaded, followed by required inputs, then limitations. Every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, it covers cost, source immutability, parameter conflicts and input limits well. The main gap is that it never says the operation is asynchronous or that get_operation should be used to retrieve the resulting image, which an agent would need to complete the task.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (source_path and source_artifact_id documented), but the description adds meaning the schema lacks: prompt must be 'self-contained' and operation_id must be 'new'. The size/quality parameters are single-value enums with defaults, so they need no prose; the mutual-exclusion rule for the two source params largely repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create one new billable image') plus the required input ('from exactly one explicit source_artifact_id or approved absolute PNG source_path'). The source-based framing implicitly separates it from generate_image, but the sibling is never named, so the agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational constraints: exactly one source, a self-contained prompt, a new operation UUID, and an explicit when-not list ('masks, URLs and multiple references are unsupported'). It stops short of naming generate_image as the alternative when no source exists, so the routing decision is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA
Idempotent

Generate one billable Azure PNG at native 1536x864/high. Supply and retain a fresh UUID operation_id; replaying identical arguments recovers the same artifact without another submission. No retries, resizing or prompt rewrites.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNo1536x864
promptYes
qualityNohigh
operation_idYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already include idempotentHint=true, so replay behavior is partly structured, but the description adds valuable detail: the operation_id must be fresh and retained, replaying identical arguments returns the same artifact without another submission, and the tool performs no retries, resizing, or prompt rewrites. The billing note is also an important side effect that annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences deliver a high density of useful information with no filler. The core outcome is front-loaded, and each clause adds a distinct constraint or behavioral detail. Nothing in the description is redundant with the structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description covers the essential operational facts: output type, dimensions, quality, billing, idempotency, and explicit non-behaviors. It does not describe the exact response envelope or authentication requirements, but it supplies enough for an agent to invoke the tool and understand the artifact contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the bare schema. It explicitly covers operation_id semantics (fresh UUID, replay behavior), names the exact accepted size and quality ('native 1536x864/high'), and implies the prompt is used verbatim via 'no prompt rewrites'. This meaningfully supplements the schema, even though prompt constraints like maxLength are not restated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate one billable Azure PNG'. It also fixes the native size and quality, making the tool's purpose unambiguous. The action clearly contrasts with sibling tools like edit_image and get_operation, so an agent can identify this as the creation tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The verb 'Generate' and the sizing/quality constraints imply the tool is for creating new images, but the description never explicitly says when to prefer it over siblings or what conditions would call for edit_image or get_operation. It provides clear operational context, such as the need for a UUID, but no direct when-to-use versus when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_capabilitiesC
Read-onlyIdempotent

Nonbillable configuration and capability diagnostics. Optional credential acquisition does NOT verify inference permission or deployment identity.

ParametersJSON Schema
NameRequiredDescriptionDefault
check_credentialsNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, which conveys safety. The description adds a critical behavioral caveat: optional credential acquisition does NOT verify inference permission or deployment identity, which is valuable for the agent to know. However, it does not elaborate on other behaviors such as what 'capability diagnostics' includes or any side effects beyond the annotation, so it earns a 3.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is clear and front-loaded with the core purpose. The caveat about credential acquisition is placed second, which is effective. It is concise without unnecessary fluff, earning a 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's relative simplicity (one optional parameter, no output schema), the description covers the essential purpose and a key behavioral caveat. However, it lacks specifics on what 'capability diagnostics' return or how to interpret results, and it does not clarify whether credentials are required for certain operations. This is a moderate gap, so a 3 is justified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameter 'check_credentials'. It mentions 'optional credential acquisition', which relates to the parameter, but it does not explicitly state that setting check_credentials to true will attempt credential acquisition or what the implications are. The parameter's meaning is partially implied but not fully specified, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource ('configuration and capability diagnostics') and a domain ('nonbillable'), but the verb 'get' is generic, and the purpose is somewhat vague. It does not clearly distinguish from siblings like get_operation, which also retrieves information. The mention of 'Nonbillable' hints at a billing context, but the overall purpose could be more specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives like get_operation. The description implies it is for diagnostics, but does not state when it should be preferred over other sibling tools, nor does it provide context for when credential checking is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_operationA
Read-onlyIdempotent

Look up a retained operation UUID and recover a committed artifact after interruption. Never submits or resubmits an image request. Unknown completion may still incur charges.

ParametersJSON Schema
NameRequiredDescriptionDefault
operation_idYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds two valuable behavioral details: it never submits or resubmits (consistent with readOnly) and warns that unknown completion may still incur charges. This charge caveat is beyond what annotations provide and is important for cost awareness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero filler. The purpose is front-loaded, the exclusion is stated immediately, and the cost warning is a final compact addition. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only tool with no output schema, the description covers what it does, when to use it, what it doesn't do, and a financial caveat. Annotations handle the safety profile. Nothing critical is missing for an agent to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention the parameter operation_id at all. However, the schema already provides format: uuid, making the parameter self-explanatory. With a single parameter and clear schema, the description adds no extra meaning but also doesn't need to. Baseline for low schema coverage is 3, and here it's adequate because the parameter is trivial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Look up' and the resource 'retained operation UUID' and 'committed artifact', and explicitly distinguishes it from submission tools by saying it never submits or resubmits an image request. This makes it easy for an agent to understand this is a retrieval/status tool, distinct from generate_image and edit_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after interruption for recovering artifacts, which provides clear context. It explicitly states it never submits or resubmits, ruling out using this tool for creation. However, it doesn't name the sibling tools directly or explicitly say 'use this instead of generate_image when you need the result of a prior call', so the guidance is strong but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changededit_image3 fields changed
      • removedInput schema / oneOf
        Removed value: -[
        -  {
        -    "required": [
        -      "source_artifact_id"
        -    ]
        -  },
        -  {
        -    "required": [
        -      "source_path"
        -    ]
        -  }
        -]
      • addedInput schema / properties / source_artifact_id / description
        Added value: +"Choose exactly one source. Omit source_path when using this artifact ID."
      • changedInput schema / properties / source_path / description
        Previous value: -"Absolute .png path inside an explicitly configured input root."New value: +"Absolute .png path inside an explicitly configured input root. Choose exactly one source; omit source_artifact_id when using this path."
  2. 4 tool updatesv0.1.0
    • First observededit_image
    • First observedgenerate_image
    • First observedget_capabilities
    • First observedget_operation

TDQS

A3.8/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: generate_image creates new images, edit_image transforms an existing artifact, get_capabilities reports config, and get_operation looks up prior operations. The descriptions explicitly delimit boundaries (e.g. get_operation never submits requests, edit_image requires a source artifact).

Naming Consistency5/5

All four names follow a strict verb_noun snake_case pattern (edit_image, generate_image, get_capabilities, get_operation). The convention is predictable and consistent throughout.

Tool Count4/5

Four tools is a small but coherent set for a focused image generation/editing server. It is slightly lean, but each tool earns its place without overlap.

Completeness4/5

Core lifecycle is covered: generate, edit, capability diagnostics, and operation retrieval for recovery after interruption. Minor gaps exist, such as no operation listing/enumeration or cleanup, but the essential workflows are present.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers