Skip to main content
Glama

Azure Image Generation MCP

Model Context Protocol (MCP) server for AI-powered image generation using Azure DALL-E 3 and FLUX models

License: MIT Node.js Version

🎨 Overview

A powerful MCP server that brings professional AI image generation to LibreChat. Generate stunning images using Azure's DALL-E 3 for photorealistic content or FLUX for creative artwork, with intelligent automatic model selection based on your prompts.

Perfect for LibreChat users who want seamless image generation capabilities powered by Azure AI Foundry models.

Related MCP server: ImageGen MCP Server

✨ Features

  • 🤖 Dual Model Support

    • DALL-E 3: Photorealistic images, portraits, and artistic content

    • FLUX (FLUX.1-Kontext-pro): Creative illustrations and flexible generation

  • 🧠 Intelligent Model Selection

    • Automatic model selection based on prompt analysis

    • FLUX as default for optimal results

    • DALL-E 3 when explicitly requested or optimal

  • 📐 Multiple Image Sizes

    • Square (1024x1024) - Perfect for social media

    • Wide (1792x1024) - Great for banners and headers

    • Tall (1024x1792) - Ideal for posters and vertical content

  • ⚙️ Customization Options

    • Quality settings (standard/HD) for DALL-E 3

    • Style options (vivid/natural) for DALL-E 3

    • Fast generation times (typically 30-60 seconds)

  • 🔌 Easy Integration

    • Works seamlessly with LibreChat

    • Compatible with MCP clients

    • Simple configuration via environment variables

📋 Prerequisites

  • Node.js >= 18.0.0

  • Azure OpenAI API access with:

    • DALL-E 3 deployment (optional)

    • FLUX deployment (FLUX.1-Kontext-pro)

  • LibreChat instance (for LibreChat integration)

🚀 Installation

npm install -g azure-image-generation-mcp

Option 2: From Source

git clone https://github.com/malikmalikayesha/azure-image-generation-mcp.git
cd azure-image-generation-mcp
npm install

Option 3: NPX (No Installation)

npx azure-image-generation-mcp

⚙️ Configuration

1. Environment Variables

Create a .env file or set environment variables:

AZURE_IMAGE_API_KEY=your_azure_api_key_here
AZURE_IMAGE_BASE_URL=https://your-endpoint.cognitiveservices.azure.com/openai/deployments

2. LibreChat Integration

Add to your librechat.yaml:

mcpServers:
  "Image Generation":
    type: stdio
    command: node
    args:
      - /path/to/azure-image-generation-server.js

    name: "Image Generation"
    displayName: "Image Generation"

    timeout: 180000      # 3 minutes for generation
    initTimeout: 60000   # 1 minute startup

    chatMenu: true       # Show in chat tools

    serverInstructions: |
      🎨 AI Image Generation Tool

      Create stunning images using DALL-E 3 or FLUX models.
      Simply describe what you want to see!

    env:
      AZURE_IMAGE_API_KEY: "${AZURE_IMAGE_API_KEY}"
      AZURE_IMAGE_BASE_URL: "${AZURE_IMAGE_BASE_URL}"

📖 Usage

In LibreChat

Simply ask the AI to generate an image:

"Generate an image of a serene mountain landscape at sunset"
"Create a modern minimalist logo for a tech startup"
"Draw a realistic portrait of a confident businesswoman"
"Make an abstract pattern with geometric shapes"

Model Selection

  • Automatic (Default): The system intelligently chooses between DALL-E 3 and FLUX

  • FLUX (Default): Used for most requests unless DALL-E is explicitly mentioned

  • DALL-E 3: Explicitly request by mentioning "DALL-E" in your prompt

Advanced Options

Specify additional parameters in your request:

"Generate a wide landscape image in HD quality using DALL-E"
Size: 1792x1024, Quality: HD, Model: DALL-E 3

"Create a tall poster with vivid colors"
Size: 1024x1792, Style: vivid

🔧 Docker Deployment (LibreChat)

If using Docker with LibreChat, add to your Dockerfile:

# Install MCP SDK dependencies
RUN npm install @modelcontextprotocol/sdk@^1.17.2

# Copy Azure image generation files
COPY azure-image-generation-server.js ./

Then ensure your docker-compose.yml includes the environment variables:

services:
  api:
    environment:
      - AZURE_IMAGE_API_KEY=${AZURE_IMAGE_API_KEY}
      - AZURE_IMAGE_BASE_URL=${AZURE_IMAGE_BASE_URL}

🛠️ API Reference

Tool: generate_image

Generates an AI image based on a text prompt.

Parameters

Parameter

Type

Required

Default

Description

prompt

string

Yes

-

Description of the image to generate

model

string

No

auto

Model selection: dall-e-3, flux, or auto

size

string

No

1024x1024

Image dimensions: 1024x1024, 1792x1024, 1024x1792

style

string

No

vivid

DALL-E style: vivid or natural

quality

string

No

standard

DALL-E quality: standard or hd

Response

Returns a structured response with:

  • Text description of the generated image

  • Base64-encoded PNG image data

  • Metadata (model used, size, generation time)

🐛 Troubleshooting

Common Issues

Images not displaying in Azure models:

  • Ensure you're using LibreChat with the MCP image rendering fix (included in LibreChat v0.7.9+)

  • Check that your librechat.yaml configuration is correct

MCP server fails to start:

  • Verify environment variables are set correctly

  • Check that Node.js version is >= 18.0.0

  • Ensure @modelcontextprotocol/sdk is installed

API errors:

  • Verify your Azure API key is valid

  • Check that the base URL points to your Azure OpenAI endpoint

  • Ensure your Azure deployment has DALL-E 3 or FLUX enabled

Generation timeout:

  • Increase timeout value in librechat.yaml (default: 180000ms)

  • Check your network connectivity to Azure

Debug Mode

Enable debug logging by checking LibreChat logs:

# Docker
docker logs librechat-api

# Local
DEBUG=* npm start

📝 Example Prompts

Photorealistic Images

"A professional headshot of a software engineer in a modern office"
"Sunset over Tokyo skyline with Mount Fuji in the distance"
"Close-up of fresh vegetables on a wooden cutting board"

Artistic & Creative

"Minimalist logo design for a coffee shop called 'Bean Dreams'"
"Watercolor painting of a cottage in a flower garden"
"Abstract geometric pattern in blues and golds"

Marketing & Design

"Modern tech startup hero banner image, wide format"
"Instagram post background with pastel gradients"
"Professional LinkedIn banner for a data scientist"

🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository

  2. Create your feature branch (git checkout -b feature/AmazingFeature)

  3. Commit your changes (git commit -m 'Add some AmazingFeature')

  4. Push to the branch (git push origin feature/AmazingFeature)

  5. Open a Pull Request

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

🙏 Acknowledgments

📬 Support


Made with ❤️ for the LibreChat community

Available Tools

1 tool
generate_imageA

🎨 Create stunning AI-generated images using Azure DALL-E 3 or FLUX models with intelligent model selection

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescribe the image you want to create in natural language. Be detailed for best results. Examples: "A serene mountain landscape at sunset", "Modern minimalist logo design", "Cute cartoon mascot for a coffee shop"
modelNoChoose AI model: "dall-e-3" (photorealistic, artistic), "flux" (creative, flexible), or "auto" (smart selection based on prompt)auto
sizeNoImage dimensions: Square (1024x1024) for social media, Wide (1792x1024) for banners, Tall (1024x1792) for posters1024x1024
styleNoVisual style (DALL-E only): "vivid" for dramatic/artistic, "natural" for realistic/subduedvivid
qualityNoImage quality (DALL-E only): "standard" for faster generation, "hd" for higher detailstandard

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions model selection and creation capabilities, it lacks critical information about rate limits, authentication requirements, cost implications, response format, or error handling. For a generative AI tool with potential costs and limitations, this represents significant gaps in behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally concise and well-structured in a single sentence that communicates the core capability, technology stack, and key feature. Every element earns its place: the emoji adds visual context, 'Create stunning AI-generated images' states the purpose, 'using Azure DALL-E 3 or FLUX models' specifies technology, and 'with intelligent model selection' highlights differentiation. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (AI image generation with multiple models and parameters) and lack of both annotations and output schema, the description is incomplete. While concise and clear about purpose, it doesn't address behavioral aspects like cost, rate limits, or response format that are crucial for such a tool. The excellent schema coverage helps, but the description alone doesn't provide sufficient context for safe and effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides no parameter-specific information beyond what's already comprehensively documented in the input schema (100% coverage). The schema includes detailed descriptions, examples, enums, and defaults for all parameters. The description adds no additional semantic context about parameters, so it meets but doesn't exceed the baseline expectation when schema coverage is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('Create stunning AI-generated images') and resources ('using Azure DALL-E 3 or FLUX models'). It distinguishes the tool's unique capability of 'intelligent model selection' which adds differentiation even without sibling tools. The description goes beyond just restating the name by specifying the technology and key feature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context through phrases like 'intelligent model selection' and mentions of specific models, but provides no explicit guidance on when to use this tool versus alternatives. There are no sibling tools mentioned, so the lack of comparative guidance is understandable, but it doesn't offer any when/when-not advice or prerequisites for successful use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.4/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion or overlap between tools. The tool's purpose is clearly defined and distinct by default.

Naming Consistency5/5

A single tool inherently has consistent naming, as there are no other tools to compare it against. The name 'generate_image' follows a clear verb_noun pattern.

Tool Count2/5

One tool is too few for a server focused on Azure image generation, as it lacks operations like listing models, checking generation status, or managing images. This minimal set limits agent capabilities and feels incomplete for the domain.

Completeness1/5

The tool set is severely incomplete for image generation; it only provides generation without supporting operations like model selection, status tracking, or image management. This will cause significant agent failures in workflows requiring more than basic generation.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/malikmalikayesha/Azure-Image-Generation-MCP'

If you have feedback or need assistance with the MCP directory API, please join our Discord server