imagen-mcp
The imagen-mcp server provides intelligent multi-provider image generation and editing via MCP tools, supporting OpenAI (gpt-image-2) and Google Gemini with automatic provider selection based on prompt content.
Core Tools:
Generate images — Create images from text prompts with auto or manual provider selection; supports custom sizes, quality tiers, aspect ratios, and output paths
Batch generation — Generate up to 50 images concurrently with per-item error isolation
Conversational refinement — Iteratively refine images across multi-turn dialogues with conversation history tracking
Edit existing images — Inpaint or modify images using OpenAI's gpt-image-2 with optional mask support
List providers — View configured providers, their capabilities, and best use cases
List conversations — Browse saved multi-turn sessions for resumption or review
List Gemini models — Query available Gemini image generation models
Estimate cost — Approximate generation cost without actually generating an image
Key Features:
Auto provider routing — OpenAI for text-heavy images, diagrams, comics, and infographics; Gemini for photorealistic portraits, product photography, and 4K output
Reference images — Up to 14 base64-encoded reference images for style/character consistency (Gemini only)
Google Search grounding — Incorporate real-time data (weather, stocks, events) into images (Gemini only)
High-resolution output — Up to 4K with Gemini; up to 3840px with OpenAI
Multiple output formats — PNG, JPEG, WebP; results as markdown or JSON
Fallback notices — Clear warnings when the optimal provider isn't configured
Generates images using Google Gemini's image models (Nano Banana Pro, Flash, Imagen 3.0) with support for photorealism, up to 4K resolution, reference images for character/style consistency, real-time data via Google Search grounding, and conversational history for iterative refinement.
Generates images using OpenAI's GPT-Image-1 model, optimized for text-heavy images like menus, infographics, comics, and diagrams with excellent text rendering capabilities.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@imagen-mcpcreate a logo for a coffee shop called 'Morning Brew' with a minimalist design"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
imagen-mcp
A Model Context Protocol (MCP) server for intelligent multi-provider image generation.
Quick Start
Version 0.4.0 is released as the immutable v0.4.0 Git tag and is not
published on PyPI. Install it from that tag or a local checkout, then run the
canonical module or console script:
git clone https://github.com/michaeljabbour/imagen-mcp.git
cd imagen-mcp
python3 -m pip install .
python -m imagen_mcp
# equivalent after installation: imagen-mcpFor a reproducible VCS install, use the release tag:
python3 -m pip install \
"imagen-mcp @ git+https://github.com/michaeljabbour/imagen-mcp.git@v0.4.0"The historical python -m src.server entry point remains available for
compatibility in 0.4.x.
See CHANGELOG.md for release details and migration notes.
1. Get an API key (at least one):
Provider | Get a key at | Environment variable |
OpenAI |
| |
Google Gemini |
|
Having both keys lets auto-selection use either provider. With only one key, soft preferences may fall back with a notice. An unavailable explicit provider pin fails closed; in auto mode, Gemini-only requirements such as reference images and Google Search grounding also fail closed rather than silently dropping the requested capability.
2. Add to your MCP client (pick one):
claude mcp add -s user imagen \
-e OPENAI_API_KEY=sk-... \
-e GEMINI_API_KEY=AI... \
-- imagen-mcpVerify it's registered:
claude mcp listReference: Claude Code MCP docs
Edit the config file:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"imagen": {
"command": "imagen-mcp",
"args": [],
"env": {
"OPENAI_API_KEY": "sk-...",
"GEMINI_API_KEY": "AI..."
}
}
}
}Restart Claude Desktop (Cmd+Q, then reopen) after editing.
Reference: Claude Desktop MCP docs
Option A — CLI command:
codex mcp add imagen -- imagen-mcpOption B — edit ~/.codex/config.toml directly:
[mcp_servers.imagen]
command = "imagen-mcp"
[mcp_servers.imagen.env]
OPENAI_API_KEY = "sk-..."
GEMINI_API_KEY = "AI..."Reference: Codex MCP docs
Edit ~/.gemini/settings.json:
{
"mcpServers": {
"imagen": {
"command": "imagen-mcp",
"args": [],
"env": {
"OPENAI_API_KEY": "sk-...",
"GEMINI_API_KEY": "AI..."
}
}
}
}Reference: Gemini CLI MCP docs
Setting | Value |
Command |
|
Args |
|
Environment |
|
3. Generate an image — ask your AI assistant:
"Generate a professional headshot with studio lighting"
That's it. The server picks a provider automatically and uses Gemini 3.1 Flash Image when Gemini is selected unless you explicitly request another Gemini model.
Related MCP server: ImageGen MCP Server
Features
Auto Provider Selection — analyzes prompts to choose the best provider
Multi-Provider Support — OpenAI gpt-image-2 and Google Gemini 3 Image
Reference Images — up to 14 images for character/style consistency (Gemini)
Real-time Data — Google Search grounding for current info (Gemini)
Conversational Refinement — iteratively refine images with context
High Resolution — up to 4K output (Gemini)
Fallback Notices — clear warnings when a prompt would benefit from a provider you haven't configured
How Auto-Selection Works
The server analyzes your prompt and routes it to the best provider:
"Create a menu card for an Italian restaurant" -> OpenAI (text rendering)
"Professional headshot with studio lighting" -> Gemini (photorealism)
"Infographic about climate change" -> OpenAI (diagram + text)
"Product shot of perfume on marble" -> Gemini (product photography)What if the best provider isn't configured? The server falls back to whatever you have and tells you:
Provider Fallback: Gemini would be better for this prompt (Photorealistic content), but it's not configured. Using OpenAI instead. Set
GEMINI_API_KEYfor better results.
You can always override auto-selection with the provider parameter:
generate_image(prompt="...", provider="openai")
generate_image(prompt="...", provider="gemini")Explicit provider pins never fall back. In auto mode, hard Gemini requirements (reference images, Google Search grounding, and recognized real-time-data requests) also fail if Gemini is unavailable. Soft quality preferences may fall back to the configured provider and include a notice.
Provider Comparison
Feature | OpenAI gpt-image-2 | Gemini 3 Image |
Text Rendering | Excellent | Good |
Photorealism | Good | Excellent |
Latency | Varies by size/quality | Varies by model/size |
Max Resolution | 3840px edge / 8,294,400 pixels | 4K |
Sizes | Constrained custom | 1K, 2K, 4K; Flash also 0.5K |
Aspect Ratios | Up to 3:1 | 10 baseline presets; 14 on Gemini 3.1 Flash/Flash Lite |
Reference Images | Via | Yes (model-specific, up to 14) |
Real-time Data | No | Yes (Google Search) |
Use OpenAI for: text-heavy images, menus, infographics, comics, diagrams
Use Gemini for: portraits, product photography, 4K output, reference images
For gpt-image-2, both edges must be multiples of 16 and no larger than
3840px, the long-to-short ratio must be at most 3:1, and total pixels must be
between 655,360 and 8,294,400. Outputs above 2560x1440 are experimental.
gpt-image-2 does not support transparent backgrounds; use an opaque output
and a downstream background-removal step.
MCP Tools
Tool | Description |
| Main tool with auto provider selection (reports progress) |
| Generate many prompts concurrently (bounded fan-out, per-item error isolation) |
| Multi-turn refinement; native MCP elicitation with dialogue fallback |
| Edit/inpaint an existing image via OpenAI gpt-image-2 |
| List active conversations and their history |
| Show available providers and capabilities |
| Query available Gemini image models |
| Approximate generation cost without generating |
All tools advertise MCP tool annotations (read-only / open-world hints) so clients can reason about their side effects.
Output Location
Images are saved to ~/Downloads/images/{provider}/ by default (openai/ or gemini/ subdirectories).
Customize with:
# Save to a specific directory (auto-generated filename)
generate_image(prompt="...", output_path="~/Desktop/logos/")
# Save to a specific file
generate_image(prompt="...", output_path="~/Desktop/logos/my-logo.png")Set OUTPUT_DIR to change the base directory globally. Logs go to {OUTPUT_DIR}/logs/.
Gemini-Specific Features
# High resolution
generate_image(prompt="...", size="4K")
# Specific model
generate_image(prompt="...", gemini_model="gemini-3.1-flash-image")
# Reference images for style/character consistency (base64 encoded)
generate_image(prompt="...", reference_images=["base64..."])
# Real-time data via Google Search
generate_image(prompt="Current weather in NYC", enable_google_search=True)Search grounding requires Markdown output and a client that renders the returned Google Search Suggestions HTML plus associated source links. JSON output fails before calling Gemini because escaped HTML is not a compliant rendered surface.
Available Models
OpenAI
Model ID | Description |
| Default image generation and editing model |
| Legacy compatibility model |
| Legacy compatibility model |
Gemini
Model ID | Description |
| Nano Banana 2; default GA model, 0.5K/1K/2K/4K |
| Nano Banana Pro; GA model, 1K/2K/4K |
| Nano Banana Lite; 1K only, no Search, up to 14 object references |
Retired *-preview IDs are rejected with an actionable GA migration message;
explicit model pins never silently change. Gemini 3.1 Flash Lite Image outputs include SynthID and C2PA
provenance metadata; downstream transforms should preserve that metadata when
the file format and processing pipeline allow it.
Architecture
flowchart TB
subgraph Clients["MCP Clients"]
CD[Claude Desktop]
CC[Claude Code CLI]
GC[Gemini CLI]
CX[Codex CLI]
end
subgraph Server["imagen-mcp Server"]
MCP[MCP Protocol Layer]
subgraph Tools["MCP Tools"]
GI[generate_image]
CI[conversational_image]
LP[list_providers]
LM[list_gemini_models]
end
subgraph Core["Core Components"]
PS[Provider Selector]
PR[Provider Registry]
end
subgraph Providers["Image Providers"]
OAI[OpenAI Provider<br/>gpt-image-2]
GEM[Gemini Provider<br/>Gemini 3.1 Flash Image]
end
end
subgraph APIs["External APIs"]
OAPI[OpenAI API]
GAPI[Google Gemini API]
end
subgraph Storage["Local Storage"]
DL[~/Downloads/images/]
end
CD & CC & GC & CX --> MCP
MCP --> Tools
GI & CI --> PS
PS --> PR
PR --> OAI & GEM
OAI --> OAPI
GEM --> GAPI
OAI & GEM --> DLEnvironment Variables
Variable | Description | Required |
| OpenAI API key | At least one API key |
| Google Gemini API key | At least one API key |
| Alias for | |
| Base directory for saved images | No (default: |
| OS-path-separator list of roots that | No |
| Delete persisted conversational history after this many inactive days; | No (default: |
| Force a default provider | No (default: |
| Default OpenAI image size | No (default: |
| Default Gemini image size | No (default: |
| Opt in to prompt enhancement; adds an assistant-model API call, latency, and cost before generation | No (default: |
| Enable Google Search grounding | No (default: |
| Read-timeout ceiling in seconds for provider calls (covers slow high-quality renders) | No (default: |
| OpenAI client-side rate limits | No (defaults: |
| Gemini client-side rate limits | No (defaults: |
| Log directory override | No |
| Log level (DEBUG, INFO, etc.) | No |
| Log full prompts | No (default: |
|
| No |
| Bind address for HTTP transports | No (default: |
Troubleshooting
"No providers available"
You need at least one API key. Set OPENAI_API_KEY or GEMINI_API_KEY in your MCP client config (see Quick Start above).
Images generate but quality isn't great for portraits/products
You're probably missing GEMINI_API_KEY. The server fell back to OpenAI and showed a warning. Add a Gemini key for better photorealistic results.
Images generate but text looks bad
You're probably missing OPENAI_API_KEY. Add an OpenAI key for better text rendering.
"imagen-mcp: command not found"
Ensure the Python environment used by your MCP client has imagen-mcp
installed from a local checkout or pinned VCS revision and its scripts directory
is on PATH. As a fallback, configure the command as python with args -m,
imagen_mcp.
Where are my images saved?
Default: ~/Downloads/images/openai/ or ~/Downloads/images/gemini/. Check the tool output for the exact path. Set OUTPUT_DIR to change this.
How do I check which providers are active?
Use the list_providers tool, or run:
python3 -c "from imagen_mcp.providers import get_provider_registry; print(get_provider_registry().list_providers())"Development
# Clone and install (runtime + dev tooling)
git clone https://github.com/michaeljabbour/imagen-mcp.git
cd imagen-mcp
pip install -e ".[dev]" # or: uv sync --extra dev
# Install pre-commit hooks (ruff + mypy)
pre-commit install
# Run the full quality gate (same as CI)
ruff format --check imagen_mcp/ src/ tests/
ruff check imagen_mcp/ src/ tests/
mypy imagen_mcp/ src/
pytest --cov=src --cov-fail-under=80
# Verify server loads
python3 -c "from imagen_mcp.server import mcp; print('Server loads')"
python3 -m imagen_mcp
# Run over HTTP instead of stdio
IMAGEN_MCP_TRANSPORT=streamable-http imagen-mcp
# Check Claude Desktop logs (macOS)
tail -f ~/Library/Logs/Claude/mcp-server-imagen.logProject Structure
imagen-mcp/
├── imagen_mcp/ # Canonical package namespace + module runner
├── src/
│ ├── server.py # Implementation + legacy import path
│ ├── config/
│ │ ├── constants.py # Provider constants
│ │ └── settings.py # Environment configuration
│ ├── providers/
│ │ ├── base.py # Abstract provider interface
│ │ ├── openai_provider.py # OpenAI implementation
│ │ ├── gemini_provider.py # Gemini implementation
│ │ ├── selector.py # Auto-selection logic
│ │ └── registry.py # Provider factory
│ └── models/
│ └── input_models.py # Pydantic input models
├── tests/
│ ├── test_selector.py # Provider selection tests
│ ├── test_providers.py # Provider unit tests
│ └── test_server.py # Server integration tests
├── .github/
│ └── workflows/
│ └── ci.yml # GitHub Actions CI
├── run.sh # Wrapper script for MCP clients
├── requirements.txt
├── CLAUDE.md
└── README.mdRequirements
mcp>=1.26.0,<2
pydantic>=2.12.3
httpx>=0.28.0
google-genai>=2.8.0
pillow>=11.0.0License
Sources
Available Tools
8 toolsconversational_imageA
Generate images conversationally with iterative refinement.
USE THIS TOOL when:
User gives a vague/incomplete prompt that needs refinement
User wants iterative refinement across multiple messages
User explicitly asks for guidance or suggestions
Dialogue Modes:
"quick": 1-2 questions, fast path
"guided": 3-5 questions, balanced (DEFAULT)
"explorer": Deep exploration with 6+ questions
"skip": Direct generation, no dialogue
Provider Selection: Same auto-selection logic as generate_image. Provider is locked for the duration of a conversation (cannot switch mid-conversation).
Usage Pattern:
Initial: "A cozy coffee shop" → System asks refinement questions
User answers questions
Image generated with refined prompt
Refine: "Add more plants" (with same conversation_id)
Continue refining as needed
Args: params: Conversational image parameters including prompt and dialogue options.
Returns: Either dialogue questions or generated image with metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=false (mutation) and destructiveHint=false. The description adds behavioral traits beyond annotations: provider locking across conversations (cannot switch mid-conversation), dialogue modes impact on interaction depth, and the usage pattern for continuation via conversation_id. These details are useful but not exhaustive (e.g., no mention of output_path creation).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, when-to-use, dialogue modes, provider selection, usage pattern, returns). It is concise and front-loaded with essential information. The usage pattern could be slightly shorter, but overall it efficiently conveys the tool's workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (conversational image generation with many parameters and multi-turn refinement), the description covers core concepts, when to use, dialogue modes, provider locking, and continuation pattern. The output schema exists, so return details are not needed. It lacks discussion of some advanced parameters (e.g., reference_images, input_image_file_id), but schema handles those. It is sufficient for an AI agent to select and invoke.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (main description does not detail individual parameters), but the schema itself has thorough descriptions for all parameters. The description adds some context for 'dialogue_mode' (enum values) but largely repeats schema info. Baseline is 3 due to high schema coverage, and no significant extra meaning is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate images conversationally with iterative refinement.' It specifies the resource (images) and the verb (generate with conversation), and distinguishes from sibling tools like 'generate_image' by emphasizing conversation and iterative refinement. The dialogue modes further clarify the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly lists 'USE THIS TOOL when:' conditions, such as vague prompts, iterative refinement, or user asking for guidance. It does not directly state when not to use it (e.g., when prompt is clear and no refinement needed), but the context effectively implies when alternatives like 'generate_image' are better. It also covers dialogue modes and provider selection, offering solid guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageA
Edit an existing image using OpenAI gpt-image-2's /images/edits endpoint.
This is the right tool for:
Image-to-image refinement (OpenAI's answer to reference images)
Inpainting with a mask (paint over regions while preserving the rest)
Sequential/cumulative edits that preserve unchanged pixels
Brand-accurate modifications to existing images
Key features of gpt-image-2 editing:
input_fidelity='high'(default) keeps unchanged pixels constant — critical for multi-step refinement where each edit should build on the last without drift.Full control over quality, background, output_format, and compression.
Supports optional PNG mask (transparent pixels are the edit region).
Typical workflow:
Generate or obtain a base image (path on disk)
Call edit_image with prompt='change the sky to sunset'
Take the output path, call edit_image again with next instruction
Repeat — each step preserves pixels outside the described change
Args: params: Edit parameters including prompt, image_path, and options.
Returns: Formatted response with edited image path and metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses key behavioral traits: uses /images/edits endpoint, default high input fidelity to preserve pixels, mask support, sequential workflow. Annotations (readOnlyHint=false, destructiveHint=false) are consistent; description adds context about mutable but non-destructive behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with bullet points and sections (key features, typical workflow). No fluff; each sentence adds valuable information. Appropriate length given tool complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers use cases, workflow, and key features comprehensively. With an output schema present, return value explanation is unnecessary. The description is self-contained for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the tool has a single 'params' object with 0% description coverage at top level, the nested schema thoroughly documents all parameters. The description adds value by explaining key parameters like input_fidelity and mask_path in context, aiding interpretation beyond schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it edits existing images using OpenAI's gpt-image-2 endpoint. It lists specific use cases (image-to-image refinement, inpainting, sequential edits) that distinguish it from sibling tools like generate_image or conversational_image, which focus on generation or conversation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance under 'This is the right tool for:' with bullet points covering refinement, inpainting, sequential edits, and brand modifications. It implies not for from-scratch generation (handled by generate_image). The workflow description further clarifies when to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_costARead-onlyIdempotent
Estimate the cost of generating an image without generating it.
Runs the same provider auto-selection as generate_image (unless you
pin a provider) and looks up an approximate price from a local pricing
table. Useful for comparing providers/qualities before committing.
The figure is a ballpark — real cost depends on live provider pricing and, for OpenAI, actual image output tokens.
Args: params: Prompt plus optional provider/quality/size/n.
Returns: A formatted cost estimate.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint and idempotentHint. The description adds behavioral details: it runs the same provider auto-selection as generate_image, uses a local pricing table, and notes the estimate is a ballpark. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: starts with purpose, then explanation, caveats, and finally Args/Returns. It is slightly verbose but front-loaded with key information. A bit more conciseness could improve it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (cost estimation with no side effects), the description covers all essential aspects: what it does, how it works, caveats, and return format (via output schema). With annotations and output schema present, no gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description's Args section briefly lists parameters but adds little new meaning beyond the schema's own descriptions. The schema already provides detailed descriptions for each parameter. Schema coverage is 0% by context definition, but the schema itself is informative, so the description does not compensate for missing schema details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: estimating the cost of generating an image without actually generating it. It uses specific verb 'estimate' and resource 'cost', and distinguishes it from siblings like generate_image by explicitly noting it does not generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use: before generating, for comparing providers/qualities. It implies not to use for actual generation, but does not explicitly state when not to use or list alternatives beyond the tool itself. However, the sibling list provides context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageB
Generate an image using the best available provider.
Automatic Provider Selection: The server analyzes your prompt and automatically selects the best provider:
OpenAI GPT-Image-1 is auto-selected for:
Text-heavy images (menus, posters, infographics)
Comics with dialogue or speech bubbles
Technical diagrams with labels
Marketing materials requiring precise text
Gemini Nano Banana Pro is auto-selected for:
Photorealistic portraits and headshots
Product photography
High resolution (4K) output
Images using reference images for consistency
Real-time data visualization (weather, stocks)
Examples:
"Create a menu card for an Italian restaurant" → OpenAI (text rendering)
"Professional headshot with studio lighting" → Gemini (photorealism)
"Infographic explaining photosynthesis" → OpenAI (diagram + text)
"Product shot of perfume floating on water" → Gemini (product photography)
Override Selection:
Set provider to 'openai' or 'gemini' to override auto-selection.
Args: params: Image generation parameters including prompt and optional settings.
Returns: Formatted response with image path and metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations exist but are minimal (readOnlyHint=false, etc.). Description adds value by explaining auto-selection behavior and provider strengths. However, it does not disclose important behaviors like file saving (implied by output_path), potential latency, or cost implications. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with headings, lists, and examples. It is front-loaded with the main purpose. However, it is somewhat lengthy with provider comparisons that could be condensed. Overall, sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (multiple providers, many parameters, auto-selection), the description explains the selection logic and provides examples. It mentions override and return format. An output schema exists to handle return details. It is fairly complete for an agent to understand the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage for the top-level param 'params' is 0% (no description), and the description only repeats 'Image generation parameters including prompt and optional settings' – adding no new meaning. While nested properties have descriptions in the schema, the description fails to compensate for the top-level lack of detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates an image using the best available provider, with specific verb 'Generate' and resource 'image'. It distinguishes from siblings by focusing on single image generation with auto-selection, while siblings like 'edit_image' and 'conversational_image' imply different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides detailed guidance on when to use each provider within the tool, but offers no guidance on when to choose this tool over sibling tools like 'edit_image' or 'generate_image_batch'. An agent would need to infer from the tool name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_image_batchA
Generate many images concurrently from a list of prompts.
Each item runs through the same auto provider selection as generate_image,
bounded by max_concurrency. Per-item failures are isolated — one bad
prompt does not fail the whole batch. Returns every result (saved paths plus
any per-item errors).
Use this instead of calling generate_image in a loop: 8 prompts that
would take ~4 minutes serially complete in roughly one generation's time
(subject to max_concurrency and provider rate limits).
Args: params: The batch (items + concurrency + optional default provider).
Returns: A formatted summary of all results.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate non-readonly, non-destructive, non-idempotent, open-world. Description adds concurrency bounds, failure isolation, rate limit dependency, and result format, providing substantial behavioral context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with summary, args, and returns. Every sentence adds value; no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers inputs, behavior, concurrency, error handling, and return format. With output schema present, it provides complete guidance for a batch image generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Top-level 'params' description is brief, but the nested schema for `BatchGenerationInput` has detailed field descriptions. Tool description adds context about batch structure and default provider, complementing the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool generates many images concurrently from a list of prompts. Distinguishes from siblings like `generate_image` by mentioning batch processing and concurrency.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises using this tool instead of calling `generate_image` in a loop, with a concrete performance example. Also describes per-item failure isolation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_conversationsARead-onlyIdempotent
List saved image generation conversations.
Returns recent conversations that can be continued for refinement. Each conversation tracks the provider used and generation history.
Args: params: Options for filtering and formatting the list.
Returns: List of conversations with metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read and idempotent. The description adds that results are 'recent' and can be 'continued for refinement', which provides some behavioral context beyond the annotations. However, it doesn't detail behavior like empty results, ordering, or response structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: four sentences total, with the core purpose in the first sentence. No extraneous information. All sentences add value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool, the description covers purpose and key traits (recent, continuable, tracks provider/history). It lacks details on pagination behavior and ordering, and the output format is only implied via the parameter. Still, it is mostly sufficient given the annotations and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description only says 'params: Options for filtering and formatting the list.' This is vague and does not add meaning beyond the input schema, which already describes each parameter. With 0% schema description coverage, the description fails to compensate, but the schema itself is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists saved image generation conversations, using specific verb 'list' and resource 'conversations'. It distinguishes from siblings as no other list tool exists. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to retrieve recent conversations for refinement, but provides no explicit guidance on when to use it versus other tools or when not to use it. No alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_gemini_modelsARead-onlyIdempotent
List available Gemini models that support image generation.
Queries the Gemini API to show which models are available for image generation with your API key. Useful for troubleshooting or choosing alternative models.
Returns: List of available Gemini image models with their capabilities.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds that it queries the Gemini API and returns capabilities, but does not provide additional behavioral details beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, with a clear purpose, action, return description, and use case. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero parameters and presence of an output schema, the description fully covers the tool's purpose and typical use case without needing further elaboration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the baseline is 4. The description does not need to add parameter semantics, and it does not attempt to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists available Gemini models that support image generation. This verb-noun pair is specific and distinguishes it from siblings like 'generate_image' or 'list_providers'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions it is useful for troubleshooting or choosing alternative models, providing clear context. It does not explicitly state when not to use it, but the context of sibling tools offers implied guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersARead-onlyIdempotent
List available image generation providers and their capabilities.
Returns a comparison of available providers including:
Which providers have API keys configured
Best use cases for each provider
Feature comparison (text rendering, resolution, etc.)
Use this to understand which provider to choose for your task.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, but the description adds value by detailing the return content: API key configuration, best use cases, and feature comparison. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with a clear lead sentence and bullet points for output details. Every sentence adds value, and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description adequately explains the tool's purpose and what the output includes, even without needing to detail parameters. Given the presence of an output schema, the description is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, and schema coverage is 100%. With zero parameters, the baseline is 4, and the description does not need to add parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'List available image generation providers and their capabilities', clearly identifying the verb (list) and resource (providers with capabilities). This distinguishes it from sibling tools like generate_image or edit_image, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises 'Use this to understand which provider to choose for your task', providing clear context for when to use this tool. While it doesn't list explicit alternatives or when-not-to-use, the context of sibling tools makes the usage well implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.3.0- First observed
conversational_image - First observed
edit_image - First observed
estimate_cost - First observed
generate_image - First observed
generate_image_batch - First observed
list_conversations - First observed
list_gemini_models - First observed
list_providers
TDQS
Scored across 8 tools
Each tool has a clearly distinct purpose: direct generation, conversational refinement, editing, batch generation, cost estimation, and listing of conversations/models/providers. Overlaps are minimal and resolved by detailed descriptions.
Most tools follow verb_noun snake_case (e.g., edit_image, generate_image), but 'conversational_image' uses an adjective instead of a verb, creating a slight inconsistency.
With 8 tools covering generation, editing, batch processing, cost estimation, and listing functions, the count is well-scoped for an image generation server without being too many or too few.
The tool surface covers core generation, editing, batch, and estimation needs. Minor gaps exist (e.g., no deletion tool or detailed image metadata viewer), but the essential workflows are supported.
Maintenance
Related MCP Connectors
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
AI visual generation agent: multi-pipeline rendering, prompt crafting, and image composition.
Image, video, music and text generation across 100+ models through one endpoint.
Generate images with any major model — one API key, one prepaid balance, one MCP.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI-powered image analysis using OpenAI's Vision API and image generation with DALL-E models. Supports image description, content analysis, comparison, editing, and creating variations with intelligent caching.3MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI image generation through multiple providers including OpenAI GPT-Image-1, Google Imagen 4, Gemini 2.5 Flash (Nano Banana), Flux 1.1, Qwen Image, and SeedDream-4, supporting various formats, sizes, and advanced features like background control and seed-based reproduction.19711MIT
- AlicenseNot gradedqualityDmaintenanceEnables conversational image generation, editing, and refinement through OpenAI models with session memory for iterative creative workflows.1MIT
- FlicenseNot gradedqualityDmaintenanceEnables image generation, editing, and refinement using Google's Gemini 2.5 Flash Image model with support for multi-image composition and style transfer.-