Gemini Image Generator MCP Server
Generates and transforms images using Google's Gemini AI models, supporting text-to-image generation and image editing with natural language prompts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gemini Image Generator MCP Servergenerate a photorealistic cat on a windowsill"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.

Gemini Image Generator MCP Server
Generate and transform images using Google's Gemini AI through the Model Context Protocol (MCP).
Features
Text-to-Image Generation - Create images from natural language prompts
Image Transformation - Modify existing images with text descriptions
Automatic Filename Generation - Smart naming based on prompts
Multi-Language Support - Automatic prompt translation to English
Related MCP server: Gemini Image MCP
Installation
Get a free API key from Google AI Studio.
git clone https://github.com/jonchun/gemini-image-mcp.git
cd gemini-image-mcp
# Using uv (recommended)
uv venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
uv pip install -e .Configuration
Add to claude_desktop_config.json:
{
"mcpServers": {
"gemini-image-mcp": {
"command": "uv",
"args": [
"--directory",
"/ABSOLUTE/PATH/TO/gemini-image-mcp",
"run",
"gemini-image-mcp"
],
"env": {
"GEMINI_API_KEY": "your-api-key-here",
"DEFAULT_OUTPUT_IMAGE_PATH": "/path/to/images"
}
}
}
}{
"gemini-image-mcp": {
"command": "uv",
"args": [
"--directory",
"/ABSOLUTE/PATH/TO/gemini-image-mcp",
"run",
"gemini-image-mcp"
],
"env": {
"GEMINI_API_KEY": "your-api-key-here",
"DEFAULT_OUTPUT_IMAGE_PATH": "/path/to/images"
}
}
}Install from smithery.ai - search for "gemini-image-mcp".
Usage
Generate Images
Generate a photorealistic sunset over mountains with purple sky
Create a British Shorthair silver tabby kitten playing with a ball of yarn
Transform Images
Prompt: Add beautiful vibrant aurora borealis (northern lights) dancing across the sky with green, purple, and blue colors

Prompt: Add soft natural sunlight streaming through a window, creating beautiful warm light rays and gentle shadows

Available Tools
generate_image_from_text
Creates an image from a text description.
Parameters:
prompt(required): Text description of the imageoutput_dir(optional): Directory to save the imagemodel(optional): Gemini model to use (defaults toGEMINI_MODELenvironment variable)
Returns: Path to the saved image file
transform_image_from_file
Transforms an existing image based on a text prompt.
Parameters:
image_file_path(required): Path to the source imageprompt(required): Description of the transformationoutput_dir(optional): Directory to save the imagemodel(optional): Gemini model to use (defaults toGEMINI_MODELenvironment variable)
Returns: Path to the transformed image file
transform_image_from_encoded
Transforms a base64-encoded image.
Parameters:
encoded_image(required): Base64 data URL (data:image/[format];base64,[data])prompt(required): Description of the transformationoutput_dir(optional): Directory to save the imagemodel(optional): Gemini model to use (defaults toGEMINI_MODELenvironment variable)
Returns: Path to the transformed image file
Configuration
Variable | Required | Default | Description |
| Yes | - | Your Gemini API key |
| No | Current directory | Default save location |
| No |
| Model to use |
| No |
| API base URL |
Development
Test the server locally:
fastmcp dev src/gemini_image_mcp/server.pyOpens MCP Inspector at http://localhost:5173/
License
MIT
Available Tools
3 toolsgenerate_image_from_textA
Generate an image from a text prompt using Gemini.
Args: prompt: Text description of the desired image. output_dir: Optional directory to save the generated image. If not provided, the image is only returned in the response (not saved to disk). model: Optional Gemini model name. If not provided, uses GEMINI_MODEL environment variable. ctx: Optional context for progress reporting.
Returns: List containing ImageContent with the generated image, and optionally TextContent with file path if saved.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| output_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses that the image is returned in the response and only saved to disk if output_dir is given, and that model falls back to the GEMINI_MODEL environment variable. It stops short of covering auth requirements, rate limits, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a one-line purpose, then organized Args/Returns sections. Efficient overall, though the Returns block partially duplicates the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
All three parameters are explained, the optional persistence behavior is confirmed, and an output schema exists so return details need not be restated. Missing only selection guidance relative to the transform siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters and it largely does: prompt, output_dir (with its disk-vs-response consequence), and model (with its env-var fallback) each get meaningful explanation beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Generate an image from a text prompt using Gemini.' This clearly distinguishes it from the sibling transform_image_* tools, which transform existing images rather than generating new ones, though it never explicitly names those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is offered. The agent must infer that this is the generation path versus the transformation siblings purely from the verb. The output_dir default behavior is described, but that is parameter semantics rather than selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_encodedA
Transform a base64-encoded image using Gemini.
Args: encoded_image: Base64 data URL (data:image/[format];base64,[data]). prompt: Text description of desired transformation. output_dir: Optional directory to save the generated image. If not provided, the image is only returned in the response (not saved to disk). model: Optional Gemini model name. If not provided, uses GEMINI_MODEL environment variable. ctx: Optional context for progress reporting.
Returns: List containing ImageContent with the transformed image, and optionally TextContent with file path if saved.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| output_dir | No | ||
| encoded_image | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose some behavior: the output is returned in the response and optionally saved to disk depending on output_dir, and it reveals the GEMINI_MODEL env fallback for the model. However, it omits whether the call consumes Gemini quota, what happens on invalid base64, image size/format limits, and whether disk writes create or overwrite files. The return-format disclosure partially compensates but key mutation/side-effect info is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Args/Returns structure is front-loaded with the purpose sentence, and each parameter line adds real information. It is somewhat verbose due to the Args/Returns scaffolding, and the model/ctx lines could be tighter, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not detail return values, though it still helpfully notes the return shape. For a 4-param tool with no annotations it covers purpose and all parameters well, but it leaves out API-key/auth requirements and failure/limit behavior that an agent might need to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate, and it does: it documents all four parameters. It specifies encoded_image as a data URL with the exact format, prompt as the desired transformation text, output_dir's conditional save behavior when omitted, and model's GEMINI_MODEL env fallback. This is exactly the compensation the low schema coverage demands.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Transform a base64-encoded image using Gemini.' It clearly distinguishes from the file-based sibling 'transform_image_from_file' via 'from_encoded'/base64 and from 'generate_image_from_text' by being a transformation of an existing image rather than generation. It doesn't explicitly name the siblings, but the base64 encoding distinction is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the 'from_encoded' scope (use it with base64 data URLs rather than local files), but the description never states when to prefer this over transform_image_from_file or generate_image_from_text. The data URL format note gives partial context. No explicit when-not or alternative tool routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_fileA
Transform an image file using Gemini.
Args: image_file_path: Path to the source image file. prompt: Text description of desired transformation. output_dir: Optional directory to save the generated image. If not provided, the image is only returned in the response (not saved to disk). model: Optional Gemini model name. If not provided, uses GEMINI_MODEL environment variable. ctx: Optional context for progress reporting.
Returns: List containing ImageContent with the transformed image, and optionally TextContent with file path if saved.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| prompt | Yes | ||
| output_dir | No | ||
| image_file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: the image is only returned in the response unless output_dir is supplied, and the model falls back to the GEMINI_MODEL environment variable. It omits error behavior, cost/latency, and whether an existing file at the target path is overwritten, but this is solid disclosure for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The docstring is front-loaded with the purpose in one line, then structured Args/Returns sections. Slightly verbose in formatting but every parameter line carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations and 0% schema coverage, the description covers inputs and return shape; an output schema exists so return details are partly redundant but helpful. The main gaps, sibling routing guidance and error/format handling, are the only omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it documents all four parameters with semantics: image_file_path as source path, prompt as transformation text, output_dir's save-vs-return consequence, and model's env-var fallback. It does not specify accepted image formats, path constraints, or whether output_dir must exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Transform an image file using Gemini') that clearly distinguishes it from transform_image_from_encoded (which takes encoded input) and generate_image_from_text (which creates rather than transforms). Differentiation comes implicitly from the name and the input-concept rather than an explicit sibling callout, so it stops just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over its siblings, e.g. 'use this when you have a path on disk; use transform_image_from_encoded when you have base64 data.' The agent must infer the routing decision entirely from the title.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
generate_image_from_text - First observed
transform_image_from_encoded - First observed
transform_image_from_file
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: text-to-image generation versus image transformation from two different input sources (encoded data or file path). The input types prevent confusion, and descriptions explicitly state the expected arguments.
All tool names follow the same snake_case verb_noun_preposition pattern (generate_image_from_text, transform_image_from_encoded, transform_image_from_file). The convention is consistent and predictable.
Three tools are well-scoped for an image generation and transformation server; each tool covers a distinct input method without redundancy. No tool feels superfluous or missing from a minimal set.
The server covers the core lifecycle: generating an image from text and transforming existing images from both encoded data and file paths. Minor gaps exist (e.g., no batch generation or targeted editing), but the primary workflows are supported.
Maintenance
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Generate images with any major model — one API key, one prepaid balance, one MCP.
Edit images over MCP with object removal, background removal, and guided generative edits.
Related MCP Servers
- AlicenseAqualityDmaintenanceAllows AI assistants to generate and transform high-quality images from text prompts using Google's Gemini model via the MCP protocol.334MIT
- FlicenseBqualityDmaintenanceEnables image generation and multi-turn editing sessions using the Gemini API within MCP-compatible environments. Users can create, modify, and configure images through natural language commands, supporting features like aspect ratio adjustments and session-based image transformations.5-
- AlicenseAqualityDmaintenanceMCP server for Google Gemini image generation with configurable model support, enabling text-to-image generation, image editing, and iterative refinement.640MIT
- FlicenseNot gradedqualityDmaintenanceProvides image generation capabilities using Google's Gemini 2.0 Flash Preview model through the MCP protocol, enabling AI assistants to generate high-quality images from text prompts.-