Gemini Image Generator MCP Server
Gemini 图像生成器 MCP 服务器
通过 MCP 协议使用 Google 的 Gemini 模型从文本提示生成高质量的图像。
概述
这款 MCP 服务器允许任何 AI 助手使用 Google 的 Gemini AI 模型生成图像。该服务器负责处理快速工程、文本转图像、文件名生成以及本地图像存储,让您能够通过任何 MCP 客户端轻松创建和管理 AI 生成的图像。
Related MCP server: Gemini Image Gen MCP Server
特征
使用 Gemini 2.0 Flash 进行文本到图像的生成
基于文本提示的图像到图像转换
支持基于文件和 base64 编码的图像
根据提示自动智能生成文件名
非英语提示的自动翻译
可配置输出路径的本地图像存储
生成的图像中严格排除文本
高分辨率图像输出
直接访问图像数据和文件路径
可用的 MCP 工具
该服务器为AI助手提供了以下MCP工具:
1. generate_image_from_text
根据文本提示描述创建新图像。
generate_image_from_text(prompt: str) -> Tuple[bytes, str]参数:
prompt:要生成的图像的文本描述
返回:
包含以下内容的元组:
原始图像数据(字节)
保存的图像文件的路径(str)
这种双重返回格式允许 AI 助手直接处理图像数据或引用保存的文件路径。
例子:
“生成山上日落的图像”
“在科幻城市中创造一只逼真的飞猪”
示例输出
该图像是使用提示生成的:
"Hi, can you create a 3d rendered image of a pig with wings and a top hat flying over a happy futuristic scifi city with lots of greenery?"
一只戴着高顶礼帽、长着翅膀的 3D 渲染猪,飞过一座充满绿意的未来科幻城市
已知问题
将此 MCP 服务器与 Claude Desktop Host 一起使用时:
性能问题:与其他方法相比,使用
transform_image_from_encoded处理时间可能会显著延长。这是由于通过 MCP 协议传输大量 base64 编码的图像数据会产生开销。路径解析问题:使用 Claude Desktop Host 时,可能无法正确解析图像路径。主机应用程序可能无法正确解释返回的文件路径,从而导致难以访问生成的图像。
为了获得最佳体验,请考虑在可能的情况下使用替代 MCP 客户端或transform_image_from_file方法。
2. transform_image_from_encoded
使用 base64 编码的图像数据根据文本提示转换现有图像。
transform_image_from_encoded(encoded_image: str, prompt: str) -> Tuple[bytes, str]参数:
encoded_image:带有格式标头的 Base64 编码图像数据(必须采用以下格式:“data:image/[format];base64,[data]”)prompt:关于如何转换图像的文本描述
返回:
包含以下内容的元组:
原始转换图像数据(字节)
已保存的转换图像文件的路径(str)
例子:
“给这片风景添加雪景”
“将背景改为海滩”
3. transform_image_from_file
根据文本提示转换现有的图像文件。
transform_image_from_file(image_file_path: str, prompt: str) -> Tuple[bytes, str]参数:
image_file_path:要转换的图像文件的路径prompt:关于如何转换图像的文本描述
返回:
包含以下内容的元组:
原始转换图像数据(字节)
已保存的转换图像文件的路径(str)
例子:
“在此图像中的人物旁边添加一只骆驼”
“让白天的场景看起来像夜晚”
示例转换
使用上面创建的飞猪图像,我们根据以下提示应用了变换:
"Add a cute baby whale flying alongside the pig"前:
后:
原始的飞猪图像加上一只可爱的小鲸鱼在它旁边飞翔
设置
先决条件
Python 3.11+
Google AI API 密钥(Gemini)
MCP 主机应用程序(Claude Desktop App、Cursor 或其他 MCP 兼容客户端)
获取 Gemini API 密钥
使用您的 Google 帐户登录
点击“创建 API 密钥”
复制新的 API 密钥以用于配置
注意:API 密钥每月提供一定额度的免费使用。您可以在 Google AI Studio 中查看使用情况。
安装
克隆存储库:
git clone https://github.com/your-username/gemini-image-generator.git
cd gemini-image-generator创建虚拟环境并安装依赖项:
# Using regular venv
python -m venv .venv
source .venv/bin/activate
pip install -e .
# Or using uv
uv venv
source .venv/bin/activate
uv pip install -e .复制示例环境文件并添加您的 API 密钥:
cp .env.example .env编辑
.env文件以包含您的 Google Gemini API 密钥和首选输出路径:
GEMINI_API_KEY="your-gemini-api-key-here"
OUTPUT_IMAGE_PATH="/path/to/save/images"配置 Claude 桌面
将以下内容添加到您的claude_desktop_config.json中:
macOS :
~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"gemini-image-generator": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/gemini-image-generator",
"run",
"server.py"
],
"env": {
"GEMINI_API_KEY": "GEMINI_API_KEY",
"OUTPUT_IMAGE_PATH": "OUTPUT_IMAGE_PATH"
}
}
}
}用法
安装并配置完成后,您可以要求 Claude 使用以下提示生成或转换图像:
生成新图像
“生成山上日落的图像”
“创作一幅未来城市景观的插画”
“画一张戴着太阳镜的猫的照片”
转换现有图像
“通过在场景中添加雪来改变这幅图像”
“编辑这张照片,让它看起来像是在晚上拍摄的”
“在这张图片的背景中添加一条飞翔的龙”
生成/转换后的图像将保存到您配置的输出路径,并在 Claude 中显示。通过更新的返回类型,AI 助手还可以直接处理图像数据,而无需访问已保存的文件。
测试
您可以通过运行 FastMCP 开发服务器来测试该应用程序:
fastmcp dev server.py此命令启动本地开发服务器,并通过http://localhost:5173/访问 MCP Inspector。MCP Inspector 提供了一个便捷的 Web 界面,您可以直接在其中测试图像生成工具,而无需使用 Claude 或其他 MCP 客户端。您可以输入文本提示,执行工具并立即查看结果,这对于开发和调试非常有帮助。
执照
MIT 许可证
Available Tools
3 toolsgenerate_image_from_textA
Generate an image based on the given text prompt using Google's Gemini model.
Args:
prompt: User's text prompt describing the desired image to generate
Returns:
Path to the generated image file using Gemini's image generation capabilities
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the model and return type (path to image file) but lacks critical details such as rate limits, authentication requirements, image format, size, quality, or error handling. This is insufficient for a generative AI tool with potential costs and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core functionality. The structured sections (Args, Returns) enhance readability, though the second sentence could be more integrated to avoid slight redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of image generation, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits (e.g., costs, latency), output specifics (e.g., file format, resolution), and error cases, leaving significant gaps for an AI agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, but the description compensates by explaining the single parameter ('prompt') as 'User's text prompt describing the desired image to generate.' This adds meaningful context beyond the schema's basic type information, clarifying the parameter's role in the generation process.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate an image') and resource ('based on the given text prompt'), using Google's Gemini model. It distinguishes from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' by specifying text-based generation rather than transformation from existing images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for text-to-image generation but does not explicitly state when to use this tool versus alternatives. It mentions the model (Gemini) but provides no guidance on prerequisites, limitations, or scenarios where other tools might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_encodedA
Transform an existing image based on the given text prompt using Google's Gemini model.
Args:
encoded_image: Base64 encoded image data with header. Must be in format:
"data:image/[format];base64,[data]"
Where [format] can be: png, jpeg, jpg, gif, webp, etc.
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| encoded_image | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the tool uses Google's Gemini model and that it saves the transformed image on the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or potential side effects of the transformation process.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear opening sentence stating the purpose, followed by well-organized sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no annotations and no output schema, the description provides good coverage of purpose, parameters, and basic behavior. It explains what the tool does, how to format inputs, and what to expect as output. The main gap is lack of information about error conditions, performance characteristics, or more detailed behavioral constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by providing detailed semantics for both parameters. It specifies the exact format required for encoded_image (including header format and supported image types) and explains what the prompt parameter should contain. This adds significant value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verb ('Transform') and resource ('an existing image'), and distinguishes it from siblings by specifying it uses encoded image data rather than text or file inputs. The mention of Google's Gemini model adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool (transforming existing images with encoded data) and implicitly distinguishes it from siblings (generate_image_from_text for text-to-image, transform_image_from_file for file-based transformation). However, it doesn't explicitly state when NOT to use this tool or mention specific prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_fileA
Transform an existing image file based on the given text prompt using Google's Gemini model.
Args:
image_file_path: Path to the image file to be transformed
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| image_file_path | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the tool saves the transformed file on the server, which is useful behavioral context. However, it lacks critical details like required permissions, file format limitations, transformation scope, error handling, or whether the operation is reversible/destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear purpose statement followed by labeled sections for Args and Returns. Every sentence adds value without redundancy, and information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 2 parameters, the description covers purpose and parameters adequately. However, for a transformation tool with potential complexity (image processing via Gemini), it lacks details about output format, file location specifics, or error cases, leaving gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining both parameters: 'image_file_path' as 'Path to the image file to be transformed' and 'prompt' as 'Text prompt describing the desired transformation or modifications'. This adds essential meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verb ('Transform') and resource ('existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The mention of using Google's Gemini model adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying it transforms 'an existing image file' and uses a 'text prompt', which differentiates it from 'generate_image_from_text' (creates new images) and 'transform_image_from_encoded' (uses encoded input). However, it doesn't explicitly state when to choose this tool over alternatives or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
generate_image_from_text - First observed
transform_image_from_encoded - First observed
transform_image_from_file
TDQS
Scored across 3 tools
The three tools have clearly distinct purposes: generate_image_from_text creates new images from text prompts, while transform_image_from_encoded and transform_image_from_file both transform existing images but differ in input format (base64 encoded vs. file path). The descriptions make these distinctions explicit, eliminating any potential confusion between generation and transformation operations.
All tools follow a consistent verb_noun_from_source naming pattern: generate_image_from_text, transform_image_from_encoded, and transform_image_from_file. This pattern clearly indicates the action (generate/transform), the target (image), and the input source (text/encoded/file), creating a predictable and readable naming convention throughout the toolset.
Three tools is a reasonable count for an image generation server, covering the core operations of generating new images and transforming existing ones. However, the scope feels slightly thin as there are no complementary tools for managing generated images (like listing, deleting, or retrieving metadata), which might limit agent workflows in production scenarios.
The server covers basic image generation and transformation operations well, but has notable gaps in image management. There are no tools for listing generated images, deleting files, retrieving image metadata, or batch operations. While the core generative AI functionality is present, the lack of lifecycle management tools creates potential dead ends for agents working with multiple images over time.
Maintenance
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Generate images, video, audio and short films with 140+ AI models from any MCP client.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Related MCP Servers
- FlicenseBqualityDmaintenanceEnables image generation and multi-turn editing sessions using the Gemini API within MCP-compatible environments. Users can create, modify, and configure images through natural language commands, supporting features like aspect ratio adjustments and session-based image transformations.5-
- AlicenseAqualityDmaintenanceEnables AI image generation, editing, and upscaling via Google Gemini and Imagen models, supporting dynamic model switching and multiple MCP-compatible clients.12MIT
- FlicenseNot gradedqualityDmaintenanceProvides image generation capabilities using Google's Gemini 2.0 Flash Preview model through the MCP protocol, enabling AI assistants to generate high-quality images from text prompts.-
- FlicenseNot gradedqualityDmaintenanceEnables AI-powered image generation using Google's Gemini 2.5 Flash Image Preview model, supporting text-to-image and image-to-image generation through the MCP interface.-