Gemini Image Generator MCP
This server enables image generation and transformation using Google's Gemini AI, integrated with Claude Desktop or other MCP-compatible clients.
Text-to-image generation: Create new images from text prompts using Gemini 2.0 Flash
Image transformation from file: Modify existing images by providing a local file path and a text prompt describing the desired changes
Image transformation from Base64 data: Edit images provided as Base64-encoded data (e.g.,
data:image/png;base64,...) using a transformation promptMultilingual support: Automatically translates and optimizes Japanese, Korean, and Chinese prompts into English for better results
AI-generated filenames: Automatically generates meaningful filenames based on the prompt used
Local storage: All generated/transformed images are automatically saved to a designated local folder
In-chat display: View generated or transformed images directly within the Claude Desktop chat interface
Enables high-quality image generation and editing using Google's Gemini 2.0 Flash model, supporting text-to-image generation, image transformation with text prompts, and multi-language prompt optimization
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gemini Image Generator MCPcreate a cute cat illustration with a hat"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini 画像生成器 MCP サーバー (修正版)
Claude DesktopでGoogle の Gemini AIを使用して高品質な画像を生成・編集できる MCP サーバーです。
🚀 主な特徴
テキストから画像生成: Gemini 2.0 Flash を使用したテキスト→画像変換
画像変換: 既存の画像をテキストプロンプトで修正
多言語対応: 日本語・韓国語・中国語プロンプトの自動英語翻訳・最適化
AI ファイル名生成: プロンプト基準でファイル名を自動生成
ローカル保存: 生成された画像を指定フォルダに自動保存
Claude チャット内表示: 生成された画像をチャット画面で直接確認
Related MCP server: NanoBanana MCP
🛠️ インストール要件
Python 3.11 以上
Google Gemini API キー
Claude Desktop またはその他 MCP 互換クライアント
📋 ステップ1: Gemini API キー発行
Google アカウントでログイン
"Create API Key" をクリック
生成された API キーをコピー(後で使用)
💾 ステップ2: MCP サーバーインストール
自動インストール(推奨)
# リポジトリクローン
git clone https://github.com/sungmin-koo-ai/GeminiImageMCP.git
cd GeminiImageMCP
# 仮想環境作成・有効化
python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# パッケージインストール
pip install -e .インストール確認
# サーバーが正常実行されるかテスト
python -m gemini_image_mcp.serverStarting Gemini Image Generator MCP server... メッセージが表示されれば成功!(Ctrl+CまたはCtrl+Z で終了)
⚙️ ステップ3: Claude Desktop 設定
設定ファイル場所
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
設定ファイル内容
{
"mcpServers": {
"gemini-image-generator": {
"command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/python",
"args": [
"-m", "gemini_image_mcp.server"
],
"env": {
"GEMINI_API_KEY": "ここに実際のAPIキーを入力",
"OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
}
}
}
}実際の設定例
{
"mcpServers": {
"gemini-image-generator": {
"command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/python",
"args": [
"-m", "gemini_image_mcp.server"
],
"env": {
"GEMINI_API_KEY": "AIzaSy...(実際のAPIキー)",
"OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
}
}
}
}スクリプト命令を使用する場合(簡単設定)
{
"mcpServers": {
"gemini-image-generator": {
"command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/gemini-image-mcp",
"env": {
"GEMINI_API_KEY": "AIzaSy...(実際のAPIキー)",
"OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
}
}
}
}※ この方法は args 指定が不要でより簡潔です
🚨 重要事項
絶対パス使用: すべてのパスは完全パスで入力
API キー置換:
ここに実際のAPIキーを入力部分を発行した実際のキーに置換画像フォルダ:
OUTPUT_IMAGE_PATHに指定したフォルダが事前に作成されている必要があります
画像保存フォルダ作成
mkdir -p ~/Pictures/ai_generated🎯 ステップ4: 実行・テスト
Claude Desktop 再起動: 設定後完全に終了して再起動
接続確認: Claude Desktop で MCP サーバーが接続されたか確認
テスト: 「猫の絵を描いて」とリクエストしてみる
📖 使用方法
画像生成
東京タワーの可愛いイラストを描いて画像変換(ファイルパス)
/Users/username/image.jpg この画像に虹を追加して画像変換(アップロード)
画像を Claude にアップロード後:
背景をレインボーブリッジの夜景にしてくれ🔧 トラブルシューティング
サーバー接続失敗
ログ確認: Claude Desktop のログフォルダで
gemini-image-generator.logを確認パス確認:
claude_desktop_config.jsonの Python パスが正確か確認権限確認: 画像保存フォルダに書き込み権限があるか確認
API キーエラー
キー有効性: Google AI Studio で API キーが有効化されているか確認
引用符確認: 設定ファイルで API キーが引用符で囲まれているか確認
手動テスト
cd ~/GeminiImageMCP
source venv/bin/activate
export GEMINI_API_KEY="実際のAPIキー"
export OUTPUT_IMAGE_PATH="~/Pictures/ai_generated"
python -m gemini_image_mcp.server📊 提供ツール
1. generate_image_from_text
機能: テキストプロンプトで新しい画像生成
入力: 画像説明テキスト
出力: 生成された画像(Claude チャット内表示 + ローカル保存)
2. transform_image_from_file
機能: ファイルパスの画像をテキストプロンプトで変換
入力: 画像ファイルパス、変換プロンプト
出力: 変換された画像(Claude チャット内表示 + ローカル保存)
3. transform_image_from_encoded
機能: Base64 エンコードされた画像をテキストプロンプトで変換
入力: Base64 画像データ、変換プロンプト
出力: 変換された画像(Claude チャット内表示 + ローカル保存)
📝 オリジナルからの相違点
この修正版は元のリポジトリの以下の問題を解決しました:
❌ 元の問題: JSON シリアル化エラー (
invalid utf-8 sequence)❌ 元の問題: MCP ツールがバイナリデータ返却により実行失敗
✅ 修正事項: ファイルパス返却で安定的な動作
✅ 修正事項: Claude Desktop で完璧に動作
✅ 修正事項: 生成された画像を Claude チャット内で直接確認可能
🤝 貢献・お問い合わせ
問題報告: GitHub Issues タブで問題を報告
📄 ライセンス
MIT License - 元のプロジェクトと同じ
ヒント: 初回設定時はステップごとに進め、問題が発生した場合はまずログファイルを確認してください! 🚀
Available Tools
3 toolsgenerate_image_from_textA
Generate an image based on the given text prompt using Google's Gemini model.
Args:
prompt: User's text prompt describing the desired image to generate
Returns:
Path to the generated image file using Gemini's image generation capabilities
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the action ('Generate') and output ('Path to the generated image file'), but does not cover critical aspects like rate limits, authentication needs, file formats, error handling, or whether the operation is idempotent, leaving significant gaps for a generative tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections for Args and Returns, and each sentence adds value. It is appropriately sized for the tool's complexity, though it could be slightly more concise by integrating the technology mention into the main purpose statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (generative AI operation), no annotations, and no output schema, the description provides basic purpose and parameter info but lacks details on behavioral traits, error cases, or output specifics beyond a path. It is minimally viable but has clear gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful context for the single parameter 'prompt' by explaining it as 'User's text prompt describing the desired image to generate', which goes beyond the schema's minimal title. With 0% schema description coverage and only one parameter, this adequately compensates, though it could include examples or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate an image'), resource ('based on the given text prompt'), and technology ('using Google's Gemini model'), distinguishing it from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' which involve transformation rather than generation from text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for image generation from text prompts but does not explicitly state when to use this tool versus its siblings. It mentions the technology (Gemini model) which provides some context, but lacks explicit guidance on alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_encodedA
Transform an existing image based on the given text prompt using Google's Gemini model.
Args:
encoded_image: Base64 encoded image data with header. Must be in format:
"data:image/[format];base64,[data]"
Where [format] can be: png, jpeg, jpg, gif, webp, etc.
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| encoded_image | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses that the tool uses Google's Gemini model and saves the transformed image to the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or error conditions that would be important for a transformation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear purpose statement followed by well-organized Args and Returns sections. Every sentence earns its place by providing essential information without redundancy. The formatting with clear section headers enhances readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a transformation tool with no annotations and no output schema, the description provides good parameter documentation and purpose clarity. However, it lacks information about the transformation process (e.g., quality, limitations, processing time), error handling, and more detailed behavioral context that would be valuable given the complexity of image transformation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by providing detailed semantic information for both parameters. It specifies the exact format required for encoded_image (including header structure and supported formats) and explains that prompt describes 'desired transformation or modifications.' This adds significant value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Transform an existing image based on the given text prompt using Google's Gemini model.' It specifies the verb (transform), resource (existing image), method (Gemini model), and distinguishes from sibling tools (transform_image_from_file handles file input instead of encoded data).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly suggests usage context by mentioning 'existing image' and the encoding format, but doesn't explicitly state when to use this tool versus alternatives like transform_image_from_file or generate_image_from_text. It provides technical prerequisites but lacks comparative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_fileA
Transform an existing image file based on the given text prompt using Google's Gemini model.
Args:
image_file_path: Path to the image file to be transformed
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| image_file_path | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that the tool saves the transformed file on the server, which is useful context beyond basic functionality. However, it lacks details on permissions, rate limits, error handling, or what transformations are possible, leaving behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with the core purpose, followed by clear sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of an AI-based image transformation tool with no annotations and no output schema, the description is moderately complete. It covers the basic operation and return value but lacks details on supported image formats, transformation limits, or error cases, which are important for such a tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant meaning beyond the input schema, which has 0% coverage. It explains that 'image_file_path' is for an existing image file to be transformed and 'prompt' describes the desired transformation, clarifying the purpose and usage of both parameters effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Transform') and resource ('an existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The phrase 'using Google's Gemini model' adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly provides usage context by mentioning 'existing image file' and the Gemini model, which suggests when to use this tool. However, it doesn't explicitly state when to choose this over sibling tools like 'transform_image_from_encoded' or 'generate_image_from_text', missing explicit alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
- First observed
generate_image_from_text - First observed
transform_image_from_encoded - First observed
transform_image_from_file
TDQS
Scored across 3 tools
The three tools have distinct purposes: one generates images from text, while the other two transform existing images but differ in input method (encoded data vs. file path). There is minor overlap between the two transform tools, as they serve similar functions with different input formats, which could cause slight confusion, but their descriptions clearly differentiate them.
All tool names follow a consistent verb_noun_from_noun pattern (e.g., generate_image_from_text, transform_image_from_encoded, transform_image_from_file). This uniformity makes the set predictable and easy to understand, with no deviations in style or convention.
With only 3 tools, the server feels slightly thin for an image generation domain, as it lacks operations like listing, deleting, or managing generated images. However, the count is reasonable for basic functionality, covering core tasks without being excessive.
The tools cover generation and transformation of images, which are key operations, but there are notable gaps. For example, there is no way to retrieve, update, or delete generated images, and no tools for batch processing or status checking, which could limit agent workflows in a full image management context.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
- lightgenOAuthapp.lightgen
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
LLM chat, text tools, image generation, editing and batch image jobs
Generate AI images, video, speech, music and presentations from Claude, ChatGPT and Cursor.
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Related MCP Servers
- AlicenseBqualityCmaintenanceEnables Claude Desktop to generate text and analyze images using Google's Gemini Pro API. Provides seamless integration between Claude and Gemini's AI capabilities through natural language commands.2MIT
- AlicenseNot gradedqualityNot gradedmaintenanceGenerates AI images using Google Imagen directly in Claude Desktop, with automatic saving and support for multiple images and custom aspect ratios.-
- AlicenseAqualityDmaintenanceEnables AI image generation, editing, composition, and style transfer in Claude conversations using Google's Gemini 2.5 Flash model. Automatically saves generated images to a local directory.42711MIT
- AlicenseNot gradedqualityDmaintenanceEnables image generation and transformation using Google Gemini AI, with automatic English translation for multilingual prompts and local saving.MIT
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/sungmin-koo-ai/GeminiImageMCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server