Skip to main content
Glama
sungmin-koo-ai

Gemini Image Generator MCP

Gemini 画像生成器 MCP サーバー (修正版)

Claude DesktopでGoogle の Gemini AIを使用して高品質な画像を生成・編集できる MCP サーバーです。

🚀 主な特徴

  • テキストから画像生成: Gemini 2.0 Flash を使用したテキスト→画像変換

  • 画像変換: 既存の画像をテキストプロンプトで修正

  • 多言語対応: 日本語・韓国語・中国語プロンプトの自動英語翻訳・最適化

  • AI ファイル名生成: プロンプト基準でファイル名を自動生成

  • ローカル保存: 生成された画像を指定フォルダに自動保存

  • Claude チャット内表示: 生成された画像をチャット画面で直接確認

Related MCP server: NanoBanana MCP

🛠️ インストール要件

  • Python 3.11 以上

  • Google Gemini API キー

  • Claude Desktop またはその他 MCP 互換クライアント

📋 ステップ1: Gemini API キー発行

  1. Google AI Studio API Keys ページ にアクセス

  2. Google アカウントでログイン

  3. "Create API Key" をクリック

  4. 生成された API キーをコピー(後で使用)

💾 ステップ2: MCP サーバーインストール

自動インストール(推奨)

# リポジトリクローン
git clone https://github.com/sungmin-koo-ai/GeminiImageMCP.git
cd GeminiImageMCP

# 仮想環境作成・有効化
python3 -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate

# パッケージインストール
pip install -e .

インストール確認

# サーバーが正常実行されるかテスト
python -m gemini_image_mcp.server

Starting Gemini Image Generator MCP server... メッセージが表示されれば成功!(Ctrl+CまたはCtrl+Z で終了)

⚙️ ステップ3: Claude Desktop 設定

設定ファイル場所

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

設定ファイル内容

{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/python",
      "args": [
        "-m", "gemini_image_mcp.server"
      ],
      "env": {
        "GEMINI_API_KEY": "ここに実際のAPIキーを入力",
        "OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
      }
    }
  }
}

実際の設定例

{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/python",
      "args": [
        "-m", "gemini_image_mcp.server"
      ],
      "env": {
        "GEMINI_API_KEY": "AIzaSy...(実際のAPIキー)",
        "OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
      }
    }
  }
}

スクリプト命令を使用する場合(簡単設定)

{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/gemini-image-mcp",
      "env": {
        "GEMINI_API_KEY": "AIzaSy...(実際のAPIキー)",
        "OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
      }
    }
  }
}

※ この方法は args 指定が不要でより簡潔です

🚨 重要事項

  1. 絶対パス使用: すべてのパスは完全パスで入力

  2. API キー置換: ここに実際のAPIキーを入力 部分を発行した実際のキーに置換

  3. 画像フォルダ: OUTPUT_IMAGE_PATH に指定したフォルダが事前に作成されている必要があります

画像保存フォルダ作成

mkdir -p ~/Pictures/ai_generated

🎯 ステップ4: 実行・テスト

  1. Claude Desktop 再起動: 設定後完全に終了して再起動

  2. 接続確認: Claude Desktop で MCP サーバーが接続されたか確認

  3. テスト: 「猫の絵を描いて」とリクエストしてみる

📖 使用方法

画像生成

東京タワーの可愛いイラストを描いて

画像変換(ファイルパス)

/Users/username/image.jpg この画像に虹を追加して

画像変換(アップロード)

画像を Claude にアップロード後:

背景をレインボーブリッジの夜景にしてくれ

🔧 トラブルシューティング

サーバー接続失敗

  1. ログ確認: Claude Desktop のログフォルダで gemini-image-generator.log を確認

  2. パス確認: claude_desktop_config.json の Python パスが正確か確認

  3. 権限確認: 画像保存フォルダに書き込み権限があるか確認

API キーエラー

  1. キー有効性: Google AI Studio で API キーが有効化されているか確認

  2. 引用符確認: 設定ファイルで API キーが引用符で囲まれているか確認

手動テスト

cd ~/GeminiImageMCP
source venv/bin/activate
export GEMINI_API_KEY="実際のAPIキー"
export OUTPUT_IMAGE_PATH="~/Pictures/ai_generated"
python -m gemini_image_mcp.server

📊 提供ツール

1. generate_image_from_text

  • 機能: テキストプロンプトで新しい画像生成

  • 入力: 画像説明テキスト

  • 出力: 生成された画像(Claude チャット内表示 + ローカル保存)

2. transform_image_from_file

  • 機能: ファイルパスの画像をテキストプロンプトで変換

  • 入力: 画像ファイルパス、変換プロンプト

  • 出力: 変換された画像(Claude チャット内表示 + ローカル保存)

3. transform_image_from_encoded

  • 機能: Base64 エンコードされた画像をテキストプロンプトで変換

  • 入力: Base64 画像データ、変換プロンプト

  • 出力: 変換された画像(Claude チャット内表示 + ローカル保存)

📝 オリジナルからの相違点

この修正版は元のリポジトリの以下の問題を解決しました:

  • 元の問題: JSON シリアル化エラー (invalid utf-8 sequence)

  • 元の問題: MCP ツールがバイナリデータ返却により実行失敗

  • 修正事項: ファイルパス返却で安定的な動作

  • 修正事項: Claude Desktop で完璧に動作

  • 修正事項: 生成された画像を Claude チャット内で直接確認可能

🤝 貢献・お問い合わせ

📄 ライセンス

MIT License - 元のプロジェクトと同じ


ヒント: 初回設定時はステップごとに進め、問題が発生した場合はまずログファイルを確認してください! 🚀

Available Tools

3 tools
generate_image_from_textA

Generate an image based on the given text prompt using Google's Gemini model.

Args:
    prompt: User's text prompt describing the desired image to generate
    
Returns:
    Path to the generated image file using Gemini's image generation capabilities
ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the model and return type (path to image file) but lacks critical details such as rate limits, authentication requirements, image format, size, quality, or error handling. This is insufficient for a generative AI tool with potential costs and constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core functionality. The structured sections (Args, Returns) enhance readability, though the second sentence could be more integrated to avoid slight redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of image generation, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits (e.g., costs, latency), output specifics (e.g., file format, resolution), and error cases, leaving significant gaps for an AI agent to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, but the description compensates by explaining the single parameter ('prompt') as 'User's text prompt describing the desired image to generate.' This adds meaningful context beyond the schema's basic type information, clarifying the parameter's role in the generation process.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Generate an image') and resource ('based on the given text prompt'), using Google's Gemini model. It distinguishes from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' by specifying text-based generation rather than transformation from existing images.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for text-to-image generation but does not explicitly state when to use this tool versus alternatives. It mentions the model (Gemini) but provides no guidance on prerequisites, limitations, or scenarios where other tools might be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_encodedA

Transform an existing image based on the given text prompt using Google's Gemini model.

Args:
    encoded_image: Base64 encoded image data with header. Must be in format:
                "data:image/[format];base64,[data]"
                Where [format] can be: png, jpeg, jpg, gif, webp, etc.
    prompt: Text prompt describing the desired transformation or modifications
    
Returns:
    Path to the transformed image file saved on the server
ParametersJSON Schema
NameRequiredDescriptionDefault
encoded_imageYes
promptYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses the tool uses Google's Gemini model and that it saves the transformed image on the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or potential side effects of the transformation process.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear opening sentence stating the purpose, followed by well-organized sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with no annotations and no output schema, the description provides good coverage of purpose, parameters, and basic behavior. It explains what the tool does, how to format inputs, and what to expect as output. The main gap is lack of information about error conditions, performance characteristics, or more detailed behavioral constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by providing detailed semantics for both parameters. It specifies the exact format required for encoded_image (including header format and supported image types) and explains what the prompt parameter should contain. This adds significant value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verb ('Transform') and resource ('an existing image'), and distinguishes it from siblings by specifying it uses encoded image data rather than text or file inputs. The mention of Google's Gemini model adds technical specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when to use this tool (transforming existing images with encoded data) and implicitly distinguishes it from siblings (generate_image_from_text for text-to-image, transform_image_from_file for file-based transformation). However, it doesn't explicitly state when NOT to use this tool or mention specific prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_fileA

Transform an existing image file based on the given text prompt using Google's Gemini model.

Args:
    image_file_path: Path to the image file to be transformed
    prompt: Text prompt describing the desired transformation or modifications
    
Returns:
    Path to the transformed image file saved on the server
ParametersJSON Schema
NameRequiredDescriptionDefault
image_file_pathYes
promptYes

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions the tool saves the transformed file on the server, which is useful behavioral context. However, it lacks critical details like required permissions, file format limitations, transformation scope, error handling, or whether the operation is reversible/destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear purpose statement followed by labeled sections for Args and Returns. Every sentence adds value without redundancy, and information is front-loaded appropriately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 2 parameters, the description covers purpose and parameters adequately. However, for a transformation tool with potential complexity (image processing via Gemini), it lacks details about output format, file location specifics, or error cases, leaving gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by explaining both parameters: 'image_file_path' as 'Path to the image file to be transformed' and 'prompt' as 'Text prompt describing the desired transformation or modifications'. This adds essential meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verb ('Transform') and resource ('existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The mention of using Google's Gemini model adds technical specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by specifying it transforms 'an existing image file' and uses a 'text prompt', which differentiates it from 'generate_image_from_text' (creates new images) and 'transform_image_from_encoded' (uses encoded input). However, it doesn't explicitly state when to choose this tool over alternatives or any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updates
    • First observedgenerate_image_from_text
    • First observedtransform_image_from_encoded
    • First observedtransform_image_from_file

TDQS

A3.8/5.0

Scored across 3 tools

Disambiguation4/5

The three tools have distinct purposes: one generates images from text, while the other two transform existing images but differ in input method (encoded data vs. file path). There is minor overlap between the two transform tools, as they serve similar functions with different input formats, which could cause slight confusion, but their descriptions clearly differentiate them.

Naming Consistency5/5

All tool names follow a consistent verb_noun_from_noun pattern (e.g., generate_image_from_text, transform_image_from_encoded, transform_image_from_file). This uniformity makes the set predictable and easy to understand, with no deviations in style or convention.

Tool Count3/5

With only 3 tools, the server feels slightly thin for an image generation domain, as it lacks operations like listing, deleting, or managing generated images. However, the count is reasonable for basic functionality, covering core tasks without being excessive.

Completeness3/5

The tools cover generation and transformation of images, which are key operations, but there are notable gaps. For example, there is no way to retrieve, update, or delete generated images, and no tools for batch processing or status checking, which could limit agent workflows in a full image management context.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers