Skip to main content
Glama
sungmin-koo-ai

Gemini Image Generator MCP

Gemini 画像生成器 MCP サーバー (修正版)

Claude DesktopでGoogle の Gemini AIを使用して高品質な画像を生成・編集できる MCP サーバーです。

🚀 主な特徴

  • テキストから画像生成: Gemini 2.0 Flash を使用したテキスト→画像変換

  • 画像変換: 既存の画像をテキストプロンプトで修正

  • 多言語対応: 日本語・韓国語・中国語プロンプトの自動英語翻訳・最適化

  • AI ファイル名生成: プロンプト基準でファイル名を自動生成

  • ローカル保存: 生成された画像を指定フォルダに自動保存

  • Claude チャット内表示: 生成された画像をチャット画面で直接確認

Related MCP server: NanoBanana MCP

🛠️ インストール要件

  • Python 3.11 以上

  • Google Gemini API キー

  • Claude Desktop またはその他 MCP 互換クライアント

📋 ステップ1: Gemini API キー発行

  1. Google AI Studio API Keys ページ にアクセス

  2. Google アカウントでログイン

  3. "Create API Key" をクリック

  4. 生成された API キーをコピー(後で使用)

💾 ステップ2: MCP サーバーインストール

自動インストール(推奨)

# リポジトリクローン
git clone https://github.com/sungmin-koo-ai/GeminiImageMCP.git
cd GeminiImageMCP

# 仮想環境作成・有効化
python3 -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate

# パッケージインストール
pip install -e .

インストール確認

# サーバーが正常実行されるかテスト
python -m gemini_image_mcp.server

Starting Gemini Image Generator MCP server... メッセージが表示されれば成功!(Ctrl+CまたはCtrl+Z で終了)

⚙️ ステップ3: Claude Desktop 設定

設定ファイル場所

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

設定ファイル内容

{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/python",
      "args": [
        "-m", "gemini_image_mcp.server"
      ],
      "env": {
        "GEMINI_API_KEY": "ここに実際のAPIキーを入力",
        "OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
      }
    }
  }
}

実際の設定例

{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/python",
      "args": [
        "-m", "gemini_image_mcp.server"
      ],
      "env": {
        "GEMINI_API_KEY": "AIzaSy...(実際のAPIキー)",
        "OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
      }
    }
  }
}

スクリプト命令を使用する場合(簡単設定)

{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "/Users/ユーザー名/GeminiImageMCP/venv/bin/gemini-image-mcp",
      "env": {
        "GEMINI_API_KEY": "AIzaSy...(実際のAPIキー)",
        "OUTPUT_IMAGE_PATH": "/Users/ユーザー名/Pictures/ai_generated"
      }
    }
  }
}

※ この方法は args 指定が不要でより簡潔です

🚨 重要事項

  1. 絶対パス使用: すべてのパスは完全パスで入力

  2. API キー置換: ここに実際のAPIキーを入力 部分を発行した実際のキーに置換

  3. 画像フォルダ: OUTPUT_IMAGE_PATH に指定したフォルダが事前に作成されている必要があります

画像保存フォルダ作成

mkdir -p ~/Pictures/ai_generated

🎯 ステップ4: 実行・テスト

  1. Claude Desktop 再起動: 設定後完全に終了して再起動

  2. 接続確認: Claude Desktop で MCP サーバーが接続されたか確認

  3. テスト: 「猫の絵を描いて」とリクエストしてみる

📖 使用方法

画像生成

東京タワーの可愛いイラストを描いて

画像変換(ファイルパス)

/Users/username/image.jpg この画像に虹を追加して

画像変換(アップロード)

画像を Claude にアップロード後:

背景をレインボーブリッジの夜景にしてくれ

🔧 トラブルシューティング

サーバー接続失敗

  1. ログ確認: Claude Desktop のログフォルダで gemini-image-generator.log を確認

  2. パス確認: claude_desktop_config.json の Python パスが正確か確認

  3. 権限確認: 画像保存フォルダに書き込み権限があるか確認

API キーエラー

  1. キー有効性: Google AI Studio で API キーが有効化されているか確認

  2. 引用符確認: 設定ファイルで API キーが引用符で囲まれているか確認

手動テスト

cd ~/GeminiImageMCP
source venv/bin/activate
export GEMINI_API_KEY="実際のAPIキー"
export OUTPUT_IMAGE_PATH="~/Pictures/ai_generated"
python -m gemini_image_mcp.server

📊 提供ツール

1. generate_image_from_text

  • 機能: テキストプロンプトで新しい画像生成

  • 入力: 画像説明テキスト

  • 出力: 生成された画像(Claude チャット内表示 + ローカル保存)

2. transform_image_from_file

  • 機能: ファイルパスの画像をテキストプロンプトで変換

  • 入力: 画像ファイルパス、変換プロンプト

  • 出力: 変換された画像(Claude チャット内表示 + ローカル保存)

3. transform_image_from_encoded

  • 機能: Base64 エンコードされた画像をテキストプロンプトで変換

  • 入力: Base64 画像データ、変換プロンプト

  • 出力: 変換された画像(Claude チャット内表示 + ローカル保存)

📝 オリジナルからの相違点

この修正版は元のリポジトリの以下の問題を解決しました:

  • 元の問題: JSON シリアル化エラー (invalid utf-8 sequence)

  • 元の問題: MCP ツールがバイナリデータ返却により実行失敗

  • 修正事項: ファイルパス返却で安定的な動作

  • 修正事項: Claude Desktop で完璧に動作

  • 修正事項: 生成された画像を Claude チャット内で直接確認可能

🤝 貢献・お問い合わせ

📄 ライセンス

MIT License - 元のプロジェクトと同じ


ヒント: 初回設定時はステップごとに進め、問題が発生した場合はまずログファイルを確認してください! 🚀

Available Tools

3 tools
generate_image_from_textA

Generate an image based on the given text prompt using Google's Gemini model.

Args:
    prompt: User's text prompt describing the desired image to generate
    
Returns:
    Path to the generated image file using Gemini's image generation capabilities
ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the action ('Generate') and output ('Path to the generated image file'), but does not cover critical aspects like rate limits, authentication needs, file formats, error handling, or whether the operation is idempotent, leaving significant gaps for a generative tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections for Args and Returns, and each sentence adds value. It is appropriately sized for the tool's complexity, though it could be slightly more concise by integrating the technology mention into the main purpose statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (generative AI operation), no annotations, and no output schema, the description provides basic purpose and parameter info but lacks details on behavioral traits, error cases, or output specifics beyond a path. It is minimally viable but has clear gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaningful context for the single parameter 'prompt' by explaining it as 'User's text prompt describing the desired image to generate', which goes beyond the schema's minimal title. With 0% schema description coverage and only one parameter, this adequately compensates, though it could include examples or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Generate an image'), resource ('based on the given text prompt'), and technology ('using Google's Gemini model'), distinguishing it from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' which involve transformation rather than generation from text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for image generation from text prompts but does not explicitly state when to use this tool versus its siblings. It mentions the technology (Gemini model) which provides some context, but lacks explicit guidance on alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_encodedA

Transform an existing image based on the given text prompt using Google's Gemini model.

Args:
    encoded_image: Base64 encoded image data with header. Must be in format:
                "data:image/[format];base64,[data]"
                Where [format] can be: png, jpeg, jpg, gif, webp, etc.
    prompt: Text prompt describing the desired transformation or modifications
    
Returns:
    Path to the transformed image file saved on the server
ParametersJSON Schema
NameRequiredDescriptionDefault
encoded_imageYes
promptYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses that the tool uses Google's Gemini model and saves the transformed image to the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or error conditions that would be important for a transformation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear purpose statement followed by well-organized Args and Returns sections. Every sentence earns its place by providing essential information without redundancy. The formatting with clear section headers enhances readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a transformation tool with no annotations and no output schema, the description provides good parameter documentation and purpose clarity. However, it lacks information about the transformation process (e.g., quality, limitations, processing time), error handling, and more detailed behavioral context that would be valuable given the complexity of image transformation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by providing detailed semantic information for both parameters. It specifies the exact format required for encoded_image (including header structure and supported formats) and explains that prompt describes 'desired transformation or modifications.' This adds significant value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Transform an existing image based on the given text prompt using Google's Gemini model.' It specifies the verb (transform), resource (existing image), method (Gemini model), and distinguishes from sibling tools (transform_image_from_file handles file input instead of encoded data).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly suggests usage context by mentioning 'existing image' and the encoding format, but doesn't explicitly state when to use this tool versus alternatives like transform_image_from_file or generate_image_from_text. It provides technical prerequisites but lacks comparative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_fileA

Transform an existing image file based on the given text prompt using Google's Gemini model.

Args:
    image_file_path: Path to the image file to be transformed
    prompt: Text prompt describing the desired transformation or modifications
    
Returns:
    Path to the transformed image file saved on the server
ParametersJSON Schema
NameRequiredDescriptionDefault
image_file_pathYes
promptYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that the tool saves the transformed file on the server, which is useful context beyond basic functionality. However, it lacks details on permissions, rate limits, error handling, or what transformations are possible, leaving behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded with the core purpose, followed by clear sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of an AI-based image transformation tool with no annotations and no output schema, the description is moderately complete. It covers the basic operation and return value but lacks details on supported image formats, transformation limits, or error cases, which are important for such a tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning beyond the input schema, which has 0% coverage. It explains that 'image_file_path' is for an existing image file to be transformed and 'prompt' describes the desired transformation, clarifying the purpose and usage of both parameters effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Transform') and resource ('an existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The phrase 'using Google's Gemini model' adds technical specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly provides usage context by mentioning 'existing image file' and the Gemini model, which suggests when to use this tool. However, it doesn't explicitly state when to choose this over sibling tools like 'transform_image_from_encoded' or 'generate_image_from_text', missing explicit alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updates
    • First observedgenerate_image_from_text
    • First observedtransform_image_from_encoded
    • First observedtransform_image_from_file

TDQS

A3.8/5.0

Scored across 3 tools

Disambiguation4/5

The three tools have distinct purposes: one generates images from text, while the other two transform existing images but differ in input method (encoded data vs. file path). There is minor overlap between the two transform tools, as they serve similar functions with different input formats, which could cause slight confusion, but their descriptions clearly differentiate them.

Naming Consistency5/5

All tool names follow a consistent verb_noun_from_noun pattern (e.g., generate_image_from_text, transform_image_from_encoded, transform_image_from_file). This uniformity makes the set predictable and easy to understand, with no deviations in style or convention.

Tool Count3/5

With only 3 tools, the server feels slightly thin for an image generation domain, as it lacks operations like listing, deleting, or managing generated images. However, the count is reasonable for basic functionality, covering core tasks without being excessive.

Completeness3/5

The tools cover generation and transformation of images, which are key operations, but there are notable gaps. For example, there is no way to retrieve, update, or delete generated images, and no tools for batch processing or status checking, which could limit agent workflows in a full image management context.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/sungmin-koo-ai/GeminiImageMCP'

If you have feedback or need assistance with the MCP directory API, please join our Discord server