Skip to main content
Glama
qhdrl12

Gemini Image Generator MCP Server

by qhdrl12

Gemini 画像ジェネレーター MCP サーバー

MCP プロトコルを介して Google の Gemini モデルを使用して、テキスト プロンプトから高品質の画像を生成します。

概要

このMCPサーバーは、あらゆるAIアシスタントがGoogleのGemini AIモデルを使用して画像を生成できるようにします。このサーバーはプロンプトエンジニアリング、テキストから画像への変換、ファイル名の生成、ローカル画像ストレージを処理するため、あらゆるMCPクライアントからAI生成画像を簡単に作成・管理できます。

Related MCP server: Gemini Image Gen MCP Server

特徴

  • Gemini 2.0 Flashを使用したテキストから画像への生成

  • テキストプロンプトに基づく画像間の変換

  • ファイルベースとbase64エンコードされた画像の両方をサポート

  • プロンプトに基づいて自動的にインテリジェントなファイル名を生成する

  • 英語以外のプロンプトの自動翻訳

  • 設定可能な出力パスを備えたローカル画像ストレージ

  • 生成された画像からテキストを厳密に除外する

  • 高解像度画像出力

  • 画像データとファイルパスの両方に直接アクセス

利用可能なMCPツール

サーバーは、AI アシスタント用に次の MCP ツールを提供します。

1. generate_image_from_text

テキストプロンプトの説明から新しい画像を作成します。

generate_image_from_text(prompt: str) -> Tuple[bytes, str]

パラメータ:

  • prompt : 生成したい画像のテキスト説明

戻り値:

  • 次の内容を含むタプル:

    • 生画像データ(バイト)

    • 保存された画像ファイルへのパス (str)

このデュアルリターン形式により、AI アシスタントは画像データを直接操作したり、保存されたファイルパスを参照したりすることができます。

例:

  • 「山に沈む夕日の画像を生成する」

  • 「SF都市でフォトリアリスティックな空飛ぶ豚を作ろう」

出力例

この画像はプロンプトを使用して生成されました:

"Hi, can you create a 3d rendered image of a pig with wings and a top hat flying over a happy futuristic scifi city with lots of greenery?"

SF都市の上空を飛ぶ豚

翼とシルクハットをつけた3Dレンダリングされた豚が、緑豊かな未来のSF都市の上を飛んでいます。

既知の問題

この MCP サーバーを Claude Desktop Host で使用する場合:

  1. パフォーマンスの問題transform_image_from_encoded使用すると、他の方法と比較して処理時間が大幅に長くなる可能性があります。これは、MCP プロトコルを介して大きな base64 エンコードされた画像データを転送する際のオーバーヘッドが原因です。

  2. パス解決の問題:Claude Desktop Host の使用時に、画像パスを正しく解決できない問題が発生する可能性があります。ホストアプリケーションが返されたファイルパスを正しく解釈できず、生成された画像にアクセスできなくなる可能性があります。

最良のエクスペリエンスを得るには、可能な場合は代替の MCP クライアントまたはtransform_image_from_fileメソッドの使用を検討してください。

2. transform_image_from_encoded

base64 でエンコードされた画像データを使用して、テキスト プロンプトに基づいて既存の画像を変換します。

transform_image_from_encoded(encoded_image: str, prompt: str) -> Tuple[bytes, str]

パラメータ:

  • encoded_image : フォーマットヘッダー付きのBase64エンコードされた画像データ(形式は「data:image/[format];base64,[data]」である必要があります)

  • prompt : 画像をどのように変換したいかのテキスト説明

戻り値:

  • 次の内容を含むタプル:

    • 変換された生画像データ(バイト)

    • 保存された変換された画像ファイルへのパス(str)

例:

  • 「この風景に雪を加えよう」

  • 「背景をビーチに変更」

3. ファイルtransform_image_from_file

テキスト プロンプトに基づいて既存の画像ファイルを変換します。

transform_image_from_file(image_file_path: str, prompt: str) -> Tuple[bytes, str]

パラメータ:

  • image_file_path : 変換する画像ファイルへのパス

  • prompt : 画像をどのように変換したいかのテキスト説明

戻り値:

  • 次の内容を含むタプル:

    • 変換された生画像データ(バイト)

    • 保存された変換された画像ファイルへのパス(str)

例:

  • 「この画像の人物の隣にラマを追加してください」

  • 「この昼間のシーンを夜のように見せましょう」

変換例

上記で作成した空飛ぶ豚の画像を使用して、次のプロンプトで変換を適用しました。

"Add a cute baby whale flying alongside the pig"

前に:SF都市の上空を飛ぶ豚

後:空飛ぶ豚と子クジラ

かわいい赤ちゃんクジラが一緒に飛んでいるオリジナルの空飛ぶ豚の画像

設定

前提条件

  • Python 3.11以上

  • Google AI API キー (Gemini)

  • MCP ホスト アプリケーション (Claude デスクトップ アプリ、カーソル、またはその他の MCP 互換クライアント)

Gemini APIキーの取得

  1. Google AI Studio APIキーページにアクセスしてください

  2. Googleアカウントでログイン

  3. 「APIキーを作成」をクリックします

  4. 設定で使用するために新しいAPIキーをコピーします

  5. 注: APIキーは、毎月一定量の無料利用を提供します。使用量はGoogle AI Studioで確認できます。

インストール

  1. リポジトリをクローンします。

git clone https://github.com/your-username/gemini-image-generator.git
cd gemini-image-generator
  1. 仮想環境を作成し、依存関係をインストールします。

# Using regular venv
python -m venv .venv
source .venv/bin/activate
pip install -e .

# Or using uv
uv venv
source .venv/bin/activate
uv pip install -e .
  1. サンプル環境ファイルをコピーし、API キーを追加します。

cp .env.example .env
  1. .envファイルを編集して、Google Gemini API キーと優先出力パスを追加します。

GEMINI_API_KEY="your-gemini-api-key-here"
OUTPUT_IMAGE_PATH="/path/to/save/images"

Claudeデスクトップの設定

claude_desktop_config.jsonに以下を追加します。

  • macOS : ~/Library/Application Support/Claude/claude_desktop_config.json

{
    "mcpServers": {
        "gemini-image-generator": {
            "command": "uv",
            "args": [
                "--directory",
                "/absolute/path/to/gemini-image-generator",
                "run",
                "server.py"
            ],
            "env": {
                "GEMINI_API_KEY": "GEMINI_API_KEY",
                "OUTPUT_IMAGE_PATH": "OUTPUT_IMAGE_PATH"
            }
        }
    }
}

使用法

インストールして設定したら、次のようなプロンプトを使用して、Claude に画像を生成または変換するように依頼できます。

新しい画像の生成

  • 「山に沈む夕日の画像を生成する」

  • 「未来都市の風景をイラストで表現する」

  • 「サングラスをかけた猫の絵を描いてください」

既存の画像の変換

  • 「シーンに雪を追加してこの画像を変形させます」

  • 「この写真を編集して夜に撮ったように見せてください」

  • 「この写真の背景に飛んでいるドラゴンを追加してください」

生成/変換された画像は、設定された出力パスに保存され、Claudeに表示されます。更新された戻り値の型により、AIアシスタントは保存されたファイルにアクセスすることなく、画像データを直接操作できるようになります。

テスト

FastMCP 開発サーバーを実行してアプリケーションをテストできます。

fastmcp dev server.py

このコマンドはローカル開発サーバーを起動し、MCP Inspector をhttp://localhost:5173/で利用できるようにします。MCP Inspector は便利なウェブインターフェースを提供しており、Claude や他の MCP クライアントを使用せずに画像生成ツールを直接テストできます。テキストプロンプトを入力してツールを実行すると、すぐに結果が表示されるため、開発やデバッグに役立ちます。

ライセンス

MITライセンス

Available Tools

3 tools
generate_image_from_textA

Generate an image based on the given text prompt using Google's Gemini model.

Args:
    prompt: User's text prompt describing the desired image to generate
    
Returns:
    Path to the generated image file using Gemini's image generation capabilities
ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the model and return type (path to image file) but lacks critical details such as rate limits, authentication requirements, image format, size, quality, or error handling. This is insufficient for a generative AI tool with potential costs and constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core functionality. The structured sections (Args, Returns) enhance readability, though the second sentence could be more integrated to avoid slight redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of image generation, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits (e.g., costs, latency), output specifics (e.g., file format, resolution), and error cases, leaving significant gaps for an AI agent to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, but the description compensates by explaining the single parameter ('prompt') as 'User's text prompt describing the desired image to generate.' This adds meaningful context beyond the schema's basic type information, clarifying the parameter's role in the generation process.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Generate an image') and resource ('based on the given text prompt'), using Google's Gemini model. It distinguishes from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' by specifying text-based generation rather than transformation from existing images.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for text-to-image generation but does not explicitly state when to use this tool versus alternatives. It mentions the model (Gemini) but provides no guidance on prerequisites, limitations, or scenarios where other tools might be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_encodedA

Transform an existing image based on the given text prompt using Google's Gemini model.

Args:
    encoded_image: Base64 encoded image data with header. Must be in format:
                "data:image/[format];base64,[data]"
                Where [format] can be: png, jpeg, jpg, gif, webp, etc.
    prompt: Text prompt describing the desired transformation or modifications
    
Returns:
    Path to the transformed image file saved on the server
ParametersJSON Schema
NameRequiredDescriptionDefault
encoded_imageYes
promptYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses the tool uses Google's Gemini model and that it saves the transformed image on the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or potential side effects of the transformation process.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear opening sentence stating the purpose, followed by well-organized sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with no annotations and no output schema, the description provides good coverage of purpose, parameters, and basic behavior. It explains what the tool does, how to format inputs, and what to expect as output. The main gap is lack of information about error conditions, performance characteristics, or more detailed behavioral constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by providing detailed semantics for both parameters. It specifies the exact format required for encoded_image (including header format and supported image types) and explains what the prompt parameter should contain. This adds significant value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verb ('Transform') and resource ('an existing image'), and distinguishes it from siblings by specifying it uses encoded image data rather than text or file inputs. The mention of Google's Gemini model adds technical specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when to use this tool (transforming existing images with encoded data) and implicitly distinguishes it from siblings (generate_image_from_text for text-to-image, transform_image_from_file for file-based transformation). However, it doesn't explicitly state when NOT to use this tool or mention specific prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_image_from_fileA

Transform an existing image file based on the given text prompt using Google's Gemini model.

Args:
    image_file_path: Path to the image file to be transformed
    prompt: Text prompt describing the desired transformation or modifications
    
Returns:
    Path to the transformed image file saved on the server
ParametersJSON Schema
NameRequiredDescriptionDefault
image_file_pathYes
promptYes

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions the tool saves the transformed file on the server, which is useful behavioral context. However, it lacks critical details like required permissions, file format limitations, transformation scope, error handling, or whether the operation is reversible/destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear purpose statement followed by labeled sections for Args and Returns. Every sentence adds value without redundancy, and information is front-loaded appropriately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 2 parameters, the description covers purpose and parameters adequately. However, for a transformation tool with potential complexity (image processing via Gemini), it lacks details about output format, file location specifics, or error cases, leaving gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully compensates by explaining both parameters: 'image_file_path' as 'Path to the image file to be transformed' and 'prompt' as 'Text prompt describing the desired transformation or modifications'. This adds essential meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verb ('Transform') and resource ('existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The mention of using Google's Gemini model adds technical specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by specifying it transforms 'an existing image file' and uses a 'text prompt', which differentiates it from 'generate_image_from_text' (creates new images) and 'transform_image_from_encoded' (uses encoded input). However, it doesn't explicitly state when to choose this tool over alternatives or any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updates
    • First observedgenerate_image_from_text
    • First observedtransform_image_from_encoded
    • First observedtransform_image_from_file

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation5/5

The three tools have clearly distinct purposes: generate_image_from_text creates new images from text prompts, while transform_image_from_encoded and transform_image_from_file both transform existing images but differ in input format (base64 encoded vs. file path). The descriptions make these distinctions explicit, eliminating any potential confusion between generation and transformation operations.

Naming Consistency5/5

All tools follow a consistent verb_noun_from_source naming pattern: generate_image_from_text, transform_image_from_encoded, and transform_image_from_file. This pattern clearly indicates the action (generate/transform), the target (image), and the input source (text/encoded/file), creating a predictable and readable naming convention throughout the toolset.

Tool Count4/5

Three tools is a reasonable count for an image generation server, covering the core operations of generating new images and transforming existing ones. However, the scope feels slightly thin as there are no complementary tools for managing generated images (like listing, deleting, or retrieving metadata), which might limit agent workflows in production scenarios.

Completeness3/5

The server covers basic image generation and transformation operations well, but has notable gaps in image management. There are no tools for listing generated images, deleting files, retrieving image metadata, or batch operations. While the core generative AI functionality is present, the lack of lifecycle management tools creates potential dead ends for agents working with multiple images over time.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers