Gemini Image Generator MCP Server
Gemini Image Generator MCP-Server
Generieren Sie mithilfe des Gemini-Modells von Google über das MCP-Protokoll hochwertige Bilder aus Textaufforderungen.
Überblick
Dieser MCP-Server ermöglicht es jedem KI-Assistenten, Bilder mithilfe des Gemini-KI-Modells von Google zu generieren. Der Server übernimmt die Eingabeaufforderung, die Text-zu-Bild-Konvertierung, die Dateinamengenerierung und die lokale Bildspeicherung. So können KI-generierte Bilder ganz einfach über jeden MCP-Client erstellt und verwaltet werden.
Related MCP server: Gemini Image Gen MCP Server
Merkmale
Text-zu-Bild-Generierung mit Gemini 2.0 Flash
Bild-zu-Bild-Transformation basierend auf Textaufforderungen
Unterstützung sowohl für dateibasierte als auch für base64-kodierte Bilder
Automatische intelligente Dateinamengenerierung basierend auf Eingabeaufforderungen
Automatische Übersetzung nicht-englischer Eingabeaufforderungen
Lokaler Bildspeicher mit konfigurierbarem Ausgabepfad
Strikter Textausschluss aus generierten Bildern
Hochauflösende Bildausgabe
Direkter Zugriff auf Bilddaten und Dateipfad
Verfügbare MCP-Tools
Der Server stellt die folgenden MCP-Tools für KI-Assistenten bereit:
1. generate_image_from_text
Erstellt ein neues Bild aus einer Textaufforderungsbeschreibung.
generate_image_from_text(prompt: str) -> Tuple[bytes, str]Parameter:
prompt: Textbeschreibung des Bildes, das Sie generieren möchten
Widerrufsfolgen:
Ein Tupel mit:
Rohbilddaten (Bytes)
Pfad zur gespeicherten Bilddatei (str)
Dieses duale Rückgabeformat ermöglicht es KI-Assistenten, entweder direkt mit den Bilddaten zu arbeiten oder auf den gespeicherten Dateipfad zu verweisen.
Beispiele:
"Erstellen Sie ein Bild eines Sonnenuntergangs über den Bergen"
"Erstellen Sie ein fotorealistisches fliegendes Schwein in einer Science-Fiction-Stadt"
Beispielausgabe
Dieses Bild wurde mit der Eingabeaufforderung generiert:
"Hi, can you create a 3d rendered image of a pig with wings and a top hat flying over a happy futuristic scifi city with lots of greenery?"
Ein 3D-gerendertes Schwein mit Flügeln und Zylinder, das über einer futuristischen Science-Fiction-Stadt voller Grün fliegt
Bekannte Probleme
Bei Verwendung dieses MCP-Servers mit Claude Desktop Host:
Leistungsprobleme : Die Verwendung von
transform_image_from_encodedkann im Vergleich zu anderen Methoden deutlich länger dauern. Dies liegt am Overhead bei der Übertragung großer base64-codierter Bilddaten über das MCP-Protokoll.Probleme bei der Pfadauflösung : Bei der Verwendung von Claude Desktop Host kann es zu Problemen mit der korrekten Auflösung von Bildpfaden kommen. Die Hostanwendung interpretiert die zurückgegebenen Dateipfade möglicherweise nicht richtig, was den Zugriff auf die generierten Bilder erschwert.
Für ein optimales Erlebnis sollten Sie nach Möglichkeit alternative MCP-Clients oder die Methode transform_image_from_file verwenden.
2. transform_image_from_encoded
Transformiert ein vorhandenes Bild basierend auf einer Textaufforderung unter Verwendung von Base64-codierten Bilddaten.
transform_image_from_encoded(encoded_image: str, prompt: str) -> Tuple[bytes, str]Parameter:
encoded_image: Base64-codierte Bilddaten mit Formatheader (müssen im Format „data:image/[format];base64,[data]“ vorliegen).prompt: Textbeschreibung, wie Sie das Bild transformieren möchten
Widerrufsfolgen:
Ein Tupel mit:
Rohe transformierte Bilddaten (Bytes)
Pfad zur gespeicherten transformierten Bilddatei (str)
Beispiel:
„Fügen Sie dieser Landschaft Schnee hinzu“
„Ändern Sie den Hintergrund in einen Strand“
3. transform_image_from_file
Transformiert eine vorhandene Bilddatei basierend auf einer Textaufforderung.
transform_image_from_file(image_file_path: str, prompt: str) -> Tuple[bytes, str]Parameter:
image_file_path: Pfad zur zu transformierenden Bilddateiprompt: Textbeschreibung, wie Sie das Bild transformieren möchten
Widerrufsfolgen:
Ein Tupel mit:
Rohe transformierte Bilddaten (Bytes)
Pfad zur gespeicherten transformierten Bilddatei (str)
Beispiele:
„Fügen Sie neben der Person in diesem Bild ein Lama hinzu.“
„Lassen Sie diese Tagesszene wie eine Nachtszene aussehen“
Beispieltransformation
Mithilfe des oben erstellten Bildes des fliegenden Schweins haben wir eine Transformation mit der folgenden Eingabeaufforderung angewendet:
"Add a cute baby whale flying alongside the pig"Vor: 
Nach:
Das Originalbild eines fliegenden Schweins mit einem süßen Walbaby, das daneben fliegt
Aufstellen
Voraussetzungen
Python 3.11+
Google AI API-Schlüssel (Gemini)
MCP-Hostanwendung (Claude Desktop App, Cursor oder andere MCP-kompatible Clients)
Abrufen eines Gemini-API-Schlüssels
Besuchen Sie die Seite mit den API-Schlüsseln von Google AI Studio
Melden Sie sich mit Ihrem Google-Konto an
Klicken Sie auf „API-Schlüssel erstellen“
Kopieren Sie Ihren neuen API-Schlüssel zur Verwendung in der Konfiguration
Hinweis: Der API-Schlüssel bietet ein bestimmtes Kontingent an kostenloser Nutzung pro Monat. Sie können Ihre Nutzung im Google AI Studio überprüfen.
Installation
Installation über Smithery
So installieren Sie Gemini Image Generator MCP für Claude Desktop automatisch über Smithery :
npx -y @smithery/cli install @qhdrl12/mcp-server-gemini-image-gen --client claudeManuelle Installation
Klonen Sie das Repository:
git clone https://github.com/your-username/gemini-image-generator.git
cd gemini-image-generatorErstellen Sie eine virtuelle Umgebung und installieren Sie Abhängigkeiten:
# Using regular venv
python -m venv .venv
source .venv/bin/activate
pip install -e .
# Or using uv
uv venv
source .venv/bin/activate
uv pip install -e .Kopieren Sie die Beispielumgebungsdatei und fügen Sie Ihren API-Schlüssel hinzu:
cp .env.example .envBearbeiten Sie die
.envDatei, um Ihren Google Gemini API-Schlüssel und den bevorzugten Ausgabepfad einzuschließen:
GEMINI_API_KEY="your-gemini-api-key-here"
OUTPUT_IMAGE_PATH="/path/to/save/images"Claude Desktop konfigurieren
Fügen Sie Ihrer claude_desktop_config.json Folgendes hinzu:
macOS :
~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"gemini-image-generator": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/gemini-image-generator",
"run",
"server.py"
],
"env": {
"GEMINI_API_KEY": "GEMINI_API_KEY",
"OUTPUT_IMAGE_PATH": "OUTPUT_IMAGE_PATH"
}
}
}
}Verwendung
Nach der Installation und Konfiguration können Sie Claude bitten, Bilder zu generieren oder zu transformieren, indem Sie Eingabeaufforderungen wie die folgenden verwenden:
Neue Bilder generieren
"Erstellen Sie ein Bild eines Sonnenuntergangs über den Bergen"
"Erstellen Sie eine Illustration einer futuristischen Stadtlandschaft"
„Machen Sie ein Bild von einer Katze mit Sonnenbrille“
Vorhandene Bilder transformieren
„Transformieren Sie dieses Bild, indem Sie der Szene Schnee hinzufügen.“
„Bearbeiten Sie dieses Foto, damit es aussieht, als wäre es nachts aufgenommen worden.“
„Fügen Sie im Hintergrund dieses Bildes einen fliegenden Drachen hinzu.“
Die generierten/transformierten Bilder werden in Ihrem konfigurierten Ausgabepfad gespeichert und in Claude angezeigt. Mit den aktualisierten Rückgabetypen können KI-Assistenten auch direkt mit den Bilddaten arbeiten, ohne auf die gespeicherten Dateien zugreifen zu müssen.
Testen
Sie können die Anwendung testen, indem Sie den FastMCP-Entwicklungsserver ausführen:
fastmcp dev server.pyDieser Befehl startet einen lokalen Entwicklungsserver und stellt den MCP Inspector unter http://localhost:5173/ bereit. Der MCP Inspector bietet eine praktische Weboberfläche, über die Sie das Bildgenerierungstool direkt testen können, ohne Claude oder einen anderen MCP-Client verwenden zu müssen. Sie können Texteingaben eingeben, das Tool ausführen und die Ergebnisse sofort sehen, was für Entwicklung und Debugging hilfreich ist.
Lizenz
MIT-Lizenz
Available Tools
3 toolsgenerate_image_from_textA
Generate an image based on the given text prompt using Google's Gemini model.
Args:
prompt: User's text prompt describing the desired image to generate
Returns:
Path to the generated image file using Gemini's image generation capabilities
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the model and return type (path to image file) but lacks critical details such as rate limits, authentication requirements, image format, size, quality, or error handling. This is insufficient for a generative AI tool with potential costs and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core functionality. The structured sections (Args, Returns) enhance readability, though the second sentence could be more integrated to avoid slight redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of image generation, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits (e.g., costs, latency), output specifics (e.g., file format, resolution), and error cases, leaving significant gaps for an AI agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, but the description compensates by explaining the single parameter ('prompt') as 'User's text prompt describing the desired image to generate.' This adds meaningful context beyond the schema's basic type information, clarifying the parameter's role in the generation process.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate an image') and resource ('based on the given text prompt'), using Google's Gemini model. It distinguishes from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' by specifying text-based generation rather than transformation from existing images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for text-to-image generation but does not explicitly state when to use this tool versus alternatives. It mentions the model (Gemini) but provides no guidance on prerequisites, limitations, or scenarios where other tools might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_encodedA
Transform an existing image based on the given text prompt using Google's Gemini model.
Args:
encoded_image: Base64 encoded image data with header. Must be in format:
"data:image/[format];base64,[data]"
Where [format] can be: png, jpeg, jpg, gif, webp, etc.
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| encoded_image | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the tool uses Google's Gemini model and that it saves the transformed image on the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or potential side effects of the transformation process.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear opening sentence stating the purpose, followed by well-organized sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no annotations and no output schema, the description provides good coverage of purpose, parameters, and basic behavior. It explains what the tool does, how to format inputs, and what to expect as output. The main gap is lack of information about error conditions, performance characteristics, or more detailed behavioral constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by providing detailed semantics for both parameters. It specifies the exact format required for encoded_image (including header format and supported image types) and explains what the prompt parameter should contain. This adds significant value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verb ('Transform') and resource ('an existing image'), and distinguishes it from siblings by specifying it uses encoded image data rather than text or file inputs. The mention of Google's Gemini model adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool (transforming existing images with encoded data) and implicitly distinguishes it from siblings (generate_image_from_text for text-to-image, transform_image_from_file for file-based transformation). However, it doesn't explicitly state when NOT to use this tool or mention specific prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_fileA
Transform an existing image file based on the given text prompt using Google's Gemini model.
Args:
image_file_path: Path to the image file to be transformed
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| image_file_path | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the tool saves the transformed file on the server, which is useful behavioral context. However, it lacks critical details like required permissions, file format limitations, transformation scope, error handling, or whether the operation is reversible/destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear purpose statement followed by labeled sections for Args and Returns. Every sentence adds value without redundancy, and information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 2 parameters, the description covers purpose and parameters adequately. However, for a transformation tool with potential complexity (image processing via Gemini), it lacks details about output format, file location specifics, or error cases, leaving gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining both parameters: 'image_file_path' as 'Path to the image file to be transformed' and 'prompt' as 'Text prompt describing the desired transformation or modifications'. This adds essential meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verb ('Transform') and resource ('existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The mention of using Google's Gemini model adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying it transforms 'an existing image file' and uses a 'text prompt', which differentiates it from 'generate_image_from_text' (creates new images) and 'transform_image_from_encoded' (uses encoded input). However, it doesn't explicitly state when to choose this tool over alternatives or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
generate_image_from_text - First observed
transform_image_from_encoded - First observed
transform_image_from_file
TDQS
Scored across 3 tools
The three tools have clearly distinct purposes: generate_image_from_text creates new images from text prompts, while transform_image_from_encoded and transform_image_from_file both transform existing images but differ in input format (base64 encoded vs. file path). The descriptions make these distinctions explicit, eliminating any potential confusion between generation and transformation operations.
All tools follow a consistent verb_noun_from_source naming pattern: generate_image_from_text, transform_image_from_encoded, and transform_image_from_file. This pattern clearly indicates the action (generate/transform), the target (image), and the input source (text/encoded/file), creating a predictable and readable naming convention throughout the toolset.
Three tools is a reasonable count for an image generation server, covering the core operations of generating new images and transforming existing ones. However, the scope feels slightly thin as there are no complementary tools for managing generated images (like listing, deleting, or retrieving metadata), which might limit agent workflows in production scenarios.
The server covers basic image generation and transformation operations well, but has notable gaps in image management. There are no tools for listing generated images, deleting files, retrieving image metadata, or batch operations. While the core generative AI functionality is present, the lack of lifecycle management tools creates potential dead ends for agents working with multiple images over time.
Maintenance
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- FlicenseBqualityDmaintenanceEnables image generation and multi-turn editing sessions using the Gemini API within MCP-compatible environments. Users can create, modify, and configure images through natural language commands, supporting features like aspect ratio adjustments and session-based image transformations.5-
- AlicenseAqualityCmaintenanceEnables AI image generation, editing, and upscaling via Google Gemini and Imagen models, supporting dynamic model switching and multiple MCP-compatible clients.12MIT
- FlicenseNot gradedqualityDmaintenanceProvides image generation capabilities using Google's Gemini 2.0 Flash Preview model through the MCP protocol, enabling AI assistants to generate high-quality images from text prompts.-
- FlicenseNot gradedqualityDmaintenanceEnables AI-powered image generation using Google's Gemini 2.5 Flash Image Preview model, supporting text-to-image and image-to-image generation through the MCP interface.-