Gemini Image MCP Server
Provides image generation, modification, analysis, batch generation, and style transfer capabilities using Google's Gemini API.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gemini Image MCP ServerGenerate an image of a sunset over mountains in watercolor style"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini Image MCP Server
A Model Context Protocol (MCP) server that provides image generation and manipulation capabilities using Google's Gemini API. This server integrates with Claude Desktop and other MCP-compatible clients to enable AI-powered image operations.
Features
Image Generation: Create images from text prompts using Gemini 2.5 Flash Image Preview
Image Modification: Modify existing images with natural language instructions
Image Analysis: Analyze images for objects, text, colors, emotions, and comprehensive insights
Batch Generation: Generate multiple images from different prompts in one operation
Style Transfer: Apply artistic styles to existing images
Rate Limiting: Built-in rate limiting to respect API quotas
Safety Settings: Configurable content safety levels
Related MCP server: Gemini Image MCP
Available Tools
1. generateImage
Generate images from text prompts with customizable options.
Parameters:
prompt(required): Text description of the image to generatewidth(optional): Image width in pixelsheight(optional): Image height in pixelsaspectRatio(optional): One of1:1,16:9,9:16,4:3,3:4style(optional): One ofrealistic,artistic,cartoon,sketch,watercolor,oil-paintingquality(optional): One ofstandard,high,ultranumberOfImages(optional): Number of images to generate (default: 1)
2. modifyImage
Modify existing images using natural language instructions.
Parameters:
imageBase64(required): Base64 encoded image datainstructions(required): Text instructions for modificationpreserveStyle(optional): Whether to preserve original artistic stylestrength(optional): Modification strength from 0 to 1
3. analyzeImage
Analyze images and extract various types of information.
Parameters:
imageBase64(required): Base64 encoded image dataanalysisType(optional): One ofdescription,objects,text,colors,emotions,comprehensivedetail(optional): Analysis detail level -low,medium,high
4. batchGenerate
Generate multiple images from different prompts efficiently.
Parameters:
prompts(required): Array of text promptsbaseOptions(optional): Shared options to apply to all generations
5. applyStyleTransfer
Apply artistic styles to existing images.
Parameters:
imageBase64(required): Base64 encoded image datastyle(required): One ofanime,renaissance,impressionist,cyberpunk,minimalist,vintage,futuristicintensity(optional): Style intensity from 0 to 100
Installation
Prerequisites
Node.js 18 or higher
Google Gemini API key (Get one here)
Setup
Clone or download the project:
git clone <repository-url>
cd gemini-image-mcpInstall dependencies:
npm installCreate environment configuration:
cp .env.example .envEdit
.envand add your Gemini API key:
GEMINI_API_KEY=your-gemini-api-key-hereBuild the project:
npm run buildUsage
With Claude Desktop
Add the server to your Claude Desktop configuration file:
On macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
On Windows: %APPDATA%/Claude/claude_desktop_config.json
{
"mcpServers": {
"gemini-image": {
"command": "node",
"args": ["/path/to/gemini-image-mcp/dist/index.js"],
"env": {
"GEMINI_API_KEY": "your-gemini-api-key-here"
}
}
}
}Standalone Usage
You can also run the server directly for testing:
npm startConfiguration Options
Environment variables you can set:
GEMINI_API_KEY(required): Your Google Gemini API keyGEMINI_MODEL(optional): Model to use (default:gemini-2.5-flash-image-preview)SAFETY_LEVEL(optional): Content safety level -LOW,MEDIUM,HIGH,BLOCK_NONE(default:MEDIUM)MAX_REQUESTS_PER_MINUTE(optional): Rate limit (default: 10)
Examples
Once integrated with Claude Desktop, you can use natural language to interact with the tools:
Image Generation
"Generate an image of a sunset over mountains in watercolor style"
Image Modification
"Take this image and add a rainbow in the sky while preserving the original style"
Image Analysis
"Analyze this image and tell me what objects you can detect with confidence scores"
Batch Generation
"Generate 3 different versions of a futuristic cityscape: one cyberpunk style, one minimalist, and one realistic"
Style Transfer
"Apply an impressionist style to this photograph with high intensity"
Development
Running in Development Mode
npm run devBuilding
npm run buildType Checking
npm run typecheckLinting
npm run lintAPI Limitations
Rate limiting is enforced based on your configuration
Image generation may take 10-30 seconds depending on complexity
Maximum image size depends on Gemini API limits
Content safety filters are applied based on your safety level setting
Troubleshooting
Common Issues
"GEMINI_API_KEY environment variable is required"
Ensure you've set the API key in your environment or Claude Desktop config
"Rate limit exceeded"
Wait for the rate limit window to reset or adjust
MAX_REQUESTS_PER_MINUTE
"Image generation failed"
Check your prompt for potentially unsafe content
Verify your API key has proper permissions
Try adjusting the safety level settings
Debug Logging
The server logs errors to stderr. Check the Claude Desktop console or your terminal for detailed error messages.
Contributing
Fork the repository
Create a feature branch
Make your changes
Run tests and linting
Submit a pull request
License
MIT License - see LICENSE file for details
Support
For issues and questions:
Check the troubleshooting section above
Review Claude Desktop MCP documentation
Create an issue in the repository
Available Tools
5 toolsanalyzeImageC
Analyze images and extract information
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | Level of detail for analysis (optional) | |
| imageBase64 | Yes | Base64 encoded image data | |
| analysisType | No | Type of analysis to perform (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. 'Analyze' implies a read-only operation, but it does not explain what the tool returns, whether it has side effects, or any operational constraints. This is a significant gap for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at one sentence with no filler, but it is under-specified. It lacks structure and does not elaborate on the various analysis types or output formats, making it minimally adequate rather than well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description must explain return values and behavior. It fails to describe the output format or mention the analysisType options, leaving the tool incomplete for effective use. The analysisType enum in the schema partially mitigates this, but the description itself is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage with descriptions for all three parameters, including enums for detail and analysisType. The description adds no additional parameter context beyond what the schema already states, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Analyze images and extract information' clearly identifies the tool's function (analyze) and resource (images), distinguishing it from sibling tools like generateImage and modifyImage. However, 'extract information' is vague and does not specify what information is extracted, limiting clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions or specific use cases, offering no context for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
applyStyleTransferC
Apply artistic styles to existing images
| Name | Required | Description | Default |
|---|---|---|---|
| style | Yes | Artistic style to apply | |
| intensity | No | Style intensity from 0 to 100 (optional) | |
| imageBase64 | Yes | Base64 encoded image data |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden of behavioral disclosure. It fails to mention whether the operation mutates the input image, returns a new image, or has side effects. No details on image format constraints, size limits, or error conditions are given, so the agent cannot anticipate the tool's behavior beyond the basic function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loaded with the core action. There is no redundancy or fluff, making it efficient. However, it may be too sparse, but for what it says, it is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three parameters (including an enum and base64 input) and lacks both annotations and an output schema. The one-line description does not explain what the tool returns, any constraints, or behavioral nuances, making it insufficient for an agent to use correctly without additional assumptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage, describing all three parameters with clear definitions. The description adds no extra meaning, such as how intensity interacts with style or the base64 encoding contract. Per the rubric, baseline 3 is appropriate because the schema handles parameter documentation, and the description does not contradict or enhance it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb+resource: 'Apply artistic styles to existing images.' It implies a transformation operation and distinguishes from generation by specifying 'existing images,' but it does not explicitly differentiate from modifyImage, which could overlap. Thus, it's clear but not fully sibling-differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like modifyImage or generateImage. It lacks any context about prerequisites, exclusions, or appropriate scenarios, leaving the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batchGenerateB
Generate multiple images from different prompts
| Name | Required | Description | Default |
|---|---|---|---|
| prompts | Yes | Array of text prompts | |
| baseOptions | No | Base options to apply to all generations (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the basic action and does not mention how the batch is processed, whether failures are handled per-prompt, if output is returned directly or asynchronously, or any side effects. This is a significant gap for a batch operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that effectively communicates the tool's purpose without unnecessary words. It is concise and well-structured, though brevity in other dimensions is penalized separately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the schema being fully documented, the description lacks important contextual information such as what the tool returns, how it handles multiple prompts (parallel vs sequential), and any prerequisites or error behavior. With no output schema and no annotations, this incompleteness is notable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema descriptions cover both parameters (prompts and baseOptions) fully, so the additional value from the description is minimal. The description's 'different prompts' aligns with the prompts array but adds no new syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Generate'), resource ('images'), and scope ('multiple', 'different prompts'), which distinguishes it from the sibling generateImage that implies single-image generation. It is specific and instantly conveys the tool's core function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for batch image generation with multiple prompts, but it does not explicitly state when to prefer this over generateImage or other siblings. There is no mention of exclusions or alternative tools, so the guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateImageC
Generate images from text prompts using Google Gemini
| Name | Required | Description | Default |
|---|---|---|---|
| style | No | Image style (optional) | |
| width | No | Image width in pixels (optional) | |
| height | No | Image height in pixels (optional) | |
| prompt | Yes | Text prompt describing the image to generate | |
| quality | No | Image quality level (optional) | |
| aspectRatio | No | Image aspect ratio (optional) | |
| numberOfImages | No | Number of images to generate (optional, default: 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, but it only mentions the use of Google Gemini. It does not explain output format (e.g., URL, base64), potential side effects (cost, rate limits), or default behaviors, making it inadequate for a generation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no superfluous words, front-loading the core action. It is perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 7 parameters, no output schema, no annotations, and a one-line description. It fails to explain return values, error handling, or typical use cases, and given the existence of sibling tools, more context is needed to ensure correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters. The description adds no additional meaning beyond the schema, justifying the baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates images from text prompts, which is a specific verb-resource combination. However, it does not distinguish this tool from batchGenerate, which also generates images, so it lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives like modifyImage, analyzeImage, or batchGenerate. There are no explicit use cases, exclusions, or comparisons, leaving the agent without context for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
modifyImageB
Modify existing images with text instructions
| Name | Required | Description | Default |
|---|---|---|---|
| strength | No | Modification strength from 0 to 1 (optional) | |
| imageBase64 | Yes | Base64 encoded image data | |
| instructions | Yes | Instructions for modifying the image | |
| preserveStyle | No | Whether to preserve the original style (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for disclosing behavior. It only says 'modify' without mentioning what the tool returns, whether it mutates the input, or any side effects. This is a significant gap for a tool that processes image data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words. It is appropriately brief but could benefit from a bit more detail, yet this dimension focuses on structure and size, which is well-handled.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (image editing with optional parameters) and lack of output schema/annotations, the description is incomplete. It does not explain what the output will be (e.g., a modified image in what format) or how the optional parameters affect behavior, leaving the agent without essential context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema description coverage is 100%, all four parameters are already documented in the schema. The description adds minimal value beyond that, merely echoing the concept of 'text instructions' (matching the 'instructions' parameter) without elaborating on 'strength' or 'preserveStyle'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Modify existing images with text instructions.' It uses a specific verb (modify) and resource (existing images), and contrasts with sibling tools like generateImage (creation) and analyzeImage (reading), making its purpose distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for editing existing images via text prompts, which differentiates it from generation or analysis. However, it does not explicitly state when not to use it or name alternative tools, leaving the guidance to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
analyzeImage - First observed
applyStyleTransfer - First observed
batchGenerate - First observed
generateImage - First observed
modifyImage
TDQS
Scored across 5 tools
Each tool targets a distinct action: generating, modifying, analyzing, batch generating, and style transfer. No overlap in purpose, and descriptions clearly differentiate them.
Most tools follow a verb+Image pattern (generateImage, modifyImage, analyzeImage), but batchGenerate and applyStyleTransfer deviate, introducing inconsistency in prefix/object placement. Still readable and predictable overall.
With 5 tools, the set is well-scoped for the server's purpose. Each tool serves a core image generation or processing function without unnecessary bloat or thinness.
The domain covers generation, modification, analysis, batch operations, and style transfer, providing a comprehensive surface for typical Gemini image workflows. No critical operations appear missing.
Maintenance
Related MCP Connectors
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate and edit images and videos with imageat.
Provides YouCam API for AI image and video editing and generation.
Resize, convert, compress, crop, thumbnail and watermark images from your AI chat.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceEnables AI agents to generate, edit, and analyze images using Google's Gemini image generation models including Nano Banana Pro (gemini-3-pro-image-preview).100 npm17MIT
- FlicenseBqualityDmaintenanceEnables image generation and multi-turn editing sessions using the Gemini API within MCP-compatible environments. Users can create, modify, and configure images through natural language commands, supporting features like aspect ratio adjustments and session-based image transformations.5-
- FlicenseNot gradedqualityCmaintenanceEnables image generation and editing using Google's Gemini models with support for model selection and custom aspect ratios. Users can generate high-quality images or modify existing ones through natural language prompts while controlling specific parameters like quality and dimensions.-
- AlicenseAqualityAmaintenanceEnables AI image generation and editing using Google's Gemini Multimodal Image APIs.61MIT