Skip to main content
Glama
Pritish053

Gemini Image MCP Server

by Pritish053

Gemini Image MCP Server

A Model Context Protocol (MCP) server that provides image generation and manipulation capabilities using Google's Gemini API. This server integrates with Claude Desktop and other MCP-compatible clients to enable AI-powered image operations.

Features

  • Image Generation: Create images from text prompts using Gemini 2.5 Flash Image Preview

  • Image Modification: Modify existing images with natural language instructions

  • Image Analysis: Analyze images for objects, text, colors, emotions, and comprehensive insights

  • Batch Generation: Generate multiple images from different prompts in one operation

  • Style Transfer: Apply artistic styles to existing images

  • Rate Limiting: Built-in rate limiting to respect API quotas

  • Safety Settings: Configurable content safety levels

Related MCP server: Gemini Image MCP

Available Tools

1. generateImage

Generate images from text prompts with customizable options.

Parameters:

  • prompt (required): Text description of the image to generate

  • width (optional): Image width in pixels

  • height (optional): Image height in pixels

  • aspectRatio (optional): One of 1:1, 16:9, 9:16, 4:3, 3:4

  • style (optional): One of realistic, artistic, cartoon, sketch, watercolor, oil-painting

  • quality (optional): One of standard, high, ultra

  • numberOfImages (optional): Number of images to generate (default: 1)

2. modifyImage

Modify existing images using natural language instructions.

Parameters:

  • imageBase64 (required): Base64 encoded image data

  • instructions (required): Text instructions for modification

  • preserveStyle (optional): Whether to preserve original artistic style

  • strength (optional): Modification strength from 0 to 1

3. analyzeImage

Analyze images and extract various types of information.

Parameters:

  • imageBase64 (required): Base64 encoded image data

  • analysisType (optional): One of description, objects, text, colors, emotions, comprehensive

  • detail (optional): Analysis detail level - low, medium, high

4. batchGenerate

Generate multiple images from different prompts efficiently.

Parameters:

  • prompts (required): Array of text prompts

  • baseOptions (optional): Shared options to apply to all generations

5. applyStyleTransfer

Apply artistic styles to existing images.

Parameters:

  • imageBase64 (required): Base64 encoded image data

  • style (required): One of anime, renaissance, impressionist, cyberpunk, minimalist, vintage, futuristic

  • intensity (optional): Style intensity from 0 to 100

Installation

Prerequisites

Setup

  1. Clone or download the project:

git clone <repository-url>
cd gemini-image-mcp
  1. Install dependencies:

npm install
  1. Create environment configuration:

cp .env.example .env
  1. Edit .env and add your Gemini API key:

GEMINI_API_KEY=your-gemini-api-key-here
  1. Build the project:

npm run build

Usage

With Claude Desktop

Add the server to your Claude Desktop configuration file:

On macOS: ~/Library/Application Support/Claude/claude_desktop_config.json On Windows: %APPDATA%/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "gemini-image": {
      "command": "node",
      "args": ["/path/to/gemini-image-mcp/dist/index.js"],
      "env": {
        "GEMINI_API_KEY": "your-gemini-api-key-here"
      }
    }
  }
}

Standalone Usage

You can also run the server directly for testing:

npm start

Configuration Options

Environment variables you can set:

  • GEMINI_API_KEY (required): Your Google Gemini API key

  • GEMINI_MODEL (optional): Model to use (default: gemini-2.5-flash-image-preview)

  • SAFETY_LEVEL (optional): Content safety level - LOW, MEDIUM, HIGH, BLOCK_NONE (default: MEDIUM)

  • MAX_REQUESTS_PER_MINUTE (optional): Rate limit (default: 10)

Examples

Once integrated with Claude Desktop, you can use natural language to interact with the tools:

Image Generation

"Generate an image of a sunset over mountains in watercolor style"

Image Modification

"Take this image and add a rainbow in the sky while preserving the original style"

Image Analysis

"Analyze this image and tell me what objects you can detect with confidence scores"

Batch Generation

"Generate 3 different versions of a futuristic cityscape: one cyberpunk style, one minimalist, and one realistic"

Style Transfer

"Apply an impressionist style to this photograph with high intensity"

Development

Running in Development Mode

npm run dev

Building

npm run build

Type Checking

npm run typecheck

Linting

npm run lint

API Limitations

  • Rate limiting is enforced based on your configuration

  • Image generation may take 10-30 seconds depending on complexity

  • Maximum image size depends on Gemini API limits

  • Content safety filters are applied based on your safety level setting

Troubleshooting

Common Issues

  1. "GEMINI_API_KEY environment variable is required"

    • Ensure you've set the API key in your environment or Claude Desktop config

  2. "Rate limit exceeded"

    • Wait for the rate limit window to reset or adjust MAX_REQUESTS_PER_MINUTE

  3. "Image generation failed"

    • Check your prompt for potentially unsafe content

    • Verify your API key has proper permissions

    • Try adjusting the safety level settings

Debug Logging

The server logs errors to stderr. Check the Claude Desktop console or your terminal for detailed error messages.

Contributing

  1. Fork the repository

  2. Create a feature branch

  3. Make your changes

  4. Run tests and linting

  5. Submit a pull request

License

MIT License - see LICENSE file for details

Support

For issues and questions:

  • Check the troubleshooting section above

  • Review Claude Desktop MCP documentation

  • Create an issue in the repository

Available Tools

5 tools
analyzeImageC

Analyze images and extract information

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNoLevel of detail for analysis (optional)
imageBase64YesBase64 encoded image data
analysisTypeNoType of analysis to perform (optional)

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. 'Analyze' implies a read-only operation, but it does not explain what the tool returns, whether it has side effects, or any operational constraints. This is a significant gap for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at one sentence with no filler, but it is under-specified. It lacks structure and does not elaborate on the various analysis types or output formats, making it minimally adequate rather than well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema and no annotations, so the description must explain return values and behavior. It fails to describe the output format or mention the analysisType options, leaving the tool incomplete for effective use. The analysisType enum in the schema partially mitigates this, but the description itself is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage with descriptions for all three parameters, including enums for detail and analysisType. The description adds no additional parameter context beyond what the schema already states, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Analyze images and extract information' clearly identifies the tool's function (analyze) and resource (images), distinguishing it from sibling tools like generateImage and modifyImage. However, 'extract information' is vague and does not specify what information is extracted, limiting clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions or specific use cases, offering no context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

applyStyleTransferC

Apply artistic styles to existing images

ParametersJSON Schema
NameRequiredDescriptionDefault
styleYesArtistic style to apply
intensityNoStyle intensity from 0 to 100 (optional)
imageBase64YesBase64 encoded image data

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears the full burden of behavioral disclosure. It fails to mention whether the operation mutates the input image, returns a new image, or has side effects. No details on image format constraints, size limits, or error conditions are given, so the agent cannot anticipate the tool's behavior beyond the basic function.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, front-loaded with the core action. There is no redundancy or fluff, making it efficient. However, it may be too sparse, but for what it says, it is well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has three parameters (including an enum and base64 input) and lacks both annotations and an output schema. The one-line description does not explain what the tool returns, any constraints, or behavioral nuances, making it insufficient for an agent to use correctly without additional assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage, describing all three parameters with clear definitions. The description adds no extra meaning, such as how intensity interacts with style or the base64 encoding contract. Per the rubric, baseline 3 is appropriate because the schema handles parameter documentation, and the description does not contradict or enhance it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb+resource: 'Apply artistic styles to existing images.' It implies a transformation operation and distinguishes from generation by specifying 'existing images,' but it does not explicitly differentiate from modifyImage, which could overlap. Thus, it's clear but not fully sibling-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like modifyImage or generateImage. It lacks any context about prerequisites, exclusions, or appropriate scenarios, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batchGenerateB

Generate multiple images from different prompts

ParametersJSON Schema
NameRequiredDescriptionDefault
promptsYesArray of text prompts
baseOptionsNoBase options to apply to all generations (optional)

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the basic action and does not mention how the batch is processed, whether failures are handled per-prompt, if output is returned directly or asynchronously, or any side effects. This is a significant gap for a batch operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that effectively communicates the tool's purpose without unnecessary words. It is concise and well-structured, though brevity in other dimensions is penalized separately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the schema being fully documented, the description lacks important contextual information such as what the tool returns, how it handles multiple prompts (parallel vs sequential), and any prerequisites or error behavior. With no output schema and no annotations, this incompleteness is notable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema descriptions cover both parameters (prompts and baseOptions) fully, so the additional value from the description is minimal. The description's 'different prompts' aligns with the prompts array but adds no new syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Generate'), resource ('images'), and scope ('multiple', 'different prompts'), which distinguishes it from the sibling generateImage that implies single-image generation. It is specific and instantly conveys the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for batch image generation with multiple prompts, but it does not explicitly state when to prefer this over generateImage or other siblings. There is no mention of exclusions or alternative tools, so the guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateImageC

Generate images from text prompts using Google Gemini

ParametersJSON Schema
NameRequiredDescriptionDefault
styleNoImage style (optional)
widthNoImage width in pixels (optional)
heightNoImage height in pixels (optional)
promptYesText prompt describing the image to generate
qualityNoImage quality level (optional)
aspectRatioNoImage aspect ratio (optional)
numberOfImagesNoNumber of images to generate (optional, default: 1)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it only mentions the use of Google Gemini. It does not explain output format (e.g., URL, base64), potential side effects (cost, rate limits), or default behaviors, making it inadequate for a generation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no superfluous words, front-loading the core action. It is perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 7 parameters, no output schema, no annotations, and a one-line description. It fails to explain return values, error handling, or typical use cases, and given the existence of sibling tools, more context is needed to ensure correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 7 parameters. The description adds no additional meaning beyond the schema, justifying the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates images from text prompts, which is a specific verb-resource combination. However, it does not distinguish this tool from batchGenerate, which also generates images, so it lacks sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like modifyImage, analyzeImage, or batchGenerate. There are no explicit use cases, exclusions, or comparisons, leaving the agent without context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

modifyImageB

Modify existing images with text instructions

ParametersJSON Schema
NameRequiredDescriptionDefault
strengthNoModification strength from 0 to 1 (optional)
imageBase64YesBase64 encoded image data
instructionsYesInstructions for modifying the image
preserveStyleNoWhether to preserve the original style (optional)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for disclosing behavior. It only says 'modify' without mentioning what the tool returns, whether it mutates the input, or any side effects. This is a significant gap for a tool that processes image data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no wasted words. It is appropriately brief but could benefit from a bit more detail, yet this dimension focuses on structure and size, which is well-handled.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (image editing with optional parameters) and lack of output schema/annotations, the description is incomplete. It does not explain what the output will be (e.g., a modified image in what format) or how the optional parameters affect behavior, leaving the agent without essential context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Since schema description coverage is 100%, all four parameters are already documented in the schema. The description adds minimal value beyond that, merely echoing the concept of 'text instructions' (matching the 'instructions' parameter) without elaborating on 'strength' or 'preserveStyle'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Modify existing images with text instructions.' It uses a specific verb (modify) and resource (existing images), and contrasts with sibling tools like generateImage (creation) and analyzeImage (reading), making its purpose distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for editing existing images via text prompts, which differentiates it from generation or analysis. However, it does not explicitly state when not to use it or name alternative tools, leaving the guidance to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv1.0.0
    • First observedanalyzeImage
    • First observedapplyStyleTransfer
    • First observedbatchGenerate
    • First observedgenerateImage
    • First observedmodifyImage

TDQS

A3.5/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct action: generating, modifying, analyzing, batch generating, and style transfer. No overlap in purpose, and descriptions clearly differentiate them.

Naming Consistency4/5

Most tools follow a verb+Image pattern (generateImage, modifyImage, analyzeImage), but batchGenerate and applyStyleTransfer deviate, introducing inconsistency in prefix/object placement. Still readable and predictable overall.

Tool Count5/5

With 5 tools, the set is well-scoped for the server's purpose. Each tool serves a core image generation or processing function without unnecessary bloat or thinness.

Completeness5/5

The domain covers generation, modification, analysis, batch operations, and style transfer, providing a comprehensive surface for typical Gemini image workflows. No critical operations appear missing.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers