Skip to main content
Glama

MCP Nano Banana

An MCP (Model Context Protocol) server that exposes Google Gemini's image generation capabilities (Nano Banana / Nano Banana Pro) as tools that Claude can use.

Installation

git clone https://github.com/Pgarciapg/mcp-nano-banana.git
cd mcp-nano-banana
npm install
npm run build

Related MCP server: Nanana AI Image Generation Server

Features

  • Text-to-Image Generation: Generate images from text prompts

  • Image Editing: Edit existing images using natural language

  • Image Composition: Combine multiple images into new compositions

  • Two Models Available:

    • nano-banana (gemini-2.5-flash-image): Fast, efficient, 1024px resolution

    • nano-banana-pro (gemini-3-pro-image-preview): Advanced, up to 4K, with thinking mode

Setup

1. Get a Gemini API Key

  1. Go to Google AI Studio

  2. Create or select a project

  3. Generate an API key

2. Set Your API Key

Add your Gemini API key to your shell profile (~/.zshrc or ~/.bashrc):

export GEMINI_API_KEY="your-api-key-here"

Then reload your shell:

source ~/.zshrc

3. Configure Claude Code

Add the server to Claude Code's MCP config. Edit ~/.claude/.mcp.json:

{
  "mcpServers": {
    "gemini-imagen": {
      "command": "node",
      "args": ["/path/to/mcp-nano-banana/dist/index.js"],
      "env": {
        "GEMINI_API_KEY": "${GEMINI_API_KEY}",
        "IMAGEN_OUTPUT_DIR": "/path/to/output/folder"
      }
    }
  }
}

4. Restart Claude Code

After configuring, restart Claude Code to load the new MCP server.

Available Tools

generate_image

Generate an image from a text prompt.

Parameters:

  • prompt (required): Text description of the image to generate

  • model: nano-banana (default) or nano-banana-pro

  • aspect_ratio: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9

  • image_size: 1K, 2K, 4K (only for nano-banana-pro)

  • filename: Optional output filename

edit_image

Edit an existing image using text prompts.

Parameters:

  • prompt (required): Description of the edit to make

  • image_path (required): Path to the input image

  • model: Model to use for editing

  • aspect_ratio: Optional aspect ratio for output

  • image_size: Resolution (only for nano-banana-pro)

  • filename: Optional output filename

compose_images

Combine multiple images into a new composition.

Parameters:

  • prompt (required): How to combine the images

  • image_paths (required): Array of paths to input images

  • model: Model to use (nano-banana-pro recommended)

  • aspect_ratio: Aspect ratio for output

  • image_size: Resolution (only for nano-banana-pro)

  • filename: Optional output filename

Usage Examples

Generate a simple image

"Generate an image of a sunset over mountains with a cabin in the foreground"

Edit an existing image

"Add a wizard hat to the cat in this image" + provide image_path

Combine multiple images

"Put the dress from the first image on the model from the second image" + provide image_paths array

Prompting Tips

  1. Be Descriptive: Describe scenes narratively, not as keyword lists

  2. Specify Style: Use photography terms for photorealistic images (lens type, lighting, angles)

  3. Include Details: Mention colors, textures, lighting, and mood

  4. Use Templates: For specific styles (product photos, logos, etc.), follow proven templates

Environment Variables

  • GEMINI_API_KEY (required): Your Google Gemini API key

  • IMAGEN_OUTPUT_DIR (optional): Directory for generated images (defaults to ./generated-images)

License

MIT

Available Tools

3 tools
compose_imagesA

Combine multiple images into a new composition.

nano-banana supports up to 3 input images. nano-banana-pro supports up to 14 input images (up to 5 humans, 6 objects).

Great for:

  • Product mockups

  • Fashion photos (dress on model)

  • Creative collages

  • Style transfer from multiple references

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescription of how to combine the images. Be specific about which elements from each image to use.
image_pathsYesArray of paths to input images to combine.
modelNoThe model to use. nano-banana-pro recommended for multi-image composition.nano-banana-pro
aspect_ratioNoThe aspect ratio of the output image.1:1
image_sizeNoThe resolution of the output (only for nano-banana-pro).2K
filenameNoOptional filename for the output image.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds valuable context beyond the schema by specifying model-specific constraints (e.g., 'nano-banana supports up to 3 input images') and use-case examples. It does not mention permissions, rate limits, or output format, but provides practical operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: the first sentence states the core purpose, followed by model constraints and use cases. Every sentence earns its place by providing essential information without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, no annotations, no output schema), the description is mostly complete. It covers purpose, constraints, and use cases well, but lacks details on output behavior (e.g., file format, error handling) and does not fully compensate for the absence of annotations and output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description does not add any parameter-specific details beyond what the schema provides (e.g., it doesn't explain 'prompt' or 'image_paths' further). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Combine multiple images into a new composition.' It uses specific verbs ('combine') and resources ('images'), and distinguishes from siblings like 'edit_image' (modify existing) and 'generate_image' (create from scratch) by focusing on multi-image composition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool through the 'Great for:' section, listing specific use cases like product mockups and fashion photos. However, it does not explicitly state when NOT to use it or name alternatives among sibling tools, which prevents a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_imageA

Edit an existing image using text prompts. Supports:

  • Adding/removing elements

  • Style transfer

  • Inpainting (changing specific parts)

  • Combining multiple images

Provide the path to an existing image and describe the changes you want.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesDescription of the edit to make. Be specific about what to change and what to preserve.
image_pathYesPath to the input image file to edit.
modelNoThe model to use for editing.nano-banana
aspect_ratioNoOptional aspect ratio for the output image.
image_sizeNoThe resolution of the output (only for nano-banana-pro).1K
filenameNoOptional filename for the output image.

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions editing capabilities, it doesn't disclose important behavioral traits like whether edits are destructive to the original file, what permissions are needed, rate limits, output format, or error conditions. The description is insufficient for a mutation tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear opening sentence, bullet points for capabilities, and a practical usage instruction. Every sentence earns its place, and information is front-loaded with the core purpose stated first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex image editing tool with 6 parameters, no annotations, and no output schema, the description is incomplete. It doesn't address behavioral aspects like file handling, error cases, or what the tool returns. While the schema covers parameters well, the overall context for proper tool invocation is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3. The description adds some value by mentioning 'path to an existing image' (reinforcing image_path) and 'describe the changes you want' (reinforcing prompt), but doesn't provide additional semantic context beyond what's already well-documented in the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('edit an existing image using text prompts') and distinguishes it from siblings by focusing on editing existing images rather than generating new ones (generate_image) or composing multiple images (compose_images). The bullet points provide concrete examples of editing capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('edit an existing image') and implies alternatives through sibling tool names, but doesn't explicitly state when NOT to use it or directly compare with compose_images/generate_image. The guidance is practical but lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generate an image from a text prompt using Google Gemini's image generation.

Models available:

  • nano-banana (gemini-2.5-flash-image): Fast, efficient, 1024px resolution. Best for high-volume tasks.

  • nano-banana-pro (gemini-3-pro-image-preview): Advanced, up to 4K resolution, with thinking mode. Best for professional assets.

Tips for better results:

  • Describe the scene narratively, don't just list keywords

  • Be specific about lighting, camera angles, and styles

  • Use photography terms for photorealistic images

  • Specify aspect ratio based on your use case

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe text prompt describing the image to generate. Be descriptive and specific.
modelNoThe model to use. nano-banana is faster, nano-banana-pro is higher quality with up to 4K.nano-banana
aspect_ratioNoThe aspect ratio of the generated image.1:1
image_sizeNoThe resolution of the output (only for nano-banana-pro). Options: 1K, 2K, 4K.1K
filenameNoOptional filename for the output image (without extension). If not provided, a timestamp-based name will be used.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context about model capabilities (resolution, speed, thinking mode) and tips for effective prompting, but it does not disclose critical behavioral traits such as rate limits, authentication requirements, cost implications, or what happens on failure (e.g., error handling). The description is informative but incomplete for safe operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and appropriately sized, with a clear purpose statement followed by model details and tips. Every sentence earns its place by providing actionable information, though the tips section could be slightly more concise. It is front-loaded with the core functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (image generation with multiple parameters) and lack of annotations or output schema, the description is moderately complete. It covers purpose, model options, and usage tips, but it lacks details on output format (e.g., file type, return structure), error conditions, and operational constraints like rate limits or costs, which are important for a generative AI tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds marginal value by reinforcing prompt specificity and mentioning aspect ratio in the tips, but it does not provide additional semantic meaning beyond what the schema descriptions already cover (e.g., the schema's prompt description says 'Be descriptive and specific,' mirroring the tips). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate an image from a text prompt using Google Gemini's image generation.' This specifies the verb ('generate'), resource ('image'), and technology ('Google Gemini's image generation'), distinguishing it from sibling tools like 'compose_images' and 'edit_image' which imply different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool (image generation from text prompts) and includes tips for better results, but it does not explicitly state when to use alternatives like 'compose_images' or 'edit_image'. The model descriptions ('Best for high-volume tasks' vs 'Best for professional assets') offer some guidance, but no explicit exclusions or comparisons to siblings are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.1/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: compose_images combines multiple images into a composition, edit_image modifies an existing image with text prompts, and generate_image creates a new image from a text prompt. There is no overlap in functionality, making it easy for an agent to select the correct tool.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (compose_images, edit_image, generate_image), using snake_case throughout. The naming is predictable and readable, with no deviations in style.

Tool Count5/5

With 3 tools, the server is well-scoped for image generation and editing tasks. Each tool serves a unique and essential function in the domain, making the count appropriate and efficient for the server's purpose.

Completeness4/5

The tool set covers core image manipulation workflows: generation, editing, and composition. However, there is a minor gap in operations like deleting or managing images, which might be needed for a full lifecycle, but agents can likely work around this given the server's focus on creation and modification.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Pgarciapg/mcp-nano-banana'

If you have feedback or need assistance with the MCP directory API, please join our Discord server