MCP Nano Banana
Exposes Google Gemini's image generation capabilities (Nano Banana models) for text-to-image generation, image editing, and image composition with support for various aspect ratios and resolutions up to 4K.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Nano Bananagenerate a photorealistic image of a cyberpunk street at night with neon signs and rain"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Nano Banana
An MCP (Model Context Protocol) server that exposes Google Gemini's image generation capabilities (Nano Banana / Nano Banana Pro) as tools that Claude can use.
Installation
git clone https://github.com/Pgarciapg/mcp-nano-banana.git
cd mcp-nano-banana
npm install
npm run buildRelated MCP server: Nanana AI Image Generation Server
Features
Text-to-Image Generation: Generate images from text prompts
Image Editing: Edit existing images using natural language
Image Composition: Combine multiple images into new compositions
Two Models Available:
nano-banana(gemini-2.5-flash-image): Fast, efficient, 1024px resolutionnano-banana-pro(gemini-3-pro-image-preview): Advanced, up to 4K, with thinking mode
Setup
1. Get a Gemini API Key
Go to Google AI Studio
Create or select a project
Generate an API key
2. Set Your API Key
Add your Gemini API key to your shell profile (~/.zshrc or ~/.bashrc):
export GEMINI_API_KEY="your-api-key-here"Then reload your shell:
source ~/.zshrc3. Configure Claude Code
Add the server to Claude Code's MCP config. Edit ~/.claude/.mcp.json:
{
"mcpServers": {
"gemini-imagen": {
"command": "node",
"args": ["/path/to/mcp-nano-banana/dist/index.js"],
"env": {
"GEMINI_API_KEY": "${GEMINI_API_KEY}",
"IMAGEN_OUTPUT_DIR": "/path/to/output/folder"
}
}
}
}4. Restart Claude Code
After configuring, restart Claude Code to load the new MCP server.
Available Tools
generate_image
Generate an image from a text prompt.
Parameters:
prompt(required): Text description of the image to generatemodel:nano-banana(default) ornano-banana-proaspect_ratio:1:1,2:3,3:2,3:4,4:3,4:5,5:4,9:16,16:9,21:9image_size:1K,2K,4K(only for nano-banana-pro)filename: Optional output filename
edit_image
Edit an existing image using text prompts.
Parameters:
prompt(required): Description of the edit to makeimage_path(required): Path to the input imagemodel: Model to use for editingaspect_ratio: Optional aspect ratio for outputimage_size: Resolution (only for nano-banana-pro)filename: Optional output filename
compose_images
Combine multiple images into a new composition.
Parameters:
prompt(required): How to combine the imagesimage_paths(required): Array of paths to input imagesmodel: Model to use (nano-banana-pro recommended)aspect_ratio: Aspect ratio for outputimage_size: Resolution (only for nano-banana-pro)filename: Optional output filename
Usage Examples
Generate a simple image
"Generate an image of a sunset over mountains with a cabin in the foreground"Edit an existing image
"Add a wizard hat to the cat in this image" + provide image_pathCombine multiple images
"Put the dress from the first image on the model from the second image" + provide image_paths arrayPrompting Tips
Be Descriptive: Describe scenes narratively, not as keyword lists
Specify Style: Use photography terms for photorealistic images (lens type, lighting, angles)
Include Details: Mention colors, textures, lighting, and mood
Use Templates: For specific styles (product photos, logos, etc.), follow proven templates
Environment Variables
GEMINI_API_KEY(required): Your Google Gemini API keyIMAGEN_OUTPUT_DIR(optional): Directory for generated images (defaults to./generated-images)
License
MIT
Available Tools
3 toolscompose_imagesA
Combine multiple images into a new composition.
nano-banana supports up to 3 input images. nano-banana-pro supports up to 14 input images (up to 5 humans, 6 objects).
Great for:
Product mockups
Fashion photos (dress on model)
Creative collages
Style transfer from multiple references
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Description of how to combine the images. Be specific about which elements from each image to use. | |
| image_paths | Yes | Array of paths to input images to combine. | |
| model | No | The model to use. nano-banana-pro recommended for multi-image composition. | nano-banana-pro |
| aspect_ratio | No | The aspect ratio of the output image. | 1:1 |
| image_size | No | The resolution of the output (only for nano-banana-pro). | 2K |
| filename | No | Optional filename for the output image. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds valuable context beyond the schema by specifying model-specific constraints (e.g., 'nano-banana supports up to 3 input images') and use-case examples. It does not mention permissions, rate limits, or output format, but provides practical operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded: the first sentence states the core purpose, followed by model constraints and use cases. Every sentence earns its place by providing essential information without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, no annotations, no output schema), the description is mostly complete. It covers purpose, constraints, and use cases well, but lacks details on output behavior (e.g., file format, error handling) and does not fully compensate for the absence of annotations and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description does not add any parameter-specific details beyond what the schema provides (e.g., it doesn't explain 'prompt' or 'image_paths' further). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Combine multiple images into a new composition.' It uses specific verbs ('combine') and resources ('images'), and distinguishes from siblings like 'edit_image' (modify existing) and 'generate_image' (create from scratch) by focusing on multi-image composition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool through the 'Great for:' section, listing specific use cases like product mockups and fashion photos. However, it does not explicitly state when NOT to use it or name alternatives among sibling tools, which prevents a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageA
Edit an existing image using text prompts. Supports:
Adding/removing elements
Style transfer
Inpainting (changing specific parts)
Combining multiple images
Provide the path to an existing image and describe the changes you want.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Description of the edit to make. Be specific about what to change and what to preserve. | |
| image_path | Yes | Path to the input image file to edit. | |
| model | No | The model to use for editing. | nano-banana |
| aspect_ratio | No | Optional aspect ratio for the output image. | |
| image_size | No | The resolution of the output (only for nano-banana-pro). | 1K |
| filename | No | Optional filename for the output image. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions editing capabilities, it doesn't disclose important behavioral traits like whether edits are destructive to the original file, what permissions are needed, rate limits, output format, or error conditions. The description is insufficient for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear opening sentence, bullet points for capabilities, and a practical usage instruction. Every sentence earns its place, and information is front-loaded with the core purpose stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex image editing tool with 6 parameters, no annotations, and no output schema, the description is incomplete. It doesn't address behavioral aspects like file handling, error cases, or what the tool returns. While the schema covers parameters well, the overall context for proper tool invocation is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description adds some value by mentioning 'path to an existing image' (reinforcing image_path) and 'describe the changes you want' (reinforcing prompt), but doesn't provide additional semantic context beyond what's already well-documented in the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('edit an existing image using text prompts') and distinguishes it from siblings by focusing on editing existing images rather than generating new ones (generate_image) or composing multiple images (compose_images). The bullet points provide concrete examples of editing capabilities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('edit an existing image') and implies alternatives through sibling tool names, but doesn't explicitly state when NOT to use it or directly compare with compose_images/generate_image. The guidance is practical but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image from a text prompt using Google Gemini's image generation.
Models available:
nano-banana (gemini-2.5-flash-image): Fast, efficient, 1024px resolution. Best for high-volume tasks.
nano-banana-pro (gemini-3-pro-image-preview): Advanced, up to 4K resolution, with thinking mode. Best for professional assets.
Tips for better results:
Describe the scene narratively, don't just list keywords
Be specific about lighting, camera angles, and styles
Use photography terms for photorealistic images
Specify aspect ratio based on your use case
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The text prompt describing the image to generate. Be descriptive and specific. | |
| model | No | The model to use. nano-banana is faster, nano-banana-pro is higher quality with up to 4K. | nano-banana |
| aspect_ratio | No | The aspect ratio of the generated image. | 1:1 |
| image_size | No | The resolution of the output (only for nano-banana-pro). Options: 1K, 2K, 4K. | 1K |
| filename | No | Optional filename for the output image (without extension). If not provided, a timestamp-based name will be used. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context about model capabilities (resolution, speed, thinking mode) and tips for effective prompting, but it does not disclose critical behavioral traits such as rate limits, authentication requirements, cost implications, or what happens on failure (e.g., error handling). The description is informative but incomplete for safe operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and appropriately sized, with a clear purpose statement followed by model details and tips. Every sentence earns its place by providing actionable information, though the tips section could be slightly more concise. It is front-loaded with the core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (image generation with multiple parameters) and lack of annotations or output schema, the description is moderately complete. It covers purpose, model options, and usage tips, but it lacks details on output format (e.g., file type, return structure), error conditions, and operational constraints like rate limits or costs, which are important for a generative AI tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds marginal value by reinforcing prompt specificity and mentioning aspect ratio in the tips, but it does not provide additional semantic meaning beyond what the schema descriptions already cover (e.g., the schema's prompt description says 'Be descriptive and specific,' mirroring the tips). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate an image from a text prompt using Google Gemini's image generation.' This specifies the verb ('generate'), resource ('image'), and technology ('Google Gemini's image generation'), distinguishing it from sibling tools like 'compose_images' and 'edit_image' which imply different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (image generation from text prompts) and includes tips for better results, but it does not explicitly state when to use alternatives like 'compose_images' or 'edit_image'. The model descriptions ('Best for high-volume tasks' vs 'Best for professional assets') offer some guidance, but no explicit exclusions or comparisons to siblings are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: compose_images combines multiple images into a composition, edit_image modifies an existing image with text prompts, and generate_image creates a new image from a text prompt. There is no overlap in functionality, making it easy for an agent to select the correct tool.
All tool names follow a consistent verb_noun pattern (compose_images, edit_image, generate_image), using snake_case throughout. The naming is predictable and readable, with no deviations in style.
With 3 tools, the server is well-scoped for image generation and editing tasks. Each tool serves a unique and essential function in the domain, making the count appropriate and efficient for the server's purpose.
The tool set covers core image manipulation workflows: generation, editing, and composition. However, there is a minor gap in operations like deleting or managing images, which might be needed for a full lifecycle, but agents can likely work around this given the server's focus on creation and modification.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
Generate images, video & speech with Nano Banana, Veo, Omni and Gemini TTS. Pay as you go.
Omni Flash and Veo video, Nano Banana images on Google Flow, from any MCP client
Turn Claude into a creative studio: DNA-locked characters, images, video, voiceover — 55 tools.
Related MCP Servers
- AlicenseBqualityCmaintenanceEnables Claude Desktop to generate text and analyze images using Google's Gemini Pro API. Provides seamless integration between Claude and Gemini's AI capabilities through natural language commands.2MIT
- AlicenseAqualityFmaintenanceEnables AI assistants to generate images from text prompts and transform existing images using Google Gemini's nano banana model through the Nanana AI service. Supports both text-to-image generation and image-to-image transformation capabilities.210310MIT
- AlicenseAqualityCmaintenanceEnables Claude and other AI assistants to generate high-quality images up to 4K resolution using Google's Gemini image models, with support for flexible aspect ratios, natural language editing, and Google Search grounding for accurate results.412MIT
- FlicenseNot gradedqualityDmaintenanceEnables image generation and prompt enhancement within Claude.ai by leveraging Google Gemini models. It allows users to create visual content in various styles like photorealistic and 3D render directly through natural language.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Pgarciapg/mcp-nano-banana'
If you have feedback or need assistance with the MCP directory API, please join our Discord server