Gemini MCP Server for Claude Desktop
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| generate_imageC | Generate an image using Google's Gemini 2.0 Flash Experimental model (with learned user preferences) |
| gemini-edit-imageC | Edit existing images using Gemini's AI image editing capabilities (with learned user preferences) |
| gemini-advanced-imageC | Generate advanced images with Gemini 2.5 Flash Image: multi-image fusion, character consistency, targeted editing, and template adherence |
| gemini-nano-banana-proB | Generate professional images with Nano Banana Pro (Gemini 3 Pro Image): 4K resolution, up to 14 reference images, advanced text rendering, character consistency, and studio-grade controls |
| gemini-chatC | Chat with Gemini AI for conversations, questions, and general assistance (with learned user preferences) |
| gemini-transcribe-audioB | Transcribe audio files to text using Gemini's multimodal capabilities (with learned user preferences) |
| gemini-code-executeC | Execute Python code using Gemini's built-in code execution sandbox (with learned user preferences) |
| gemini-analyze-videoB | Analyze video files using Gemini's multimodal video understanding capabilities (with learned user preferences) |
| gemini-analyze-imageC | Analyze images using Gemini's multimodal vision capabilities (with learned user preferences) |
| gemini-upload-fileA | Upload files to Gemini File API (up to 2GB) for use in subsequent operations. Files persist for 48 hours. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 10 tools
Most tools have distinct purposes (e.g., analyze vs. generate vs. transcribe), but there is notable overlap between gemini-advanced-image, gemini-nano-banana-pro, and generate_image—all focused on image generation with varying model specifications. This could cause confusion for an agent trying to select the right image generation tool.
Nine of the ten tools follow a consistent gemini-verb-noun pattern (e.g., gemini-analyze-image), which is clear and predictable. However, generate_image deviates from this pattern by omitting the gemini prefix, creating a minor inconsistency in the naming scheme.
With 10 tools, the count is well-scoped for a Gemini AI server, covering key multimodal capabilities like image analysis, video analysis, chat, code execution, and file handling. Each tool appears to serve a specific function without unnecessary duplication, making the set appropriately sized.
The toolset provides broad coverage for interacting with Gemini's multimodal features, including image generation/editing, audio/video analysis, chat, and file uploads. A minor gap is the lack of a dedicated tool for text-based document analysis or summarization, but core workflows are well-supported.