gemini-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| GEMINI_API_KEY | Yes | Your Google Gemini API key | |
| GEMINI_INPUT_DIR | No | Directory to resolve bare input-image filenames against | |
| GEMINI_OUTPUT_DIR | No | Default directory for generated images | current working directory |
| GEMINI_TIMEOUT_MS | No | Upstream request timeout in ms | 60000 |
| GEMINI_IMAGE_MODEL | No | Override the default image model | gemini-3.1-flash-image |
| GEMINI_HEARTBEAT_MS | No | Progress-notification cadence in ms | 10000 |
| GEMINI_CHAIN_RETRY_MS | No | How long to wait out interactions-store lag when a chained call 404s | 120000 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| gemini_list_modelsA | List the Gemini image-generation models available to your API key (Nano Banana / Nano Banana Pro family), and the current default model. |
| gemini_healthcheckA | Resolves the credential the way real tools do, then makes one authenticated request to generativelanguage.googleapis.com. Reports which source supplied the credential, whether generativelanguage.googleapis.com accepted it, the round-trip time, and a plain-English hint distinguishing 'no credential' from 'credential rejected' from 'a generativelanguage.googleapis.com-side problem'. Read-only; never returns the credential itself. Call this when a real tool fails and you want to know which hop broke. |
| gemini_image_generateA | Generate image(s) from a text prompt with a Gemini image model (Nano Banana / Nano Banana Pro). If the result will likely be refined iteratively, prefer gemini_interact (multi-turn) as the entry point. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE). |
| gemini_image_editA | Edit or compose images: provide one or more input images (paths or base64), plus a text instruction. For a SERIES of successive edits to the same image, prefer gemini_interact (multi-turn) — it keeps edit context and avoids re-processing the full image each round; use gemini_image_edit for one-off edits or composing multiple distinct inputs. Gemini over-preserves the input; there is no edit-strength control — for large structural changes, reroll with a different |
| gemini_image_setA | Generate a consistent SET of images: a master image from master_prompt, then one image per scene that references the master so the subject/style stays consistent. Provide |
| gemini_interactA | Preferred tool for iterative refinement of a single image — multi-turn generation and editing via the Interactions API. To refine, pass the returned interaction |
| gemini_video_generateA | Generate a short video (~10s) via the Gemini omni model: text→video, image→video / reference→video (supply reference image[s]), interpolate between two stills (pass first frame then last frame as images), or continue a prior video (task: "edit" or "extend" + previous_interaction_id / continue_last; extensions add ~3-10s each, to ~40s total). Cost scales with |
| gemini_music_generateA | Generate music from a text prompt (mood, genre, instruments, structure, or lyrics inline) via a Lyria model: lyria-3-clip-preview (30s instrumental clip, default, cheapest), lyria-3.5 (full-length song with vocals) or lyria-3-pro-preview (longer-form). Output is MP3, written to disk (or returned inline). Single-turn: a track cannot be refined by a follow-up call, so put the whole brief in the prompt. Runs long — give it a |
| gemini_get_resultA | Retrieve a generation started with |
| gemini_token_usageA | Token usage for this session so far — what every generation has cost in tokens, added up. Call it before and after a workflow and subtract to get that workflow's cost; call it after a single generation for that call's. Reports tokens AND an estimated USD cost, priced per call against each call's own model and stamped with the date its rates were read (override with GEMINI_RATE_CARD). Note there is no account-balance endpoint to query — Google Cloud is post-paid and its billing data lags by hours — so this is the accurate way to attribute spend to a call. |
| gemini_upload_fileA | Upload a file — an image, reference photo, picture, screenshot, video or audio clip — to the Gemini Files API ONCE, and get back a reusable |
| gemini_list_filesA | List files, images and photos currently uploaded to the Gemini Files API under this API key, with their reusable |
| gemini_delete_fileA | Delete an uploaded file, image or photo (by file_uri) from the Gemini Files API before its ~48h expiry. Any tool call still referencing it will then fail with a generic 404, so delete only references you are finished with. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 13 tools
The four image-related tools (gemini_image_generate, gemini_image_edit, gemini_image_set, gemini_interact) have overlapping purposes, though the detailed descriptions with cross-references help guide selection. The other tools are clearly distinct, so the set is not highly ambiguous, but agents may still struggle to choose the right image tool at first.
All tools share a consistent 'gemini_' prefix and snake_case, but the word order is mixed: some use verb_noun (list_models, upload_file), others noun_verb (image_generate, music_generate), and a few are compound nouns or bare verbs (healthcheck, interact). This mixed convention is readable but not uniform.
13 tools is well within the ideal 3-15 range for a media generation server. Each tool covers a distinct part of the workflow: generation, editing, async retrieval, file management, model listing, health check, and usage tracking. No tool feels redundant or missing.
The tool surface covers the full lifecycle for media generation: create (image/video/music), edit, refine, set generation, async results, file upload/list/delete, model listing, and cost tracking. There are no obvious dead ends or critical missing operations for the stated purpose.