Skip to main content
Glama

Generate Image

generate_image

Generate or edit images using Google Gemini from prompts, reference photos, or videos. Remove backgrounds for transparent PNG cutouts in one call, with cost and file path returned.

Instructions

Generate or edit images using Google Gemini. Provide just a prompt for text-to-image generation. Add image file paths to edit or use reference images. Set removeBackground to get a transparent PNG cutout in one call (local AI matte; works on any subject, no extra API cost). Returns the saved file path, model used, token counts, and estimated cost. Advanced inputs (video-to-image, thinking depth, image-search grounding): see the README's Advanced Features section.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNoSeed for reproducible generation. Same seed + prompt + model = same image.
modelNoGemini image model ID. Defaults to the configured default (gemini-3.1-flash-lite-image). Validated at request time against the models your API key supports (discovered at startup). Common: gemini-3.1-flash-lite-image (cheapest, 1K), gemini-3.1-flash-image (fast, grounding, 512-4K), gemini-3-pro-image (best quality, up to 4K), gemini-2.5-flash-image (legacy, 1K; shuts down 2026-10-02).
imagesNoFile paths to input/reference images for editing. Omit for text-to-image generation. Per-model reference limits vary (gemini-3.1-flash-lite-image up to 14; others less) — the API enforces.
promptYesText description of the image to generate, or editing instruction when images are provided
videosNoFile paths to input videos (mp4/mov/webm/etc, max 500MB each). The model watches the video and creates a NEW image from what it understood — thumbnails, posters, summary art. Not a frame grabber. gemini-3.1-flash family only; not combinable with sessionId.
filenameNoBase name for the saved file (e.g. 'hero-banner'). Extension added automatically. Duplicates get a version suffix (hero-banner-v2). Omit for auto-generated name.
groundingNoSearch grounding. 'web' = Google Search for real-world accuracy. 'web+image' adds image results (gemini-3.1-flash-image only; response includes searchEntryPointHtml which ToS requires displaying). Not supported on gemini-3.1-flash-lite-image.
outputDirNoDirectory to save the image. Defaults to config file outputDir, OUTPUT_DIR env var, or ~/gemini-images
sessionIdNoContinue a multi-turn edit. Pass the sessionId from a previous response to refine that image across calls — the server keeps the prior turns as context.
subfolderNoSubfolder within the output directory (e.g. 'landing-page'). Created automatically.
resolutionNoImage resolution. Defaults to config value or 1K. 512 only on gemini-3.1-flash-image; 2K/4K on gemini-3.1-flash-image and gemini-3-pro-image; gemini-3.1-flash-lite-image and gemini-2.5-flash-image are 1K.
aspectRatioNoImage aspect ratio (defers to the API — unsupported values are rejected by Gemini). Defaults to config value or 1:1. Current models support: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, plus 1:4, 4:1, 1:8, 8:1 on gemini-3.1-flash-image.
thinkingLevelNoThinking depth (gemini-3.1-flash family). Default MINIMAL = fast/cheap. Use HIGH for text-heavy or diagram/infographic renders.
removeBackgroundNoReturn a transparent PNG cutout in one call. Omit for a normal opaque image. Default mode 'auto' runs a local AI matte (no extra API cost; first use downloads a ~one-time model). Supplying `color` implies chroma and `threshold` implies threshold — these override the 'auto' default.
useSearchGroundingNoLegacy alias for grounding: 'web'. Prefer the grounding parameter.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed7 schema fields changedv0.6.2
    • addedInput schema / properties / grounding
      Added value: +{
      +  "description": "Search grounding. 'web' = Google Search for real-world accuracy. 'web+image' adds image results (gemini-3.1-flash-image only; response includes searchEntryPointHtml which ToS requires displaying). Not supported on gemini-3.1-flash-lite-image.",
      +  "enum": [
      +    "web",
      +    "web+image"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / images / description
      Previous value: -"File paths to input/reference images for editing. Omit for text-to-image generation. Max references vary by model (gemini-3.1-flash-image ~14, gemini-3-pro-image ~11)."New value: +"File paths to input/reference images for editing. Omit for text-to-image generation. Per-model reference limits vary (gemini-3.1-flash-lite-image up to 14; others less) — the API enforces."
    • changedInput schema / properties / model / description
      Previous value: -"Gemini image model ID. Defaults to the configured default (gemini-2.5-flash-image). Validated at request time against the models your API key supports (discovered at startup). Common: gemini-3.1-flash-image (fast, grounding, 512-4K), gemini-3-pro-image (best quality, up to 4K), gemini-2.5-flash-image (cheapest, 1K; shuts down 2026-10-02)."New value: +"Gemini image model ID. Defaults to the configured default (gemini-3.1-flash-lite-image). Validated at request time against the models your API key supports (discovered at startup). Common: gemini-3.1-flash-lite-image (cheapest, 1K), gemini-3.1-flash-image (fast, grounding, 512-4K), gemini-3-pro-image (best quality, up to 4K), gemini-2.5-flash-image (legacy, 1K; shuts down 2026-10-02)."
    • changedInput schema / properties / resolution / description
      Previous value: -"Image resolution. Defaults to config value or 1K. 512 only on gemini-3.1-flash-image; 1K/2K/4K on gemini-3.x image models; gemini-2.5-flash-image is 1K."New value: +"Image resolution. Defaults to config value or 1K. 512 only on gemini-3.1-flash-image; 2K/4K on gemini-3.1-flash-image and gemini-3-pro-image; gemini-3.1-flash-lite-image and gemini-2.5-flash-image are 1K."
    • addedInput schema / properties / thinkingLevel
      Added value: +{
      +  "description": "Thinking depth (gemini-3.1-flash family). Default MINIMAL = fast/cheap. Use HIGH for text-heavy or diagram/infographic renders.",
      +  "enum": [
      +    "MINIMAL",
      +    "HIGH"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / useSearchGrounding / description
      Previous value: -"Enable Google Search grounding for real-world accuracy. Supported on the gemini-3.x image models; the API rejects it on models that don't support it."New value: +"Legacy alias for grounding: 'web'. Prefer the grounding parameter."
    • addedInput schema / properties / videos
      Added value: +{
      +  "description": "File paths to input videos (mp4/mov/webm/etc, max 500MB each). The model watches the video and creates a NEW image from what it understood — thumbnails, posters, summary art. Not a frame grabber. gemini-3.1-flash family only; not combinable with sessionId.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "maxItems": 3,
      +  "type": "array"
      +}
  2. Changed9 schema fields changedv0.5.0
    • changedInput schema / properties / aspectRatio / description
      Previous value: -"Image aspect ratio. Defaults to config value or 1:1"New value: +"Image aspect ratio (defers to the API — unsupported values are rejected by Gemini). Defaults to config value or 1:1. Current models support: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, plus 1:4, 4:1, 1:8, 8:1 on gemini-3.1-flash-image."
    • removedInput schema / properties / aspectRatio / enum
      Removed value: -[
      -  "1:1",
      -  "16:9",
      -  "9:16",
      -  "3:2",
      -  "2:3",
      -  "4:3",
      -  "3:4",
      -  "21:9"
      -]
    • changedInput schema / properties / images / description
      Previous value: -"File paths to input/reference images for editing (max 14). Omit for text-to-image generation"New value: +"File paths to input/reference images for editing. Omit for text-to-image generation. Max references vary by model (gemini-3.1-flash-image ~14, gemini-3-pro-image ~11)."
    • changedInput schema / properties / model / description
      Previous value: -"Gemini model ID. Defaults to gemini-2.5-flash-image. Options: gemini-2.5-flash-image, gemini-3-pro-image-preview, gemini-3.1-flash-image-preview"New value: +"Gemini image model ID. Defaults to the configured default (gemini-2.5-flash-image). Validated at request time against the models your API key supports (discovered at startup). Common: gemini-3.1-flash-image (fast, grounding, 512-4K), gemini-3-pro-image (best quality, up to 4K), gemini-2.5-flash-image (cheapest, 1K; shuts down 2026-10-02)."
    • addedInput schema / properties / removeBackground
      Added value: +{
      +  "description": "Return a transparent PNG cutout in one call. Omit for a normal opaque image. Default mode 'auto' runs a local AI matte (no extra API cost; first use downloads a ~one-time model). Supplying `color` implies chroma and `threshold` implies threshold — these override the 'auto' default.",
      +  "properties": {
      +    "color": {
      +      "description": "Chroma-key target hex (chroma mode only). Default #00FF00.",
      +      "pattern": "^#?[0-9a-fA-F]{6}$",
      +      "type": "string"
      +    },
      +    "mode": {
      +      "description": "How to cut out the background. 'auto' (default) = local AI semantic matte (BiRefNet): best quality, works on ANY subject incl. green/yellow/glass/reflective, no special prompt, no extra API cost. On first use the matte engine ('@huggingface/transformers') auto-installs (a one-time pause; set GEMINI_IMAGE_AUTO_INSTALL=0 to disable, then it falls back with install instructions). 'chroma' = generate on a green screen then HSV-key it (zero-dependency, instant, but can damage green/yellow/reflective subjects — prefer 'auto' for those). 'threshold' = generate on white then remove white (line art / logos).",
      +      "enum": [
      +        "auto",
      +        "chroma",
      +        "threshold"
      +      ],
      +      "type": "string"
      +    },
      +    "threshold": {
      +      "description": "White brightness cutoff 0-255 (threshold mode only). Default 240.",
      +      "maximum": 255,
      +      "minimum": 0,
      +      "type": "integer"
      +    },
      +    "tolerance": {
      +      "description": "Chroma hue match tolerance 0-255 (chroma mode only). Default 80.",
      +      "maximum": 255,
      +      "minimum": 0,
      +      "type": "integer"
      +    }
      +  },
      +  "type": "object"
      +}
    • changedInput schema / properties / resolution / description
      Previous value: -"Image resolution. Defaults to config value or 1K. 2K/4K only on gemini-3-pro and gemini-3.1-flash. gemini-2.5-flash is 1K only."New value: +"Image resolution. Defaults to config value or 1K. 512 only on gemini-3.1-flash-image; 1K/2K/4K on gemini-3.x image models; gemini-2.5-flash-image is 1K."
    • changedInput schema / properties / resolution / enum
      Previous value: -[
      -  "1K",
      -  "2K",
      -  "4K"
      -]New value: +[
      +  "512",
      +  "1K",
      +  "2K",
      +  "4K"
      +]
    • changedInput schema / properties / sessionId / description
      Previous value: -"Continue a multi-turn editing session. Pass the sessionId from a previous response to refine the image iteratively. The server preserves conversation history."New value: +"Continue a multi-turn edit. Pass the sessionId from a previous response to refine that image across calls — the server keeps the prior turns as context."
    • changedInput schema / properties / useSearchGrounding / description
      Previous value: -"Enable Google Search grounding for real-world accuracy. Available on gemini-3.1-flash-image-preview."New value: +"Enable Google Search grounding for real-world accuracy. Supported on the gemini-3.x image models; the API rejects it on models that don't support it."
  3. First observedv0.4.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses return values ('Returns the saved file path, model used, token counts, and estimated cost'), notes the background-removal behavior ('local AI matte; works on any subject, no extra API cost'), and points to the README for advanced inputs. It could add side-effect details like file-writing behavior or the one-time model auto-install, but the description is reasonably transparent for a generation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with no wasted words. The purpose is front-loaded, the primary workflows are covered, the standout feature (removeBackground) is highlighted, return values are summarized, and a pointer to the README handles advanced topics. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 15 parameters and no output schema, the description covers the essentials: purpose, core workflows, a key feature, and return information. The schema fully documents all parameters, so the README pointer for advanced inputs (video-to-image, thinking depth, grounding) is acceptable. The only notable gap is the lack of any mention of the sibling tool process_image, but that falls under usage guidance rather than completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds workflow-level context (prompt-only, editing, removeBackground) but does not add meaning beyond what the schema already provides for individual parameters. It adds value in organizing usage, but not in explaining parameter semantics further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate or edit images using Google Gemini'. This clearly communicates the tool's core function and differentiates it from a generic image utility. However, it does not explicitly distinguish itself from the sibling tool process_image, so it misses the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear, actionable usage context: 'Provide just a prompt for text-to-image generation. Add image file paths to edit or use reference images. Set removeBackground to get a transparent PNG cutout in one call.' This tells an agent how to choose among the main modes of the tool. It does not mention when to prefer process_image instead, so no exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools