Generate Image
generate_imageGenerate or edit images using Google Gemini from prompts, reference photos, or videos. Remove backgrounds for transparent PNG cutouts in one call, with cost and file path returned.
Instructions
Generate or edit images using Google Gemini. Provide just a prompt for text-to-image generation. Add image file paths to edit or use reference images. Set removeBackground to get a transparent PNG cutout in one call (local AI matte; works on any subject, no extra API cost). Returns the saved file path, model used, token counts, and estimated cost. Advanced inputs (video-to-image, thinking depth, image-search grounding): see the README's Advanced Features section.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for reproducible generation. Same seed + prompt + model = same image. | |
| model | No | Gemini image model ID. Defaults to the configured default (gemini-3.1-flash-lite-image). Validated at request time against the models your API key supports (discovered at startup). Common: gemini-3.1-flash-lite-image (cheapest, 1K), gemini-3.1-flash-image (fast, grounding, 512-4K), gemini-3-pro-image (best quality, up to 4K), gemini-2.5-flash-image (legacy, 1K; shuts down 2026-10-02). | |
| images | No | File paths to input/reference images for editing. Omit for text-to-image generation. Per-model reference limits vary (gemini-3.1-flash-lite-image up to 14; others less) — the API enforces. | |
| prompt | Yes | Text description of the image to generate, or editing instruction when images are provided | |
| videos | No | File paths to input videos (mp4/mov/webm/etc, max 500MB each). The model watches the video and creates a NEW image from what it understood — thumbnails, posters, summary art. Not a frame grabber. gemini-3.1-flash family only; not combinable with sessionId. | |
| filename | No | Base name for the saved file (e.g. 'hero-banner'). Extension added automatically. Duplicates get a version suffix (hero-banner-v2). Omit for auto-generated name. | |
| grounding | No | Search grounding. 'web' = Google Search for real-world accuracy. 'web+image' adds image results (gemini-3.1-flash-image only; response includes searchEntryPointHtml which ToS requires displaying). Not supported on gemini-3.1-flash-lite-image. | |
| outputDir | No | Directory to save the image. Defaults to config file outputDir, OUTPUT_DIR env var, or ~/gemini-images | |
| sessionId | No | Continue a multi-turn edit. Pass the sessionId from a previous response to refine that image across calls — the server keeps the prior turns as context. | |
| subfolder | No | Subfolder within the output directory (e.g. 'landing-page'). Created automatically. | |
| resolution | No | Image resolution. Defaults to config value or 1K. 512 only on gemini-3.1-flash-image; 2K/4K on gemini-3.1-flash-image and gemini-3-pro-image; gemini-3.1-flash-lite-image and gemini-2.5-flash-image are 1K. | |
| aspectRatio | No | Image aspect ratio (defers to the API — unsupported values are rejected by Gemini). Defaults to config value or 1:1. Current models support: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, plus 1:4, 4:1, 1:8, 8:1 on gemini-3.1-flash-image. | |
| thinkingLevel | No | Thinking depth (gemini-3.1-flash family). Default MINIMAL = fast/cheap. Use HIGH for text-heavy or diagram/infographic renders. | |
| removeBackground | No | Return a transparent PNG cutout in one call. Omit for a normal opaque image. Default mode 'auto' runs a local AI matte (no extra API cost; first use downloads a ~one-time model). Supplying `color` implies chroma and `threshold` implies threshold — these override the 'auto' default. | |
| useSearchGrounding | No | Legacy alias for grounding: 'web'. Prefer the grounding parameter. |