Skip to main content
Glama

Generate an image with Gemini

gemini_generate_image

Generates images from text prompts via Gemini web app, saving full-resolution results to disk and returning file paths. Continue chat for iterative edits.

Instructions

Sends a prompt to gemini.google.com in a headless Chrome signed in with the local Google profile, waits for the generated image(s), saves them to disk and returns the file paths. conversation="continue" keeps the existing chat thread (so you can iterate: "make it darker"); conversation="new" starts a fresh chat.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
promptYesThe image prompt, or a follow-up edit when continuing.
save_dirNoDirectory for the saved images. Default: /root/.gemini-image-mcp/images
conversationNo"new" starts a fresh Gemini chat; "continue" stays in the current thread.continue
timeout_secondsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the execution environment (headless Chrome signed in with the local Google profile), the auth/profile requirement, the blocking wait for generated images, disk persistence, and the return shape (file paths). It does not mention failure modes, timeouts, or rate limiting, which would be needed for a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no padding. The first front-loads the mechanism and output behavior; the second explains the conversation toggle with a concrete iteration example. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no output schema, the description covers what the tool does, the auth/environment dependency, and the return value (file paths). It is nearly complete, with only timeout behavior and error handling left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the baseline is 3, but the description adds real meaning: it clarifies the conversation enum with concrete iteration semantics and the "make it darker" example, and confirms output goes to disk as file paths. The prompt's follow-up-edit behavior is also reinforced. timeout_seconds and save_dir get no added context, keeping it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (generate images via gemini.google.com) and explains the mechanism: headless Chrome with the local Google profile, saving images to disk and returning file paths. It does not name the sibling tools (gemini_new_conversation, gemini_session_status) or explicitly contrast with them, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the conversation parameter semantics: "continue" for iteration ("make it darker") and "new" for a fresh chat. However, there is no explicit guidance on when to choose this tool versus gemini_new_conversation, nor any prerequisites or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.