Skip to main content
Glama

images_generate

Generates a PNG image from a text prompt using Gemini 2.5 Flash Image. Returns a file_id consumable by messages.send(attachments=[...]) and other file-aware tools. Supports up to 12 reference image file_ids for subject-consistent edits and composition (use file IDs from the [ATTACHMENTS] block, files.search, or search.files). Latency: ~8-10s per image. Output: 1024×1024 PNG.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
promptYesText description of the image to generate (3-4000 chars).
aspect_ratioNoOutput aspect ratio.1:1
in_workspaceNoRun this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.
reference_file_idsNoOptional list of up to 3 file_ids whose images should be used as visual references (for edits, subject consistency, or composition). Files must be image MIME types (image/png, image/jpeg, image/webp, image/gif).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / in_workspace
      Added value: +{
      +  "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.",
      +  "type": "integer"
      +}
  2. Added
  3. Removed
  4. First observed

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (readOnly=false, idempotent=false, destructive=false), and the description adds real operational context beyond them: latency (~8-10s per image), output dimensions (1024x1024 PNG), and the return value's downstream use. The only weakness is the misleading 'up to 12 reference file_ids' claim, but that is a schema conflict rather than an annotation conflict.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and return contract, and each sentence carries information. Minor redundancy in stating PNG twice (opening line and 'Output: 1024x1024 PNG') keeps it just short of exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by describing the returned file_id and how to consume it, plus latency and dimensions. The only completeness gap is the inconsistent reference-count claim, which could lead an agent to pass an invalid number of references.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 and the description adds only marginal value (purpose of reference IDs, e.g. subject-consistent edits/composition). It also introduces a factual error: the description says 'up to 12 reference image file_ids' while the schema caps the array at 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific verb (Generates), resource (PNG image), input (text prompt) and underlying model, so an agent can distinguish it from videos_generate and images_search without opening the schema. Nothing tautological, no ambiguity about output type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains how the output is consumed (file_id usable in messages.send(attachments=[...])) and where to source reference IDs ([ATTACHMENTS] block, files.search, search.files), which is genuinely actionable. However it never states when to pick this over images_search or videos_generate, so sibling differentiation for usage is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.