Skip to main content
Glama

Create Image

create_image

Generate an image for a social post, thumbnail or ad creative. Models include GPT Image 2.5 Flare / Sunburst, GPT Image 2, Nano Banana Pro / 2 / 2 Lite and Seedream 5.0 Pro — see list_models for each model's inputs. Attaching images in inputs edits / uses them as references. Priced live (get_price). Returns generation.id — poll with wait_for_image. Renders a live preview in app-capable hosts.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoMode key from list_models. Omit to pick it from the attached media.
modelNoModel key from list_models (category "image"). Default: nano-banana-2.
inputsNoWaveSpeed input fields for the chosen model and mode, exactly as list_models shows them (e.g. aspect_ratio, resolution, duration, generate_audio, image, last_image, reference_images). Media fields take URLs; files hosted elsewhere are imported into the user's library automatically. Omitted fields use the model's defaults.
promptYesWhat to generate
maxCreditsNoRefuse to start if the live price is above this. Pass the credits get_price returned.
idempotencyKeyNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the async lifecycle (returns generation.id, poll with wait_for_image), live pricing and credit refusal via maxCredits, automatic import of externally hosted media, and live preview in app-capable hosts. It omits auth requirements, rate limits, and idempotency behavior, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by tight, information-dense clauses covering models, inputs, pricing, and polling with no filler. The model enumeration is slightly long, but every sentence carries actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, nested-input tool with no output schema and no annotations, the description covers the full call lifecycle: model selection, input construction, cost guardrails, return value, and follow-up polling. Only permission/prerequisite context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3, and the description adds genuine meaning beyond it: attaching `images` inside `inputs` means edit/reference, `mode`/`model` keys come from list_models, and omitted fields fall back to model defaults. This enriches the nested `inputs` object that the schema only partly documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate an image') and names concrete use cases (social post, thumbnail, ad creative). It is clear what the tool does, but it does not explicitly differentiate itself from core siblings like create_thumbnail or create_image_and_schedule, which overlap with the stated use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing for the workflow: see list_models for per-model inputs, get_price for live pricing, and wait_for_image to poll. It also explains that attaching `images` in inputs switches to an edit/reference flow. It stops short of stating when NOT to use it versus the scheduling/thumbnail variants.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources