Skip to main content
Glama

Generate or edit an image

generate_image

Turn text prompts into images, or edit and combine supplied photos, returning the result inline or as a URL in MCP clients.

Instructions

Text to image, or an edit / combination when images are given. Returns the image (inline when small) and its URL. The free tier runs studio-image-v1 without a key; other models need TANVO_API_KEY. Credits are charged per run and refunded if it fails.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
waitNoWait for the image (usually under a minute)
modelNoModel id from list_modelsstudio-image-v1
aspectNo1:1, 16:9, 9:16, 4:3, 3:4 … as the model allows (default: the model's)
formatNo
imagesNoPhotos as local file paths or public https URLs. Local files are uploaded for you.
promptYesWhat to make: subject, setting, light and style. For edits, say what to change and what to keep.
save_toNoOptional folder to download the results into (defaults to TANVO_OUTPUT_DIR when set)
resolutionNo1K, 2K or 4K where the model supports it (default: the model's)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint, so the description carries the behavioral burden and does it well: it discloses the return form (inline when small, plus URL), the authentication requirement (TANVO_API_KEY for non-free models), and the billing behavior (credits charged per run, refunded on failure). That is exactly the cost/auth/error context an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero filler, front-loaded with the core capability and mode switch before dropping into auth and billing details. Every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the return-value explanation (inline image plus URL) and covers setup and cost. The one gap is async behavior: `wait` defaults to true, but the description never tells the agent what to do when wait=false (poll get_generation), which is a meaningful omission for a long-running generation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 88%, so the schema already documents prompt, images, model, aspect, resolution, save_to, wait and format. The description adds one genuinely new semantic — that `images` turns the call into an edit/combination — but contributes nothing about `wait`, `save_to` defaults, or aspect/resolution constraints beyond the schema. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause names a specific verb pair (generate/edit) and resource (image), and explicitly states the condition that switches modes: an edit or combination 'when `images` are given.' It is immediately distinguishable from siblings like generate_video, generate_music, and generate_from_app.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the two usage modes via the `images` condition and notes the free-tier model, which is useful routing context. But it never says when to prefer this over siblings (e.g. generate_from_app) nor names list_models as the way to discover valid model ids, and it is silent on the non-blocking path (wait=false, then get_generation).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.