Skip to main content
Glama

edit_image

Edit images or create new ones from references using AI models; apply masks for local changes and save results with cost details.

Instructions

Edit an image (or create one from reference images) with an OpenRouter image model.

Call list_image_models first to choose a model, picking one whose inputs column says "text+image". Options are checked against the model before anything is sent.

The input images are uploaded to a third-party model provider through OpenRouter: get the user's OK before sending client or project images to third-party providers. Each call costs money on the user's OpenRouter account; the result reports the cost.

All paths must be absolute (or start with "~"); relative paths are rejected. Results are saved next to the first input image (never overwriting it) with a JSON sidecar, unless output_dir is given. The result lists the saved paths, notes, failures, model, provider, time and cost, plus JPEG previews.

Args: prompt: The change to make, e.g. "make it dusk with warm window light". model: Model id from list_image_models. images: Absolute paths of the input images. The first is the image being edited; any others are references. mask_path: Optional absolute path of a mask for the first image: white = change, black = keep. The edit is composited locally, so only the white area changes, but there may be a visible lighting seam at the mask edge where new and old pixels meet; a soft mask edge or mask_feather_px helps. Requires fit="preserve". A masked edit also saves the model's image before the blend as <result>.unmasked.png, so remask_image can redo the blend with a different mask for free. mask_feather_px: Blur radius in pixels for the mask edge; by default it is chosen from the image size. fit: "preserve" (default) returns each result at exactly the first input's pixel size: the server picks the model's closest aspect ratio, then scales and crops back (padding the input first if no ratio is close). With fit="preserve" don't pass aspect_ratio or size. "model" keeps the size the model returns and lets you choose aspect_ratio. n: Number of images, 1 to 10. Above the model's max n, the server splits them into several calls (each billed) and says so in the notes. aspect_ratio: One of the model's aspect ratios (only with fit="model"). resolution: One of the model's resolution tiers. size: Exact pixel size such as "1024x1024" (only with fit="model"). quality: One of the model's quality values. seed: Integer for repeatable results, for models that support it. background: One of the model's background values, e.g. "transparent". output_format: One of the model's output formats, e.g. "png". output_dir: Folder to save into instead of next to the first input. Must be an absolute path or start with "~". filename_prefix: File name stem; defaults to the first input's name. provider_options: Provider passthrough options as a flat dict, e.g. {"moderation": "low"}. Keys must be in the model's passthrough parameters (see get_image_model); the server sends each key to every provider that allows it. Any other key is rejected before anything is spent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nNo
fitNopreserve
seedNo
sizeNo
modelYes
imagesYes
promptYes
qualityNo
mask_pathNo
backgroundNo
output_dirNo
resolutionNo
aspect_ratioNo
output_formatNo
filename_prefixNo
mask_feather_pxNo
provider_optionsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: third-party upload and consent requirement, per-call monetary cost billed to the user's OpenRouter account, absolute-path enforcement with relative paths rejected, non-overwriting save behavior plus JSON sidecar, and the n-above-max splitting into multiple billed calls. It even discloses an edge-case artifact (visible lighting seam at the mask edge) and how to mitigate it, plus the <result>.unmasked.png artifact saved for remask_image.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line purpose, then a prerequisites/behavior paragraph, then a clean per-argument Args block that mirrors the parameter order. It is long, but the length is largely justified by 17 undocumented parameters; a few behavioral details (mask seam, previews) could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by enumerating the returned payload (saved paths, notes, failures, model, provider, time, cost, JPEG previews) and where files land. Given 17 params, cost implications, and a third-party data flow, nothing an agent needs to call this correctly appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 17 parameters, so the description must compensate and it does so for every argument: prompt example, images ordering semantics (first = edited, rest = references), mask white/black convention, mask_feather_px default derivation, fit modes and their interactions, n range and billing split, provider_options key validation. This adds substantial meaning beyond the bare type/default schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (edit an image, or create one from reference images) with the engine named (OpenRouter image model). It differentiates itself from siblings by pointing to list_image_models for model choice, get_image_model for passthrough params, and remask_image for re-blending, so an agent can separate it from generate_image without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit prerequisites (call list_image_models first, pick a model whose inputs column says text+image), explicit consent rule (get the user's OK before sending client/project images to third-party providers), and explicit when-nots (don't pass aspect_ratio or size with fit="preserve"; mask_path requires fit="preserve"). It also routes to remask_image for free re-blends, covering alternatives rather than leaving them to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.