Skip to main content
Glama
skelly-77
by skelly-77

edit_image

Edit images or create new ones from reference images via OpenRouter models, with optional masks and saved results.

Instructions

Edit an image (or create one from reference images) with an OpenRouter image model.

Call list_image_models first to choose a model, picking one whose inputs column says "text+image". Options are checked against the model before anything is sent.

The input images are uploaded to a third-party model provider through OpenRouter: get the user's OK before sending client or project images to third-party providers. Each call costs money on the user's OpenRouter account; the result reports the cost.

All paths must be absolute (or start with "~"); relative paths are rejected. Results are saved next to the first input image (never overwriting it) with a JSON sidecar, unless output_dir is given. The result lists the saved paths, notes, failures, model, provider, time and cost, plus JPEG previews.

Args: prompt: The change to make, e.g. "make it dusk with warm window light". model: Model id from list_image_models. images: Absolute paths of the input images. The first is the image being edited; any others are references. mask_path: Optional absolute path of a mask for the first image: white = change, black = keep. The edit is composited locally, so only the white area changes, but there may be a visible lighting seam at the mask edge where new and old pixels meet; a soft mask edge or mask_feather_px helps. Requires fit="preserve". mask_feather_px: Blur radius in pixels for the mask edge; by default it is chosen from the image size. fit: "preserve" (default) returns each result at exactly the first input's pixel size: the server picks the model's closest aspect ratio, then scales and crops back (padding the input first if no ratio is close). With fit="preserve" don't pass aspect_ratio or size. "model" keeps the size the model returns and lets you choose aspect_ratio. n: Number of images, 1 to 10. Above the model's max n, the server splits them into several calls (each billed) and says so in the notes. aspect_ratio: One of the model's aspect ratios (only with fit="model"). resolution: One of the model's resolution tiers. size: Exact pixel size such as "1024x1024" (only with fit="model"). quality: One of the model's quality values. seed: Integer for repeatable results, for models that support it. background: One of the model's background values, e.g. "transparent". output_format: One of the model's output formats, e.g. "png". output_dir: Folder to save into instead of next to the first input. Must be an absolute path or start with "~". filename_prefix: File name stem; defaults to the first input's name. provider_options: Provider passthrough options as a flat dict, e.g. {"moderation": "low"}. Keys must be in the model's passthrough parameters (see get_image_model); the server sends each key to every provider that allows it. Any other key is rejected before anything is spent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nNo
fitNopreserve
seedNo
sizeNo
modelYes
imagesYes
promptYes
qualityNo
mask_pathNo
backgroundNo
output_dirNo
resolutionNo
aspect_ratioNo
output_formatNo
filename_prefixNo
mask_feather_pxNo
provider_optionsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden and does so: third-party upload path, monetary cost per call on the user's account, absolute-path enforcement, non-overwriting save behavior with JSON sidecar, n splitting into multiple billed calls, provider_options validation before spending, and the mask seam artifact. This is well beyond what the schema declares.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then the critical prerequisites, then args. It is long, but with 17 undocumented parameters the length is largely earned; a few arg entries (e.g. resolution/quality) could be compressed further.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description describes the return payload (saved paths, notes, failures, model, provider, time, cost, JPEG previews) and covers the full parameter surface, consent, cost, and side effects. Nothing material an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 17 parameters, and the description compensates almost completely: white=change/black=keep mask semantics, feather default derivation, fit=preserve crop/pad behavior, aspect_ratio/size only under fit='model', seed repeatability caveat, and array ordering ('first is the image being edited; any others are references').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('edit an image') and broadens scope to 'create one from reference images' with a named backend ('OpenRouter image model'). An agent can tell this apart from generate_image in practice, but the sibling is never named and the create-from-references overlap is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit prerequisites and ordering: 'Call `list_image_models` first to choose a model, picking one whose inputs column says "text+image"', plus a consent gate ('get the user's OK before sending client or project images') and a conditional constraint ('With fit="preserve" don't pass aspect_ratio or size'). It states when-not for several parameters, not just when.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.