Skip to main content
Glama

grok_image_edit

Edit images by prompt in two modes: change only named regions from a prior generation, or apply whole-image instructions to any uploaded image. Returns new image; task IDs chain for refinement.

Instructions

Edit an image with Grok Imagine Image 2.0 (4 credits). TWO MODES: (1) Region mode — pass task_id (a prior Grok 2.0 generation) + mask_indexs from grok_segment_map (run it first, free, pick regions by NAME): only those regions change. (2) Whole-image mode (NEW Aug 2026) — pass image_urls (ANY uploaded/external image, e.g. from upload_file) + aspect_ratio + prompt: instruction-based edit of the full image, no segmentation. Returns a new full image; result task_ids chain back into segment/edit for iterative refinement. Downloads to kie/assets/raw/.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
promptYesRegion mode: what the masked region(s) should become plus what to preserve. Whole-image mode: the edit instruction for the full image.
task_idNoRegion mode: source task ID — a Grok Image 2.0 generation (or a previous grok_image_edit result). Mutually exclusive with image_urls.
filenameNoOutput filename. Auto-generated if omitted.
image_urlsNoWhole-image mode: 1-5 public URL(s) of the image(s) to edit — any image, not just Grok generations (upload local files with upload_file first). Mutually exclusive with task_id.
mask_indexsNoRegion mode only: region indices from grok_segment_map (e.g. [1] or [0, 2]). Field name matches kie's API spelling.
aspect_ratioNoWhole-image mode: required output aspect ratio (auto keeps the input's shape).
download_dirNoAbsolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changedv5.3.0
    • addedInput schema / properties / aspect_ratio
      Added value: +{
      +  "description": "Whole-image mode: required output aspect ratio (auto keeps the input's shape).",
      +  "enum": [
      +    "auto",
      +    "1:1",
      +    "2:3",
      +    "3:2",
      +    "16:9",
      +    "9:16"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / image_urls
      Added value: +{
      +  "description": "Whole-image mode: 1-5 public URL(s) of the image(s) to edit — any image, not just Grok generations (upload local files with upload_file first). Mutually exclusive with task_id.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "maxItems": 5,
      +  "type": "array"
      +}
    • changedInput schema / properties / mask_indexs / description
      Previous value: -"Region indices to edit, from grok_segment_map (e.g. [1] or [0, 2]). Field name matches kie's API spelling."New value: +"Region mode only: region indices from grok_segment_map (e.g. [1] or [0, 2]). Field name matches kie's API spelling."
    • changedInput schema / properties / prompt / description
      Previous value: -"What the masked region(s) should become, plus what to preserve (e.g. \"change the background to a sunset beach, keep the apple unchanged\")"New value: +"Region mode: what the masked region(s) should become plus what to preserve. Whole-image mode: the edit instruction for the full image."
    • changedInput schema / properties / task_id / description
      Previous value: -"Source task ID — a Grok Image 2.0 generation (or a previous grok_image_edit result)"New value: +"Region mode: source task ID — a Grok Image 2.0 generation (or a previous grok_image_edit result). Mutually exclusive with image_urls."
    • changedInput schema / required
      Previous value: -[
      -  "task_id",
      -  "prompt",
      -  "mask_indexs"
      -]New value: +[
      +  "prompt"
      +]
  2. Addedv4.8.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does much of it: it states the credit cost (4), the dependency chain on grok_segment_map, chaining semantics for returned task_ids, and the file download location. It does not cover auth requirements, rate limits, or failure behavior, leaving modest gaps for a zero-annotation mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with action, cost, and the two modes as numbered branches; each sentence carries routing or behavioral payload. Density is justified by the two-mode complexity and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the return (a new full image) and how returned task_ids chain into further segment/edit calls. It covers both modes and the prerequisite tool, but could say more about failure modes or what happens when neither mode's parameters are supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter, including mutual exclusivity and the aspect_ratio default. The description adds mode-grouping context that helps interpret the parameters together, but no format or syntax detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Edit an image with Grok Imagine Image 2.0') and explicitly enumerates its two operating modes with the parameters that select each. An agent can distinguish it from siblings like grok_segment_map or generate_image without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly prescribes when to use each mode (region mode = task_id + mask_indexs; whole-image mode = image_urls + aspect_ratio + prompt) and gives an ordering dependency, instructing the agent to run grok_segment_map first and pick regions by NAME. This is a routing-quality guideline, not just context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.