Skip to main content
Glama

generate-image

Create Foundry VTT art assets from text prompts with presets for tokens, portraits, icons, props, and illustrations. Returns a finished PNG on disk with file path, dimensions, and estimated spend.

Instructions

Generate one Foundry art asset from a prompt via the Gemini image API. kind picks the model tier, aspect, size, framing text, and post-processing; the result is a finished PNG on disk. READ IT before showing anyone: count limbs per creature, check for duplicated spell effects or props, stray signatures, and reference faces on the wrong figure; obvious flaws are one edit-image call away. Every kind runs on flash by default; tier: "pro" refuses without confirmPro: true and states the cost. Returns the file path, dimensions, and estimated spend.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
kindYesPurpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro.
slugYesKebab-cased into the filename: <kind>-<slug>-<id>.png.
tierNoDefault flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.
promptYesWhat a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it.
footprintNoProps only: grid cells wide x tall, e.g. "2x1" for a table (default "1x1"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: "TC_Anvil 02_2x1.png".
confirmProNoRequired true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it.
referencesNoReference images, attached in this order. Bind each in the prompt by its label or 1-based index ("Image 2 is Morgash").
creatureSizeNoTokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changedv1.1.0
    • addedInput schema / properties / creatureSize
      Added value: +{
      +  "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.",
      +  "enum": [
      +    "medium",
      +    "large"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / footprint
      Added value: +{
      +  "description": "Props only: grid cells wide x tall, e.g. \"2x1\" for a table (default \"1x1\"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: \"TC_Anvil 02_2x1.png\".",
      +  "pattern": "^\\d{1,2}x\\d{1,2}$",
      +  "type": "string"
      +}
    • changedInput schema / properties / kind / description
      Previous value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
    • changedInput schema / properties / kind / enum
      Previous value: -[
      -  "icon",
      -  "token",
      -  "portrait",
      -  "illustration"
      -]New value: +[
      +  "icon",
      +  "token",
      +  "prop",
      +  "portrait",
      +  "illustration"
      +]
    • changedInput schema / properties / references / items / properties / role / description
      Previous value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role."
    • changedInput schema / properties / references / items / properties / role / enum
      Previous value: -[
      -  "character",
      -  "style"
      -]New value: +[
      +  "character",
      +  "style",
      +  "pose"
      +]
  2. Changed12 schema fields changedv1.0.0
    • removedInput schema / properties / batch
      Removed value: -{
      -  "default": 6,
      -  "description": "Draft mode only: images per batch.",
      -  "maximum": 8,
      -  "minimum": 1,
      -  "type": "integer"
      -}
    • addedInput schema / properties / confirmPro
      Added value: +{
      +  "description": "Required true with tier: \"pro\". Offer pro to the owner as an option for portraits and illustrations (\"pro is available for a bit extra\"); never assume it.",
      +  "type": "boolean"
      +}
    • removedInput schema / properties / denoise
      Removed value: -{
      -  "default": 0.7,
      -  "description": "Refine mode only. 0.7 (pinned by test) keeps the scene skeleton in dev style; ~0.55 clones composition but inherits the draft rendering style.",
      -  "maximum": 0.95,
      -  "minimum": 0.3,
      -  "type": "number"
      -}
    • changedInput schema / properties / kind / description
      Previous value: -"Purpose preset — fixes generation and output resolution. No raw dimensions."New value: +"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
    • changedInput schema / properties / kind / enum
      Previous value: -[
      -  "handout",
      -  "scene-background",
      -  "portrait",
      -  "token"
      -]New value: +[
      +  "icon",
      +  "token",
      +  "portrait",
      +  "illustration"
      +]
    • removedInput schema / properties / mode
      Removed value: -{
      -  "default": "draft",
      -  "description": "draft: fast klein batch for curation. final: dev-quality render from the prompt alone, finished at output resolution. refine: dev img2img over sourceImage (a picked draft) — keeps its scene skeleton, re-renders in dev style, finished at output resolution.",
      -  "enum": [
      -    "draft",
      -    "final",
      -    "refine"
      -  ],
      -  "type": "string"
      -}
    • changedInput schema / properties / prompt / description
      Previous value: -"The full image prompt."New value: +"What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it."
    • addedInput schema / properties / references
      Added value: +{
      +  "description": "Reference images, attached in this order. Bind each in the prompt by its label or 1-based index (\"Image 2 is Morgash\").",
      +  "items": {
      +    "properties": {
      +      "label": {
      +        "description": "Short name used to bind the reference in the prompt, e.g. \"Morgash\".",
      +        "type": "string"
      +      },
      +      "path": {
      +        "description": "Absolute path of a PNG/JPEG on disk.",
      +        "minLength": 1,
      +        "type": "string"
      +      },
      +      "role": {
      +        "description": "character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice).",
      +        "enum": [
      +          "character",
      +          "style"
      +        ],
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "path",
      +      "role"
      +    ],
      +    "type": "object"
      +  },
      +  "maxItems": 14,
      +  "type": "array"
      +}
    • removedInput schema / properties / seed
      Removed value: -{
      -  "description": "Fixed seed; random when omitted.",
      -  "minimum": 0,
      -  "type": "integer"
      -}
    • changedInput schema / properties / slug / description
      Previous value: -"Short kebab-case subject name used in output filenames, e.g. \"smugglers-cove\"."New value: +"Kebab-cased into the filename: <kind>-<slug>-<id>.png."
    • removedInput schema / properties / sourceImage
      Removed value: -{
      -  "description": "Refine mode only (required there): absolute path of the picked draft PNG.",
      -  "type": "string"
      -}
    • addedInput schema / properties / tier
      Added value: +{
      +  "description": "Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. \"pro\" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.",
      +  "enum": [
      +    "flash",
      +    "pro"
      +  ],
      +  "type": "string"
      +}
  3. First observedv0.1.0

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool writes a PNG to disk, that pro tier refuses without confirmPro and states cost, that every kind defaults to flash, and that post-processing (alpha cut, framing, plate) is done automatically. It also warns to inspect the result before showing anyone. It does not mention rate limits or failure modes, but the behavioral traits that matter for invocation are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: a one-sentence purpose, a QA warning, a tier/cost note, and a return summary. It front-loads the core purpose and the most important behavioral caveat (read before showing). It is longer than ideal, but every sentence adds operational value; the only minor redundancy is repeating flash default and confirmPro, which also appear in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema and no annotations, the description is quite complete: it covers the output (file path, dimensions, estimated spend), the tier gating, the QA expectation, and the per-kind behavior. It does not describe error cases or what happens when references are invalid, but the schema already documents reference roles and limits. The description is sufficient for an agent to invoke the tool correctly in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema: it explains that kind picks model tier, aspect, size, framing text, and post-processing; it clarifies that icons/tokens get framing/background appended automatically; it explains the pro tier cost and confirmPro requirement; and it gives concrete output dimensions. This goes beyond the schema's field descriptions, so a 4 is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate one Foundry art asset from a prompt via the Gemini image API.' It names the output (finished PNG on disk) and distinguishes the tool from siblings by mentioning edit-image as a follow-up for fixing flaws. The kind parameter further clarifies the five asset types, so an agent can tell this apart from edit-image or cutout-image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool (generate a new asset) and when to use edit-image ('obvious flaws are one edit-image call away'). It also gives usage context for pro tier: 'refuses without confirmPro: true and states the cost,' and instructs to offer pro only as an option to the owner. This is strong routing guidance relative to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.