Skip to main content
Glama

fvtt-mcp-imagegen

An MCP server that lets Claude make art for your Foundry VTT table using Google's Gemini image models (Nano Banana). Four tools, one API key, no GPU.

Claude Code drives it. The server is registered as imagegen, and its illustration-builder skill grounds every piece in what your world already says and shows. Claude writes the prompt, the server renders it, Claude looks at the result and fixes what is wrong, and the finished file lands on disk ready to upload into Foundry.

What you can make

  • Item, spell, and feature icons. Twenty icons from one style line come back as one matching set. About 7 cents each.

  • Tokens. Top-down, full-body, cut to transparency, centred on a square so Foundry's scale 1.0 is right: 512 px, or 1024 px for Large and bigger creatures. Humanoids face up at the camera; beasts are seen along the back. No shadow, and nothing ever clips off the edge: a render with a sword or wing cut by the frame is redone, and refused if it clips twice. Hand it one of your existing tokens as a style reference and new ones match the angle and look. About 7 cents each.

  • Token refreshes. Point it at an old, low-res token and it repaints it at higher quality while keeping the pose, silhouette, design, and colours your players know. It can also turn one creature into another in the same pose (a rat's token into a squirrel).

  • Token redresses. Start from a token you already have as the prototype and change who it is: "make this a Sharran cleric, black armour, a morningstar instead of the club". The pose, camera angle, and cut stay; the costume, gear, and figure change to match the actor it is now for. One stock token becomes the specific NPC in your world without a word of prompt about framing.

  • Token restyles. The same trick as battlemaps, for a whole bestiary. Tokens gathered from different packs (Immortal Nights, Forgotten Adventures, a few you drew yourself) do not match on one map. Hand it one token as the style reference and run the rest through, and they come back in one consistent look, each still recognisably the creature it was.

  • Props. Map dressing placed as tiles: furniture, barrels, trees, anvils. Seen straight down, object only, cut to transparency at the tile size (300 px per grid cell, from a footprint like "2x1"). Refreshing an existing prop returns it at the file's exact pixel size in the same spot on its canvas, so it replaces the old one one for one.

  • Battlemap restyles. Repaint a map you bought (Tom Cartos, Mad Cartographer, anything) in one consistent painted style, with every wall, door, and object exactly where it was, so the walls and lights already traced over it in Foundry still fit. The result lands on the source's own pixel grid. An old low-res map comes back at a whole-number upscale (a 1125×1500 map at 3375×4500). Every render is checked against the original, and one whose layout moved is redone once, then refused. Maps ship as WebP, about 4 MB where a PNG would be 40 to 70. About 15 cents a map.

  • Overland maps. A regional or world map from a sourcebook or a sketch, repainted in the house style with the geography locked and every label, marker, road, compass rose, and scale bar painted out. Names go back on afterwards by script, spelled right and movable, instead of being left to the model. A busy map takes two or three passes to come clean; each is about 15 cents.

  • Portraits. Actor sheet art at 3:4. Hand it a previous portrait or two as style references and the new one matches your table's look.

  • Illustrations. Player handouts and scene splashes at 2560×1600. Hand it your party's portraits as character references and they keep their faces in group scenes.

  • Illustrated session records. With fvtt-mcp-sessionscribe alongside, the recap and GM notes it writes after a session come back illustrated: Claude picks the moments from the record, checks where each one happened against the transcript, attaches the portraits and tokens of everyone in the scene so the same faces recur from one week to the next, and drops the finished plates into the journal and the recap. The worked example is the concluded Greenrest campaign: every recap carries eight or nine plates, and the same shelf of references went on to illustrate the campaign's book.

  • Edits. Change one thing about an existing image and keep the rest: swap a weapon, recolor a cloak, add a scar, fix an extra limb.

  • Cutouts. Knock the background off any token image you already have.

Related MCP server: FoundryVTT MCP Server

The tools

tool

what it does

generate-image

Render one asset from a prompt. kind is icon, token, prop, portrait, or illustration; it picks the model, aspect, size, framing, and post-processing for you. Optional references (character, style, or pose) and, for tokens, creatureSize (medium or large); for props, footprint ("2x1").

edit-image

Apply one instruction to an existing image and keep everything else. Token edits are prompted the way you would type in the Gemini app and re-cut automatically. Takes creatureSize too. kind: "battlemap" restyles a bought map with its layout locked (edit only; generate-image refuses it, since a map painted from words has no walls). kind: "overland" does the same for a regional map and paints the lettering out.

cutout-image

Cut a token's background to alpha and deliver it on a square canvas.

imagegen-status

Key present, models reachable, estimated spend this session.

Every call returns the file path, the pixel size, and an estimated cost.

Models and cost

Everything runs on Nano Banana 2.1 (gemini-nano-banana-2.1) by default: roughly 3 cents for an icon or token, 5 cents for a portrait, 11 cents for an illustration.

Nano Banana Pro (Gemini 3 Pro Image) is available for about double at 4K and four times at 1K. It is stronger on crowded multi-figure scenes and images with legible text. Claude will not use it unless you say so: a Pro call refuses without an explicit confirm and tells you the price first.

Prices are Google's published per-image rates and may change. The server keeps a running estimate; your actual bill is in the Google Cloud console.

Requirements

  • Node.js 22 or newer.

  • A Gemini API key from Google AI Studio with billing enabled. Prepaid credit with auto-reload off is a sensible ceiling.

  • Python 3 with Pillow and numpy for the token cutout. rembg is optional and adds an AI matte fallback for busy backgrounds (first use downloads a ~176 MB model).

Install

git clone https://github.com/Txpple/fvtt-mcp-imagegen.git
cd fvtt-mcp-imagegen
npm install
npm run build
cp .env.example .env

Put your key in .env:

GEMINI_API_KEY=your-key
IMAGEGEN_OUTPUT_DIR=C:\path\where\renders\should\land

Register the server with Claude Code (user scope, so it is available in every project), then restart Claude Code:

claude mcp add -s user imagegen -- node /absolute/path/to/fvtt-mcp-imagegen/dist/index.js

Or copy .mcp.json.example and set absolute paths.

Using it

Ask Claude for what you want in table terms:

  • "Make icons for these six items."

  • "This goblin needs a token; use my existing orc token as the style reference."

  • "Illustrate the party arriving at the ruined mill at dusk; here are their portraits."

  • "Change this token's cloak to forest green."

  • "This old token looks rough; give it an updated painterly pass."

  • "Make this bandit token a Sharran cleric: black armour, morningstar, keep the pose."

  • "These twelve tokens come from three different packs; restyle them all to match this one."

  • "Illustrate last night's recap; the party's portraits are in this folder."

  • "Cut the background off this token."

Claude reads every render before showing it to you and fixes obvious flaws (an extra limb, a duplicated spell effect) with one edit. Files are named <kind>-<slug>-<id>.png so they drop straight into a Foundry asset folder.

With a Foundry MCP server: art grounded in your world

This server only makes pictures. Pair it with a Foundry MCP server such as fvtt-mcp-dnd5e and Claude can read your world before it prompts and put the result back when it is done. Then you can ask for things like:

  • "Make a new token for the dragon in the Wyrmwood." Claude pulls the actor's stat block and bio, opens the token your world already uses for a similar creature so the angle and line style match, renders, cuts to alpha, and can assign it to the actor.

  • "Illustrate the party walking into the dragon's lair for the first time." Claude screenshots the battlemap, finds where the dragon's token is placed, reads the plot notes for the room, attaches the party's portraits so the faces hold, and paints the view from the doors down the hall to the dais.

  • "Illustrate three cool moments from the last few sessions." Claude reads the session recaps and GM notes, picks the scenes, checks which actors were present and what they were carrying that night, and renders each one.

  • "Give this actor a portrait." Claude reads the bio, looks at the existing token so the hair and skin match canon, renders at 3:4, and can set it as the sheet portrait.

  • "Icons for every item in this compendium folder." One shared style line, one call per item, uploaded as a set.

The handoff is files on disk: this server writes them, the Foundry server uploads them (upload-asset, set-actor-art, add-journal-image). Nothing here talks to Foundry directly, so either half works on its own.

With a campaign repo: the same faces every week

Consistency across a campaign comes from a small file, not from luck. A campaign repo keeps an art shelf (art/SHELF.md): one approved portrait per player character with the short phrase that binds it in a prompt ("a bone-white orc in battered steel plate"), two or three finished pieces that carry the house look, the finish words that keep portraits matte and the palette muted, and a note of which older files are superseded and must never be attached again. The illustration-builder skill reads the shelf before every piece, attaches the anchors as character references and the shelf as style references, and proposes a shelf if the repo has none yet.

That is how Session Scribe's records stay illustrated by the same people session after session, how a token redress lands as the actor your world already knows, and how a book assembled at the end of a campaign reads as one artist's work. Greenrest's shelf, its approved art, and the chronicle built from them are public in fvtt-campaign-greenrest.

How it works

Claude ──MCP──> fvtt-mcp-imagegen ──HTTPS──> Gemini image API
                      │
                      ├── sharp: convert, crop, resize
                      └── token_cutout.py: chroma key or rembg → alpha
  • The API returns a JPEG; the server converts to PNG and applies the kind's post-processing (512 square for icons, 16:9 to 16:10 crop for illustrations, cutout for tokens).

  • Tokens are rendered on a flat chroma plate, then keyed out. Magenta is the default because soft edges keep a trace of the plate and a magenta trace reads as a dark outline on warm subjects (skin, hair, fur, leather) where green reads as an olive fringe; green or blue takes over for purple, pink, or violet subjects. A magenta-composited preview is written beside every cut so the edge can be checked.

  • Before a token is cut, the server checks the plate's outer edge for subject pixels. A clipped render is redone once and refused if it clips again.

  • A battlemap is sent padded to the nearest aspect the API renders (a mirrored margin), and the render is mapped back onto the source's pixel grid. The API scales its input to cover the output and trims the excess, so a "3:4" render is 1792×2400, not 1800×2400. Then both images are reduced to edge strength and compared tile by tile (phase correlation): a restyle changes colour and texture, edges stay put. More than 3% of tiles moved by over 0.4% of the long side means the layout drifted. A checkerboard of source and result is written beside every map for the eye.

  • A render the image safety filter blocks is retried once; the filter is not consistent on the same input.

  • No local models, no fine-tuning, no ComfyUI. Style comes from reference images you attach.

Development

npm test          # offline unit suite; nothing hits the live API
npm run typecheck
npm run check     # biome
npm run knip

Part of Open Roll 5e

fvtt-mcp-imagegen is one of the three MCP servers in Open Roll 5e, a suite of Foundry VTT modules and Claude Code tooling built for one D&D 5e table and shared. The other servers:

  • fvtt-mcp-dnd5e: builds D&D 5e content in a live Foundry world from Claude Code: a stat block becomes a complete NPC, a map image a walled and lit scene, an adventure its journals, tables and handouts.

  • fvtt-mcp-sessionscribe: turns a session's Discord recording and Foundry chat log into its record. Its end-to-end session-scribe skill drives the server from the Craig link to a speaker-labelled transcript, a fully illustrated player recap, combat statistics, GM notes and a party snapshot.

The modules, each of which installs and works on its own and none of which needs another:

  • Open Roll 5e: Autoexplore: lets a scene start fully explored, so the whole map shows through the fog of war while tokens still need line of sight.

  • Open Roll 5e: Battle Flow: combat automation for dnd5e 2024 rules: a hit rolls and applies its own damage, saves resolve themselves, reactions hold, and concentration is tracked. Every rule that touches a fight in the 2024 core books, Heroes of Faerûn, Arcana Unleashed and Ravenloft: The Horrors Within.

  • Open Roll 5e: Combat Plus: automates the chores of running a fight: combat music, an initiative gate, an out-of-turn movement block, defeated marking at 0 HP and turn alerts.

  • Open Roll 5e: Errata: corrects, in memory, bugs in the premium D&D 2024 books, the dnd5e system and Foundry itself, each fix held until the vendor ships its own.

  • Open Roll 5e: FX Studio: visual and sound effects for dnd5e, played from what actually happened at the table, with about a thousand stock FX and a window for authoring your own.

  • Open Roll 5e: Loot Shelf: loot chests and merchant shelves that players can take from, buy from and sell to without owning them, with a receipt for every trade.

  • Open Roll 5e: Open Server: for hosted worlds: clears the startup pause so players can play before the GM arrives, and gives any user a landing scene of their own.

  • Open Roll 5e: Party Stash: makes a dnd5e Group actor's inventory a working party stash: drags move instead of copying, coin moves through a dialog, and every transfer posts a receipt.

  • Open Roll 5e: Soundscape: background sound for scenes: random one-shots with silence between them, seamless crossfaded loops, day and night gating, and quiet during combat.

Issues are welcome on every repo in the family; pull requests are not accepted, since each is one author's design for one table, shared because it might suit yours. How they fit together is mapped in fvtt-suite-openroll5e.

License

MIT. See LICENSE.

Available Tools

4 tools
artificer-statusA

Health check: API key present, which image models the key can reach, the output directory, and estimated session spend by tier. Call this first on a cold start.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It effectively discloses that this is an informational read of several system states and gives useful specifics. However, it does not describe the return format, failure behavior, or whether the call itself has any side effects or costs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences. The first packs the full scope of the health check into a comma-separated list, and the second delivers the call guidance. No filler words or repeated schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless health check with no annotations and no output schema, the description is nearly sufficient. It states what is checked and when to call it, which is enough to select and invoke the tool. The only gap is that it doesn't explicitly describe the shape of the returned status report, but that is minor for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so there is no parameter meaning to convey. The description adds nothing about parameters because none exist; the baseline of 4 for zero-parameter tools applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Health check') and enumerates the exact resources inspected: API key presence, image model reachability, output directory, and session spend. This clearly distinguishes it from the sibling image generation/editing tools without needing to inspect their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs when to use the tool: 'Call this first on a cold start.' This gives clear usage context. It does not name alternatives or state when-not-to-use, but the distinct nature of the image-operation siblings makes the exclusion obvious enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cutout-imageA

Knock the background off a token image to real alpha and deliver it centred on a 512 square so Foundry scale 1.0 is right. Writes a magenta-composited *_preview.png beside it: READ THAT before trusting the edge. Returns coverage and residual-key numbers; a cut outside sane coverage falls back to the rembg AI matte automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoSquare canvas edge; default 512 (Foundry scale 1.0). 0 keeps the source canvas.
trimNoTighten to the subject before fitting (default true); false letterboxes as-is.
colorNoChroma key colour: "green", "magenta", "blue", or #RRGGBB. Omit to sample the corners.
erodeNoShrink the matte N px to eat a fringe.
methodNoauto (default): chroma if the plate is a flat key colour, with a rembg fallback when the cut fails verification. chroma: flat green/blue/magenta/solid plates, instant. rembg: AI matte for busy backgrounds, hair, and soft edges (first use downloads a ~176 MB model).
outputNoAbsolute output path (.png). Default: next to the source as <name>-cut.png.
padPctNoTransparent margin, % of the edge (4).
dropShadowNoAdd the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24).
keepShadowNochroma only: keep a cast shadow on the plate.
sourceImageYesAbsolute path of the image to cut (PNG/JPEG/WebP).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and meets it well. It discloses that a *_preview.png file is written, that the tool returns coverage and residual-key numbers, that unsafe cuts automatically fall back to rembg, and the schema additionally notes the ~176 MB model download on first rembg use. Nothing about the tool's side effects is hidden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two dense sentences with no filler. The primary purpose is front-loaded, followed immediately by the most important safety caveat (verify with the preview), then fallback behavior. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter image-processing tool with no output schema and no annotations, the description plus rich schema covers nearly everything: purpose, output file, verification step, return metrics, and failure fallback. It is slightly jargon-heavy ('sane coverage', 'residual-key numbers') and does not define thresholds, but an agent can still select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The prose does reinforce some semantics like 'Foundry scale 1.0' and the preview-file caveat, but it does not add meaning to individual parameters beyond what the schema already provides. The schema's rich parameter descriptions carry the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Knock the background off a token image'), then states the exact output contract: real alpha, centred on a 512 square, Foundry scale 1.0. It clearly distinguishes this tool from siblings like generate-image and edit-image by focusing on background removal and alpha output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete operational guidance: read the magenta-composited preview before trusting edges, and be aware of automatic fallback to rembg when the cut fails verification. The method parameter further explains when to use chroma vs rembg. It does not explicitly compare against sibling tools, but those are not close alternatives, so this is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit-imageA

Edit an existing image with one instruction while keeping identity, pose, angle, and style. Flash for every kind (pro was no better at fixes and re-cropped once). Tokens are prompted light (your instruction as you would type it in the Gemini app, plus a keep-face/hair/angle line and "remove any cast shadow"), put back on a chroma plate keyed to the token's own colours, and cut to alpha on the 512 square in the same call. "give this an updated painterly style" restyles a world token in place. Props (kind "prop") get object-only wording (no figures added) and come back cut at the source file's exact pixel size, ready for the same tile slot. Returns the new file path, dimensions, and estimated spend.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesPurpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro.
slugYesKebab-cased into the filename: <kind>-<slug>-<id>.png.
tierNoDefault flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.
confirmProNoRequired true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it.
referencesNoOptional extra references (attached after the source; indexes start at 2).
instructionYesThe change, and only the change: "replace the greatsword with a war maul crackling with violet energy". For a flaw-fix pass, name every flaw precisely in one instruction ("the left peryton has four legs; give it two", "remove the second fireball") and end with "keep everything else identical". Everything else is kept by the tool's own wording.
sourceImageYesAbsolute path of the image to edit (PNG/JPEG/WebP).
creatureSizeNoTokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses preservation of identity/pose/angle/style, default flash behavior, auto chroma-plating/cutting for tokens, prop pixel-size preservation, and returns new file path, dimensions, and estimated spend. It stops short of stating auth or side effects on the source file, but 'new file path' implies non-destructive output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is a good front-loaded summary, but the rest is a rambling, run-on explanation mixing model behavior, usage tips, and process internals ('put back on a chroma plate keyed to the token's own colours'). Several phrases ('pro was no better at fixes and re-cropped once', 'give this an updated painterly style') are vague and could be cut or reorganized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 parameters, no output schema, and no annotations, yet the description covers return information and kind-specific behavior. It is not complete enough for an agent to fully predict behavior—e.g., it doesn't state whether the original is preserved, what happens on failure, or when pro is genuinely preferred despite the flash recommendation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already very detailed, so the baseline is 3. The description adds value beyond the schema by giving instruction-phrasing guidance for tokens ('prompted light ... plus a keep-face/hair/angle line') and for props ('object-only wording'), which helps the agent construct a correct instruction value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Edit an existing image with one instruction while keeping identity, pose, angle, and style.' This clearly separates the tool from generate-image (creation) and cutout-image (isolation), and the rest of the description fills in kind-specific behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: it is for editing an existing image, flash is the default, tokens/props have specific wording guidance, and the returned payload is described. It does not explicitly name sibling tools or state when not to use this tool, but the 'existing image' framing makes the primary use case unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate-imageA

Generate one Foundry art asset from a prompt via the Gemini image API. kind picks the model tier, aspect, size, framing text, and post-processing; the result is a finished PNG on disk. READ IT before showing anyone: count limbs per creature, check for duplicated spell effects or props, stray signatures, and reference faces on the wrong figure; obvious flaws are one edit-image call away. Every kind runs on flash by default; tier: "pro" refuses without confirmPro: true and states the cost. Returns the file path, dimensions, and estimated spend.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesPurpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro.
slugYesKebab-cased into the filename: <kind>-<slug>-<id>.png.
tierNoDefault flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.
promptYesWhat a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it.
footprintNoProps only: grid cells wide x tall, e.g. "2x1" for a table (default "1x1"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: "TC_Anvil 02_2x1.png".
confirmProNoRequired true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it.
referencesNoReference images, attached in this order. Bind each in the prompt by its label or 1-based index ("Image 2 is Morgash").
creatureSizeNoTokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool writes a PNG to disk, that pro tier refuses without confirmPro and states cost, that every kind defaults to flash, and that post-processing (alpha cut, framing, plate) is done automatically. It also warns to inspect the result before showing anyone. It does not mention rate limits or failure modes, but the behavioral traits that matter for invocation are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: a one-sentence purpose, a QA warning, a tier/cost note, and a return summary. It front-loads the core purpose and the most important behavioral caveat (read before showing). It is longer than ideal, but every sentence adds operational value; the only minor redundancy is repeating flash default and confirmPro, which also appear in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema and no annotations, the description is quite complete: it covers the output (file path, dimensions, estimated spend), the tier gating, the QA expectation, and the per-kind behavior. It does not describe error cases or what happens when references are invalid, but the schema already documents reference roles and limits. The description is sufficient for an agent to invoke the tool correctly in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema: it explains that kind picks model tier, aspect, size, framing text, and post-processing; it clarifies that icons/tokens get framing/background appended automatically; it explains the pro tier cost and confirmPro requirement; and it gives concrete output dimensions. This goes beyond the schema's field descriptions, so a 4 is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate one Foundry art asset from a prompt via the Gemini image API.' It names the output (finished PNG on disk) and distinguishes the tool from siblings by mentioning edit-image as a follow-up for fixing flaws. The kind parameter further clarifies the five asset types, so an agent can tell this apart from edit-image or cutout-image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool (generate a new asset) and when to use edit-image ('obvious flaws are one edit-image call away'). It also gives usage context for pro tier: 'refuses without confirmPro: true and states the cost,' and instructs to offer pro only as an option to the owner. This is strong routing guidance relative to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.1.0
    • Changedcutout-image1 field changed
      • addedInput schema / properties / dropShadow
        Added value: +{
        +  "description": "Add the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24).",
        +  "type": "boolean"
        +}
    • Changededit-image5 fields changed
      • addedInput schema / properties / creatureSize
        Added value: +{
        +  "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.",
        +  "enum": [
        +    "medium",
        +    "large"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / kind / description
        Previous value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
      • changedInput schema / properties / kind / enum
        Previous value: -[
        -  "icon",
        -  "token",
        -  "portrait",
        -  "illustration"
        -]New value: +[
        +  "icon",
        +  "token",
        +  "prop",
        +  "portrait",
        +  "illustration"
        +]
      • changedInput schema / properties / references / items / properties / role / description
        Previous value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role."
      • changedInput schema / properties / references / items / properties / role / enum
        Previous value: -[
        -  "character",
        -  "style"
        -]New value: +[
        +  "character",
        +  "style",
        +  "pose"
        +]
    • Changedgenerate-image6 fields changed
      • addedInput schema / properties / creatureSize
        Added value: +{
        +  "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.",
        +  "enum": [
        +    "medium",
        +    "large"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / footprint
        Added value: +{
        +  "description": "Props only: grid cells wide x tall, e.g. \"2x1\" for a table (default \"1x1\"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: \"TC_Anvil 02_2x1.png\".",
        +  "pattern": "^\\d{1,2}x\\d{1,2}$",
        +  "type": "string"
        +}
      • changedInput schema / properties / kind / description
        Previous value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
      • changedInput schema / properties / kind / enum
        Previous value: -[
        -  "icon",
        -  "token",
        -  "portrait",
        -  "illustration"
        -]New value: +[
        +  "icon",
        +  "token",
        +  "prop",
        +  "portrait",
        +  "illustration"
        +]
      • changedInput schema / properties / references / items / properties / role / description
        Previous value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role."
      • changedInput schema / properties / references / items / properties / role / enum
        Previous value: -[
        -  "character",
        -  "style"
        -]New value: +[
        +  "character",
        +  "style",
        +  "pose"
        +]
  2. 4 tool updatesv1.0.0
    • Addedcutout-image
    • Addededit-image
    • Changedgenerate-image12 fields changed
      • removedInput schema / properties / batch
        Removed value: -{
        -  "default": 6,
        -  "description": "Draft mode only: images per batch.",
        -  "maximum": 8,
        -  "minimum": 1,
        -  "type": "integer"
        -}
      • addedInput schema / properties / confirmPro
        Added value: +{
        +  "description": "Required true with tier: \"pro\". Offer pro to the owner as an option for portraits and illustrations (\"pro is available for a bit extra\"); never assume it.",
        +  "type": "boolean"
        +}
      • removedInput schema / properties / denoise
        Removed value: -{
        -  "default": 0.7,
        -  "description": "Refine mode only. 0.7 (pinned by test) keeps the scene skeleton in dev style; ~0.55 clones composition but inherits the draft rendering style.",
        -  "maximum": 0.95,
        -  "minimum": 0.3,
        -  "type": "number"
        -}
      • changedInput schema / properties / kind / description
        Previous value: -"Purpose preset — fixes generation and output resolution. No raw dimensions."New value: +"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
      • changedInput schema / properties / kind / enum
        Previous value: -[
        -  "handout",
        -  "scene-background",
        -  "portrait",
        -  "token"
        -]New value: +[
        +  "icon",
        +  "token",
        +  "portrait",
        +  "illustration"
        +]
      • removedInput schema / properties / mode
        Removed value: -{
        -  "default": "draft",
        -  "description": "draft: fast klein batch for curation. final: dev-quality render from the prompt alone, finished at output resolution. refine: dev img2img over sourceImage (a picked draft) — keeps its scene skeleton, re-renders in dev style, finished at output resolution.",
        -  "enum": [
        -    "draft",
        -    "final",
        -    "refine"
        -  ],
        -  "type": "string"
        -}
      • changedInput schema / properties / prompt / description
        Previous value: -"The full image prompt."New value: +"What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it."
      • addedInput schema / properties / references
        Added value: +{
        +  "description": "Reference images, attached in this order. Bind each in the prompt by its label or 1-based index (\"Image 2 is Morgash\").",
        +  "items": {
        +    "properties": {
        +      "label": {
        +        "description": "Short name used to bind the reference in the prompt, e.g. \"Morgash\".",
        +        "type": "string"
        +      },
        +      "path": {
        +        "description": "Absolute path of a PNG/JPEG on disk.",
        +        "minLength": 1,
        +        "type": "string"
        +      },
        +      "role": {
        +        "description": "character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice).",
        +        "enum": [
        +          "character",
        +          "style"
        +        ],
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "path",
        +      "role"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 14,
        +  "type": "array"
        +}
      • removedInput schema / properties / seed
        Removed value: -{
        -  "description": "Fixed seed; random when omitted.",
        -  "minimum": 0,
        -  "type": "integer"
        -}
      • changedInput schema / properties / slug / description
        Previous value: -"Short kebab-case subject name used in output filenames, e.g. \"smugglers-cove\"."New value: +"Kebab-cased into the filename: <kind>-<slug>-<id>.png."
      • removedInput schema / properties / sourceImage
        Removed value: -{
        -  "description": "Refine mode only (required there): absolute path of the picked draft PNG.",
        -  "type": "string"
        -}
      • addedInput schema / properties / tier
        Added value: +{
        +  "description": "Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. \"pro\" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.",
        +  "enum": [
        +    "flash",
        +    "pro"
        +  ],
        +  "type": "string"
        +}
    • Removedupscale-image
  3. 3 tool updatesv0.1.0
    • First observedartificer-status
    • First observedgenerate-image
    • First observedupscale-image

TDQS

A4.2/5.0

Scored across 4 tools

Disambiguation4/5

The tools are mostly distinct: generate-image creates new assets, edit-image modifies existing ones, cutout-image handles background removal, and artificer-status is a health check. However, edit-image and generate-image could be slightly confused since both can produce a final PNG, though the descriptions clarify the difference (new vs. existing).

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (generate-image, edit-image, cutout-image, artificer-status). The pattern is uniform and predictable, making it easy for an agent to infer the action and target.

Tool Count5/5

With only 4 tools, the server is tightly scoped to the image generation and processing workflow. Each tool serves a clear, necessary function, and the count is appropriate for the server's stated purpose—creating and preparing Foundry art assets.

Completeness3/5

The set covers the core lifecycle: generate, edit, cutout, and status. However, there are gaps—no explicit upload to Foundry, no batch processing, and no tool to list or delete existing assets. The documentation mentions previews and fallbacks, but the surface is missing endpoints for managing the output directory or integrating with Foundry beyond saving to disk.

Maintenance

ActivityMaintained
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    A
    maintenance
    Connects Claude Desktop to Foundry VTT for AI-powered campaign management, enabling natural language interaction with game data including quest creation, character management, compendium searches, and dice rolling. Provides 20 MCP tools for seamless integration between Claude and your tabletop RPG sessions.
    72
    -
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Integrates with FoundryVTT tabletop gaming sessions, allowing AI assistants to query game data, roll dice, generate content (NPCs, loot, encounters), manage combat, and provide tactical suggestions through natural language.
    7 npm
    -