Skip to main content
Glama

fvtt-mcp-artificer

Un servidor de Model Context Protocol para generación de imágenes específico para Foundry, para arte de mesa de D&D, impulsado por Claude Code. Envuelve una instancia headless de ComfyUI en la máquina local y expone un pequeño conjunto de herramientas orientadas a Foundry para que Claude pueda redactar prompts, generar lotes de ilustraciones, curar los resultados mirándolos de verdad y entregar los ganadores al pipeline de Foundry en su servidor hermano, fvtt-mcp-molten5e (upload-asset → set-actor-art / add-journal-image / fondos de escena).

El ciclo completo es lo bastante rápido como para ser conversacional en el hardware objetivo (RTX 5090): un lote de borradores de 6 imágenes se genera en ~10 s, un render terminado de 2560×1600 en ~19 s, y el ciclo completo prompt → borrador → curar → final → diario en el mundo se probó de extremo a extremo en notes/m3-loop-proof.md.

Por qué esta forma

Esto no es un puente genérico de ComfyUI, por decreto. Las herramientas hablan el vocabulario de Foundry — retratos de actor, tokens, handouts, fondos de escena — y se ajustan libremente para el trabajo con Foundry. Y se mantiene separado de fvtt-mcp-molten5e: ese servidor está limitado a la creación de contenido de Foundry y no debe acoplarse a la generación de imágenes. Este servidor nunca habla con el puente de Foundry; el traspaso entre ambos son archivos en disco más las herramientas de subida de molten5e.

La misma filosofía de casa que el resto de la familia: las herramientas hacen, las skills deciden. La corrección (ejecución del workflow, dimensiones, el pipeline de upscaling, convenciones de archivos) reside en las herramientas probadas; el criterio (elaboración del prompt, gusto de curación, qué actor o diario recibe el arte, estilo de casa) reside en una skill posterior illustration-builder.

Claude ──MCP──> fvtt-mcp-artificer ──HTTP──> ComfyUI (headless, local)
                      │
                      └── pinned workflow JSONs (draft / final / final-refine / upscale)

ComfyUI se ejecuta headless en modo API; este servidor envía los workflow JSON fijados que están en workflows/ — nunca grafos de forma libre — con prompt/seed/batch/preset sustituidos (por ID de nodo, protegido por aserciones de deriva de tipo de clase). Las herramientas devuelven rutas de archivo absolutas; Claude lee los PNG directamente para curar.

Related MCP server: FoundryVTT MCP Server

Presets por finalidad, no dimensiones sin procesar

generate-image recibe un kind, no ancho/alto:

kind

genera a

salida final

handout / scene-background

1536×960

2560×1600

portrait

1024×1280

2048×2560

token

1024×1024

2048×2048

Pipeline de resolución (fijado): nunca generar al tamaño de salida — la composición se degrada más allá de ~1.5 MP. Genera a la resolución nativa del preset, upscaling por modelo ×4 con 4x-UltraSharp, downsampling lanczos al tamaño final, todo en un único grafo fijado.

Modelos

  • FLUX.1-dev fp8 (Comfy-Org all-in-one) — renders de calidad final, ~15–19 s por render terminado.

  • FLUX.2-klein 4B (Apache 2.0) — el modelo de borradores: 4 pasos, ~1–2 s/imagen, lote de 6–8, elige y vuelve a renderizar con dev.

  • 4x-UltraSharp — el upscaler de modelo en la cola del pipeline.

Herramientas

herramienta

qué hace

generate-image

kind + prompt + slug, con mode: draft (lote de klein para curar), final (render con dev desde el prompt, terminado a la resolución de salida) o refine (dev img2img con denoise 0.7 sobre un borrador elegido — conserva su esqueleto de escena, re-renderiza en estilo dev, terminado a la resolución de salida). Devuelve rutas absolutas de PNG nombradas <kind>-<slug>-<seed>_NNNNN_.png.

upscale-image

Termina una imagen existente (normalmente un borrador que ganó la curación directamente) a través de la cola de upscaling hasta la resolución de salida de su kind.

artificer-status

Comprobación de salud: accesibilidad/versión de ComfyUI, VRAM, profundidad de la cola, integridad de los workflow fijados, presencia de los modelos requeridos.

El contrato de sustitución entre las herramientas y los grafos fijados está documentado en workflows/README.md.

Requisitos

  • Windows + GPU NVIDIA. Desarrollado y probado en una RTX 5090 (Blackwell necesita PyTorch con CUDA 12.8+; la compilación portable actual de ComfyUI lo incluye).

  • ComfyUI (compilación portable independiente, v0.34+) ejecutándose headless — consulta scripts/launch-comfyui.ps1 para la convención de lanzamiento (modo API en 127.0.0.1:8188, con --output-directory fijado).

  • Los archivos de modelo en el árbol models/ de ComfyUI (~33 GB, todos sin restricciones de acceso): checkpoints/flux1-dev-fp8.safetensors, diffusion_models/flux-2-klein-4b.safetensors, text_encoders/qwen_3_4b.safetensors, vae/flux2-vae.safetensors, upscale_models/4x-UltraSharp.safetensors. artificer-status informa de cualquier modelo que falte.

  • Node.js 22+ para el propio servidor MCP.

Compilación

npm install
npm run build

Pruebas: npm test (suite de pruebas unitarias sin conexión; los workflow JSON fijados son los fixtures). La suite en vivo — ComfyUI real, renders reales — está detrás de npm run test:integration. Puertas de calidad: npm run typecheck, npm run check (biome), npm run knip.

Integración con Claude Code

La convención de la casa registra el servidor en el alcance de usuario (nuevas herramientas MCP ⇒ reinicia Claude Code). O usa la CLI:

claude mcp add -s user artificer -- node D:/path/to/fvtt-mcp-artificer/dist/index.js

o copia .mcp.json.example a un .mcp.json que Claude Code lea (o combínalo en mcpServers de ~/.claude.json) con rutas absolutas. En Windows, apunta command a la ruta completa de node.exe si Node no está en PATH.

Configuración

Copia .env.example a .env (ignorado por git):

  • COMFY_URL — la instancia headless (por defecto http://127.0.0.1:8188).

  • COMFY_OUTPUT_DIR — debe coincidir con el --output-directory con el que se lanzó ComfyUI; el servidor lee los PNG generados directamente de esta ruta.

  • ARTIFICER_TIMEOUT_MS — límite de espera por trabajo (por defecto 300000).

Licencia

Licencia MIT — consulta LICENSE para más detalles.

Available Tools

4 tools
artificer-statusA

Health check: API key present, which image models the key can reach, the output directory, and estimated session spend by tier. Call this first on a cold start.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It effectively discloses that this is an informational read of several system states and gives useful specifics. However, it does not describe the return format, failure behavior, or whether the call itself has any side effects or costs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences. The first packs the full scope of the health check into a comma-separated list, and the second delivers the call guidance. No filler words or repeated schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless health check with no annotations and no output schema, the description is nearly sufficient. It states what is checked and when to call it, which is enough to select and invoke the tool. The only gap is that it doesn't explicitly describe the shape of the returned status report, but that is minor for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so there is no parameter meaning to convey. The description adds nothing about parameters because none exist; the baseline of 4 for zero-parameter tools applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Health check') and enumerates the exact resources inspected: API key presence, image model reachability, output directory, and session spend. This clearly distinguishes it from the sibling image generation/editing tools without needing to inspect their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs when to use the tool: 'Call this first on a cold start.' This gives clear usage context. It does not name alternatives or state when-not-to-use, but the distinct nature of the image-operation siblings makes the exclusion obvious enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cutout-imageA

Knock the background off a token image to real alpha and deliver it centred on a 512 square so Foundry scale 1.0 is right. Writes a magenta-composited *_preview.png beside it: READ THAT before trusting the edge. Returns coverage and residual-key numbers; a cut outside sane coverage falls back to the rembg AI matte automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeNoSquare canvas edge; default 512 (Foundry scale 1.0). 0 keeps the source canvas.
trimNoTighten to the subject before fitting (default true); false letterboxes as-is.
colorNoChroma key colour: "green", "magenta", "blue", or #RRGGBB. Omit to sample the corners.
erodeNoShrink the matte N px to eat a fringe.
methodNoauto (default): chroma if the plate is a flat key colour, with a rembg fallback when the cut fails verification. chroma: flat green/blue/magenta/solid plates, instant. rembg: AI matte for busy backgrounds, hair, and soft edges (first use downloads a ~176 MB model).
outputNoAbsolute output path (.png). Default: next to the source as <name>-cut.png.
padPctNoTransparent margin, % of the edge (4).
dropShadowNoAdd the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24).
keepShadowNochroma only: keep a cast shadow on the plate.
sourceImageYesAbsolute path of the image to cut (PNG/JPEG/WebP).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and meets it well. It discloses that a *_preview.png file is written, that the tool returns coverage and residual-key numbers, that unsafe cuts automatically fall back to rembg, and the schema additionally notes the ~176 MB model download on first rembg use. Nothing about the tool's side effects is hidden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two dense sentences with no filler. The primary purpose is front-loaded, followed immediately by the most important safety caveat (verify with the preview), then fallback behavior. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter image-processing tool with no output schema and no annotations, the description plus rich schema covers nearly everything: purpose, output file, verification step, return metrics, and failure fallback. It is slightly jargon-heavy ('sane coverage', 'residual-key numbers') and does not define thresholds, but an agent can still select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The prose does reinforce some semantics like 'Foundry scale 1.0' and the preview-file caveat, but it does not add meaning to individual parameters beyond what the schema already provides. The schema's rich parameter descriptions carry the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Knock the background off a token image'), then states the exact output contract: real alpha, centred on a 512 square, Foundry scale 1.0. It clearly distinguishes this tool from siblings like generate-image and edit-image by focusing on background removal and alpha output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete operational guidance: read the magenta-composited preview before trusting edges, and be aware of automatic fallback to rembg when the cut fails verification. The method parameter further explains when to use chroma vs rembg. It does not explicitly compare against sibling tools, but those are not close alternatives, so this is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit-imageA

Edit an existing image with one instruction while keeping identity, pose, angle, and style. Flash for every kind (pro was no better at fixes and re-cropped once). Tokens are prompted light (your instruction as you would type it in the Gemini app, plus a keep-face/hair/angle line and "remove any cast shadow"), put back on a chroma plate keyed to the token's own colours, and cut to alpha on the 512 square in the same call. "give this an updated painterly style" restyles a world token in place. Props (kind "prop") get object-only wording (no figures added) and come back cut at the source file's exact pixel size, ready for the same tile slot. Returns the new file path, dimensions, and estimated spend.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesPurpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro.
slugYesKebab-cased into the filename: <kind>-<slug>-<id>.png.
tierNoDefault flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.
confirmProNoRequired true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it.
referencesNoOptional extra references (attached after the source; indexes start at 2).
instructionYesThe change, and only the change: "replace the greatsword with a war maul crackling with violet energy". For a flaw-fix pass, name every flaw precisely in one instruction ("the left peryton has four legs; give it two", "remove the second fireball") and end with "keep everything else identical". Everything else is kept by the tool's own wording.
sourceImageYesAbsolute path of the image to edit (PNG/JPEG/WebP).
creatureSizeNoTokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses preservation of identity/pose/angle/style, default flash behavior, auto chroma-plating/cutting for tokens, prop pixel-size preservation, and returns new file path, dimensions, and estimated spend. It stops short of stating auth or side effects on the source file, but 'new file path' implies non-destructive output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is a good front-loaded summary, but the rest is a rambling, run-on explanation mixing model behavior, usage tips, and process internals ('put back on a chroma plate keyed to the token's own colours'). Several phrases ('pro was no better at fixes and re-cropped once', 'give this an updated painterly style') are vague and could be cut or reorganized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 parameters, no output schema, and no annotations, yet the description covers return information and kind-specific behavior. It is not complete enough for an agent to fully predict behavior—e.g., it doesn't state whether the original is preserved, what happens on failure, or when pro is genuinely preferred despite the flash recommendation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already very detailed, so the baseline is 3. The description adds value beyond the schema by giving instruction-phrasing guidance for tokens ('prompted light ... plus a keep-face/hair/angle line') and for props ('object-only wording'), which helps the agent construct a correct instruction value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Edit an existing image with one instruction while keeping identity, pose, angle, and style.' This clearly separates the tool from generate-image (creation) and cutout-image (isolation), and the rest of the description fills in kind-specific behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: it is for editing an existing image, flash is the default, tokens/props have specific wording guidance, and the returned payload is described. It does not explicitly name sibling tools or state when not to use this tool, but the 'existing image' framing makes the primary use case unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate-imageA

Generate one Foundry art asset from a prompt via the Gemini image API. kind picks the model tier, aspect, size, framing text, and post-processing; the result is a finished PNG on disk. READ IT before showing anyone: count limbs per creature, check for duplicated spell effects or props, stray signatures, and reference faces on the wrong figure; obvious flaws are one edit-image call away. Every kind runs on flash by default; tier: "pro" refuses without confirmPro: true and states the cost. Returns the file path, dimensions, and estimated spend.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesPurpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro.
slugYesKebab-cased into the filename: <kind>-<slug>-<id>.png.
tierNoDefault flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.
promptYesWhat a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it.
footprintNoProps only: grid cells wide x tall, e.g. "2x1" for a table (default "1x1"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: "TC_Anvil 02_2x1.png".
confirmProNoRequired true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it.
referencesNoReference images, attached in this order. Bind each in the prompt by its label or 1-based index ("Image 2 is Morgash").
creatureSizeNoTokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool writes a PNG to disk, that pro tier refuses without confirmPro and states cost, that every kind defaults to flash, and that post-processing (alpha cut, framing, plate) is done automatically. It also warns to inspect the result before showing anyone. It does not mention rate limits or failure modes, but the behavioral traits that matter for invocation are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: a one-sentence purpose, a QA warning, a tier/cost note, and a return summary. It front-loads the core purpose and the most important behavioral caveat (read before showing). It is longer than ideal, but every sentence adds operational value; the only minor redundancy is repeating flash default and confirmPro, which also appear in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema and no annotations, the description is quite complete: it covers the output (file path, dimensions, estimated spend), the tier gating, the QA expectation, and the per-kind behavior. It does not describe error cases or what happens when references are invalid, but the schema already documents reference roles and limits. The description is sufficient for an agent to invoke the tool correctly in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema: it explains that kind picks model tier, aspect, size, framing text, and post-processing; it clarifies that icons/tokens get framing/background appended automatically; it explains the pro tier cost and confirmPro requirement; and it gives concrete output dimensions. This goes beyond the schema's field descriptions, so a 4 is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate one Foundry art asset from a prompt via the Gemini image API.' It names the output (finished PNG on disk) and distinguishes the tool from siblings by mentioning edit-image as a follow-up for fixing flaws. The kind parameter further clarifies the five asset types, so an agent can tell this apart from edit-image or cutout-image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool (generate a new asset) and when to use edit-image ('obvious flaws are one edit-image call away'). It also gives usage context for pro tier: 'refuses without confirmPro: true and states the cost,' and instructs to offer pro only as an option to the owner. This is strong routing guidance relative to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.1.0
    • Changedcutout-image1 field changed
      • addedInput schema / properties / dropShadow
        Added value: +{
        +  "description": "Add the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24).",
        +  "type": "boolean"
        +}
    • Changededit-image5 fields changed
      • addedInput schema / properties / creatureSize
        Added value: +{
        +  "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.",
        +  "enum": [
        +    "medium",
        +    "large"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / kind / description
        Previous value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
      • changedInput schema / properties / kind / enum
        Previous value: -[
        -  "icon",
        -  "token",
        -  "portrait",
        -  "illustration"
        -]New value: +[
        +  "icon",
        +  "token",
        +  "prop",
        +  "portrait",
        +  "illustration"
        +]
      • changedInput schema / properties / references / items / properties / role / description
        Previous value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role."
      • changedInput schema / properties / references / items / properties / role / enum
        Previous value: -[
        -  "character",
        -  "style"
        -]New value: +[
        +  "character",
        +  "style",
        +  "pose"
        +]
    • Changedgenerate-image6 fields changed
      • addedInput schema / properties / creatureSize
        Added value: +{
        +  "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.",
        +  "enum": [
        +    "medium",
        +    "large"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / footprint
        Added value: +{
        +  "description": "Props only: grid cells wide x tall, e.g. \"2x1\" for a table (default \"1x1\"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: \"TC_Anvil 02_2x1.png\".",
        +  "pattern": "^\\d{1,2}x\\d{1,2}$",
        +  "type": "string"
        +}
      • changedInput schema / properties / kind / description
        Previous value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
      • changedInput schema / properties / kind / enum
        Previous value: -[
        -  "icon",
        -  "token",
        -  "portrait",
        -  "illustration"
        -]New value: +[
        +  "icon",
        +  "token",
        +  "prop",
        +  "portrait",
        +  "illustration"
        +]
      • changedInput schema / properties / references / items / properties / role / description
        Previous value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role."
      • changedInput schema / properties / references / items / properties / role / enum
        Previous value: -[
        -  "character",
        -  "style"
        -]New value: +[
        +  "character",
        +  "style",
        +  "pose"
        +]
  2. 4 tool updatesv1.0.0
    • Addedcutout-image
    • Addededit-image
    • Changedgenerate-image12 fields changed
      • removedInput schema / properties / batch
        Removed value: -{
        -  "default": 6,
        -  "description": "Draft mode only: images per batch.",
        -  "maximum": 8,
        -  "minimum": 1,
        -  "type": "integer"
        -}
      • addedInput schema / properties / confirmPro
        Added value: +{
        +  "description": "Required true with tier: \"pro\". Offer pro to the owner as an option for portraits and illustrations (\"pro is available for a bit extra\"); never assume it.",
        +  "type": "boolean"
        +}
      • removedInput schema / properties / denoise
        Removed value: -{
        -  "default": 0.7,
        -  "description": "Refine mode only. 0.7 (pinned by test) keeps the scene skeleton in dev style; ~0.55 clones composition but inherits the draft rendering style.",
        -  "maximum": 0.95,
        -  "minimum": 0.3,
        -  "type": "number"
        -}
      • changedInput schema / properties / kind / description
        Previous value: -"Purpose preset — fixes generation and output resolution. No raw dimensions."New value: +"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."
      • changedInput schema / properties / kind / enum
        Previous value: -[
        -  "handout",
        -  "scene-background",
        -  "portrait",
        -  "token"
        -]New value: +[
        +  "icon",
        +  "token",
        +  "portrait",
        +  "illustration"
        +]
      • removedInput schema / properties / mode
        Removed value: -{
        -  "default": "draft",
        -  "description": "draft: fast klein batch for curation. final: dev-quality render from the prompt alone, finished at output resolution. refine: dev img2img over sourceImage (a picked draft) — keeps its scene skeleton, re-renders in dev style, finished at output resolution.",
        -  "enum": [
        -    "draft",
        -    "final",
        -    "refine"
        -  ],
        -  "type": "string"
        -}
      • changedInput schema / properties / prompt / description
        Previous value: -"The full image prompt."New value: +"What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it."
      • addedInput schema / properties / references
        Added value: +{
        +  "description": "Reference images, attached in this order. Bind each in the prompt by its label or 1-based index (\"Image 2 is Morgash\").",
        +  "items": {
        +    "properties": {
        +      "label": {
        +        "description": "Short name used to bind the reference in the prompt, e.g. \"Morgash\".",
        +        "type": "string"
        +      },
        +      "path": {
        +        "description": "Absolute path of a PNG/JPEG on disk.",
        +        "minLength": 1,
        +        "type": "string"
        +      },
        +      "role": {
        +        "description": "character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice).",
        +        "enum": [
        +          "character",
        +          "style"
        +        ],
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "path",
        +      "role"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 14,
        +  "type": "array"
        +}
      • removedInput schema / properties / seed
        Removed value: -{
        -  "description": "Fixed seed; random when omitted.",
        -  "minimum": 0,
        -  "type": "integer"
        -}
      • changedInput schema / properties / slug / description
        Previous value: -"Short kebab-case subject name used in output filenames, e.g. \"smugglers-cove\"."New value: +"Kebab-cased into the filename: <kind>-<slug>-<id>.png."
      • removedInput schema / properties / sourceImage
        Removed value: -{
        -  "description": "Refine mode only (required there): absolute path of the picked draft PNG.",
        -  "type": "string"
        -}
      • addedInput schema / properties / tier
        Added value: +{
        +  "description": "Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. \"pro\" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.",
        +  "enum": [
        +    "flash",
        +    "pro"
        +  ],
        +  "type": "string"
        +}
    • Removedupscale-image
  3. 3 tool updatesv0.1.0
    • First observedartificer-status
    • First observedgenerate-image
    • First observedupscale-image

TDQS

A4.2/5.0

Scored across 4 tools

Disambiguation4/5

The tools are mostly distinct: generate-image creates new assets, edit-image modifies existing ones, cutout-image handles background removal, and artificer-status is a health check. However, edit-image and generate-image could be slightly confused since both can produce a final PNG, though the descriptions clarify the difference (new vs. existing).

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (generate-image, edit-image, cutout-image, artificer-status). The pattern is uniform and predictable, making it easy for an agent to infer the action and target.

Tool Count5/5

With only 4 tools, the server is tightly scoped to the image generation and processing workflow. Each tool serves a clear, necessary function, and the count is appropriate for the server's stated purpose—creating and preparing Foundry art assets.

Completeness3/5

The set covers the core lifecycle: generate, edit, cutout, and status. However, there are gaps—no explicit upload to Foundry, no batch processing, and no tool to list or delete existing assets. The documentation mentions previews and fallbacks, but the surface is missing endpoints for managing the output directory or integrating with Foundry beyond saving to disk.

Maintenance

ActivityMaintained
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    A
    maintenance
    Connects Claude Desktop to Foundry VTT for AI-powered campaign management, enabling natural language interaction with game data including quest creation, character management, compendium searches, and dice rolling. Provides 20 MCP tools for seamless integration between Claude and your tabletop RPG sessions.
    72
    -
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Integrates with FoundryVTT tabletop gaming sessions, allowing AI assistants to query game data, roll dice, generate content (NPCs, loot, encounters), manage combat, and provide tactical suggestions through natural language.
    7 npm
    -