fvtt-mcp-artificer
fvtt-mcp-artificer
Сервер Model Context Protocol для генерации изображений, специфичный для Foundry, для артов настольных D&D, управляемый Claude Code. Он оборачивает headless-инстанс ComfyUI на локальной машине и предоставляет небольшой набор инструментов, заточенных под Foundry, чтобы Claude мог составлять промпты, генерировать пакеты иллюстраций, курировать результаты, буквально глядя на них, и передавать победителей в конвейер Foundry через соседний сервер fvtt-mcp-molten5e (upload-asset → set-actor-art / add-journal-image / фоновые изображения сцен).
Весь цикл достаточно быстр для диалогового режима на целевом оборудовании (RTX 5090): пакет из 6 черновых изображений готов за ~10 с, готовый рендер 2560×1600 — за ~19 с, а полный цикл промпт → черновик → курирование → финал → внутриигровой журнал подтверждён от начала до конца в notes/m3-loop-proof.md.
Почему такая форма
Это не универсальный мост к ComfyUI, по замыслу. Инструменты оперируют словарём Foundry — портреты актёров, токены, раздаточные материалы, фоновые изображения сцен — и свободно дорабатываются под задачи Foundry. И он остаётся отдельным от fvtt-mcp-molten5e: тот сервер предназначен для создания контента Foundry и не должен быть связан с генерацией изображений. Этот сервер никогда не общается с Foundry-мостом; передача между ними — файлы на диске плюс инструменты загрузки molten5e.
Та же домашняя философия, что и у всего семейства: инструменты делают, навыки решают. Корректность (исполнение рабочих процессов, размеры, конвейер апскейла, соглашения о файлах) живёт в протестированных инструментах; суждение (мастерство промптов, вкус при курировании, какой актёр или журнал получит арт, домашний стиль) — в последующем навыке illustration-builder.
Claude ──MCP──> fvtt-mcp-artificer ──HTTP──> ComfyUI (headless, local)
│
└── pinned workflow JSONs (draft / final / final-refine / upscale)ComfyUI работает в headless-режиме через API; этот сервер отправляет закреплённые JSON рабочих процессов из workflows/ — никогда не произвольные графы — с подстановкой prompt/seed/batch/preset (по node id, под защитой проверок на дрейф типов классов). Инструменты возвращают абсолютные пути к файлам; Claude читает PNG напрямую, чтобы курировать.
Related MCP server: FoundryVTT MCP Server
Пресеты по назначению, а не сырые размеры
generate-image принимает kind, а не ширину/высоту:
kind | генерируется в | готовый результат |
| 1536×960 | 2560×1600 |
| 1024×1280 | 2048×2560 |
| 1024×1024 | 2048×2048 |
Конвейер разрешения (зафиксирован): никогда не генерируйте в выходном размере — композиция ухудшается после ~1.5 МП. Генерируйте в нативном разрешении пресета, увеличивайте моделью ×4 через 4x-UltraSharp, уменьшайте ланцошем до готового размера — всё в одном закреплённом графе.
Модели
FLUX.1-dev fp8 (Comfy-Org all-in-one) — рендеры финального качества: ~15–19 с на готовый результат.
FLUX.2-klein 4B (Apache 2.0) — черновая модель: 4 шага, ~1–2 с/изображение, пакет 6–8, отбор и повторный рендер с помощью dev.
4x-UltraSharp — модель-апскейлер в хвосте конвейера.
Инструменты
инструмент | что делает |
| принимает |
| Доводит существующее изображение (обычно черновик, сразу прошедший курирование) через хвост апскейла до выходного разрешения его |
| Проверка состояния: доступность/версия ComfyUI, VRAM, глубина очереди, целостность закреплённых рабочих процессов, наличие требуемых моделей. |
Контракт подстановки между инструментами и закреплёнными графами документирован в workflows/README.md.
Требования
Windows + NVIDIA GPU. Разработано и подтверждено на RTX 5090 (Blackwell требует PyTorch с CUDA 12.8+; текущая портативная сборка ComfyUI поставляется с ним).
ComfyUI (автономная портативная сборка, v0.34+), запускаемый в headless-режиме, — см.
scripts/launch-comfyui.ps1о соглашении по запуску (режим API на127.0.0.1:8188, фиксированный--output-directory).Файлы моделей в дереве
models/ComfyUI (~33 ГБ, все без ограничений):checkpoints/flux1-dev-fp8.safetensors,diffusion_models/flux-2-klein-4b.safetensors,text_encoders/qwen_3_4b.safetensors,vae/flux2-vae.safetensors,upscale_models/4x-UltraSharp.safetensors.artificer-statusсообщает обо всём отсутствующем.Node.js 22+ для самого MCP-сервера.
Сборка
npm install
npm run buildТесты: npm test (офлайн-набор модульных тестов; в качестве фикстур выступают закреплённые JSON рабочих процессов). Живой набор — реальный ComfyUI, реальные рендеры — запускается через npm run test:integration. Проверки качества: npm run typecheck, npm run check (biome), npm run knip.
Подключение к Claude Code
По принятому в проекте соглашению сервер регистрируется в области пользователя (новые инструменты MCP ⇒ перезапустите Claude Code). Либо используйте CLI:
claude mcp add -s user artificer -- node D:/path/to/fvtt-mcp-artificer/dist/index.jsили скопируйте .mcp.json.example в .mcp.json, который читает Claude Code (или объедините с mcpServers в ~/.claude.json) с абсолютными путями. В Windows укажите в command полный путь к node.exe, если Node нет в PATH.
Конфигурация
Скопируйте .env.example в .env (игнорируется git):
COMFY_URL— headless-инстанс (по умолчаниюhttp://127.0.0.1:8188).COMFY_OUTPUT_DIR— должен соответствовать--output-directory, с которым был запущен ComfyUI; сервер читает сгенерированные PNG прямо из этого пути.ARTIFICER_TIMEOUT_MS— максимальное время ожидания на задание (по умолчанию 300000).
Лицензия
Лицензия MIT — подробнее см. в LICENSE.
Available Tools
4 toolsartificer-statusA
Health check: API key present, which image models the key can reach, the output directory, and estimated session spend by tier. Call this first on a cold start.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It effectively discloses that this is an informational read of several system states and gives useful specifics. However, it does not describe the return format, failure behavior, or whether the call itself has any side effects or costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences. The first packs the full scope of the health check into a comma-separated list, and the second delivers the call guidance. No filler words or repeated schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless health check with no annotations and no output schema, the description is nearly sufficient. It states what is checked and when to call it, which is enough to select and invoke the tool. The only gap is that it doesn't explicitly describe the shape of the returned status report, but that is minor for such a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so there is no parameter meaning to convey. The description adds nothing about parameters because none exist; the baseline of 4 for zero-parameter tools applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Health check') and enumerates the exact resources inspected: API key presence, image model reachability, output directory, and session spend. This clearly distinguishes it from the sibling image generation/editing tools without needing to inspect their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use the tool: 'Call this first on a cold start.' This gives clear usage context. It does not name alternatives or state when-not-to-use, but the distinct nature of the image-operation siblings makes the exclusion obvious enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cutout-imageA
Knock the background off a token image to real alpha and deliver it centred on a 512 square so Foundry scale 1.0 is right. Writes a magenta-composited *_preview.png beside it: READ THAT before trusting the edge. Returns coverage and residual-key numbers; a cut outside sane coverage falls back to the rembg AI matte automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Square canvas edge; default 512 (Foundry scale 1.0). 0 keeps the source canvas. | |
| trim | No | Tighten to the subject before fitting (default true); false letterboxes as-is. | |
| color | No | Chroma key colour: "green", "magenta", "blue", or #RRGGBB. Omit to sample the corners. | |
| erode | No | Shrink the matte N px to eat a fringe. | |
| method | No | auto (default): chroma if the plate is a flat key colour, with a rembg fallback when the cut fails verification. chroma: flat green/blue/magenta/solid plates, instant. rembg: AI matte for busy backgrounds, hair, and soft edges (first use downloads a ~176 MB model). | |
| output | No | Absolute output path (.png). Default: next to the source as <name>-cut.png. | |
| padPct | No | Transparent margin, % of the edge (4). | |
| dropShadow | No | Add the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24). | |
| keepShadow | No | chroma only: keep a cast shadow on the plate. | |
| sourceImage | Yes | Absolute path of the image to cut (PNG/JPEG/WebP). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and meets it well. It discloses that a *_preview.png file is written, that the tool returns coverage and residual-key numbers, that unsafe cuts automatically fall back to rembg, and the schema additionally notes the ~176 MB model download on first rembg use. Nothing about the tool's side effects is hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense sentences with no filler. The primary purpose is front-loaded, followed immediately by the most important safety caveat (verify with the preview), then fallback behavior. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter image-processing tool with no output schema and no annotations, the description plus rich schema covers nearly everything: purpose, output file, verification step, return metrics, and failure fallback. It is slightly jargon-heavy ('sane coverage', 'residual-key numbers') and does not define thresholds, but an agent can still select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The prose does reinforce some semantics like 'Foundry scale 1.0' and the preview-file caveat, but it does not add meaning to individual parameters beyond what the schema already provides. The schema's rich parameter descriptions carry the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Knock the background off a token image'), then states the exact output contract: real alpha, centred on a 512 square, Foundry scale 1.0. It clearly distinguishes this tool from siblings like generate-image and edit-image by focusing on background removal and alpha output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete operational guidance: read the magenta-composited preview before trusting edges, and be aware of automatic fallback to rembg when the cut fails verification. The method parameter further explains when to use chroma vs rembg. It does not explicitly compare against sibling tools, but those are not close alternatives, so this is a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit-imageA
Edit an existing image with one instruction while keeping identity, pose, angle, and style. Flash for every kind (pro was no better at fixes and re-cropped once). Tokens are prompted light (your instruction as you would type it in the Gemini app, plus a keep-face/hair/angle line and "remove any cast shadow"), put back on a chroma plate keyed to the token's own colours, and cut to alpha on the 512 square in the same call. "give this an updated painterly style" restyles a world token in place. Props (kind "prop") get object-only wording (no figures added) and come back cut at the source file's exact pixel size, ready for the same tile slot. Returns the new file path, dimensions, and estimated spend.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro. | |
| slug | Yes | Kebab-cased into the filename: <kind>-<slug>-<id>.png. | |
| tier | No | Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro. | |
| confirmPro | No | Required true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it. | |
| references | No | Optional extra references (attached after the source; indexes start at 2). | |
| instruction | Yes | The change, and only the change: "replace the greatsword with a war maul crackling with violet energy". For a flaw-fix pass, name every flaw precisely in one instruction ("the left peryton has four legs; give it two", "remove the second fireball") and end with "keep everything else identical". Everything else is kept by the tool's own wording. | |
| sourceImage | Yes | Absolute path of the image to edit (PNG/JPEG/WebP). | |
| creatureSize | No | Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses preservation of identity/pose/angle/style, default flash behavior, auto chroma-plating/cutting for tokens, prop pixel-size preservation, and returns new file path, dimensions, and estimated spend. It stops short of stating auth or side effects on the source file, but 'new file path' implies non-destructive output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a good front-loaded summary, but the rest is a rambling, run-on explanation mixing model behavior, usage tips, and process internals ('put back on a chroma plate keyed to the token's own colours'). Several phrases ('pro was no better at fixes and re-cropped once', 'give this an updated painterly style') are vague and could be cut or reorganized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 8 parameters, no output schema, and no annotations, yet the description covers return information and kind-specific behavior. It is not complete enough for an agent to fully predict behavior—e.g., it doesn't state whether the original is preserved, what happens on failure, or when pro is genuinely preferred despite the flash recommendation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and already very detailed, so the baseline is 3. The description adds value beyond the schema by giving instruction-phrasing guidance for tokens ('prompted light ... plus a keep-face/hair/angle line') and for props ('object-only wording'), which helps the agent construct a correct instruction value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Edit an existing image with one instruction while keeping identity, pose, angle, and style.' This clearly separates the tool from generate-image (creation) and cutout-image (isolation), and the rest of the description fills in kind-specific behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: it is for editing an existing image, flash is the default, tokens/props have specific wording guidance, and the returned payload is described. It does not explicitly name sibling tools or state when not to use this tool, but the 'existing image' framing makes the primary use case unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate-imageA
Generate one Foundry art asset from a prompt via the Gemini image API. kind picks the model tier, aspect, size, framing text, and post-processing; the result is a finished PNG on disk. READ IT before showing anyone: count limbs per creature, check for duplicated spell effects or props, stray signatures, and reference faces on the wrong figure; obvious flaws are one edit-image call away. Every kind runs on flash by default; tier: "pro" refuses without confirmPro: true and states the cost. Returns the file path, dimensions, and estimated spend.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro. | |
| slug | Yes | Kebab-cased into the filename: <kind>-<slug>-<id>.png. | |
| tier | No | Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro. | |
| prompt | Yes | What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it. | |
| footprint | No | Props only: grid cells wide x tall, e.g. "2x1" for a table (default "1x1"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: "TC_Anvil 02_2x1.png". | |
| confirmPro | No | Required true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it. | |
| references | No | Reference images, attached in this order. Bind each in the prompt by its label or 1-based index ("Image 2 is Morgash"). | |
| creatureSize | No | Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool writes a PNG to disk, that pro tier refuses without confirmPro and states cost, that every kind defaults to flash, and that post-processing (alpha cut, framing, plate) is done automatically. It also warns to inspect the result before showing anyone. It does not mention rate limits or failure modes, but the behavioral traits that matter for invocation are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: a one-sentence purpose, a QA warning, a tier/cost note, and a return summary. It front-loads the core purpose and the most important behavioral caveat (read before showing). It is longer than ideal, but every sentence adds operational value; the only minor redundancy is repeating flash default and confirmPro, which also appear in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema and no annotations, the description is quite complete: it covers the output (file path, dimensions, estimated spend), the tier gating, the QA expectation, and the per-kind behavior. It does not describe error cases or what happens when references are invalid, but the schema already documents reference roles and limits. The description is sufficient for an agent to invoke the tool correctly in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema: it explains that kind picks model tier, aspect, size, framing text, and post-processing; it clarifies that icons/tokens get framing/background appended automatically; it explains the pro tier cost and confirmPro requirement; and it gives concrete output dimensions. This goes beyond the schema's field descriptions, so a 4 is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Generate one Foundry art asset from a prompt via the Gemini image API.' It names the output (finished PNG on disk) and distinguishes the tool from siblings by mentioning edit-image as a follow-up for fixing flaws. The kind parameter further clarifies the five asset types, so an agent can tell this apart from edit-image or cutout-image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool (generate a new asset) and when to use edit-image ('obvious flaws are one edit-image call away'). It also gives usage context for pro tier: 'refuses without confirmPro: true and states the cost,' and instructs to offer pro only as an option to the owner. This is strong routing guidance relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.1.0- Changed
cutout-image1 field changed- added
Input schema / properties / dropShadowAdded value: +{ + "description": "Add the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24).", + "type": "boolean" +}
- Changed
edit-image5 fields changed- added
Input schema / properties / creatureSizeAdded value: +{ + "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.", + "enum": [ + "medium", + "large" + ], + "type": "string" +} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "icon", - "token", - "portrait", - "illustration" -]New value: +[ + "icon", + "token", + "prop", + "portrait", + "illustration" +] - changed
Input schema / properties / references / items / properties / role / descriptionPrevious value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role." - changed
Input schema / properties / references / items / properties / role / enumPrevious value: -[ - "character", - "style" -]New value: +[ + "character", + "style", + "pose" +]
- Changed
generate-image6 fields changed- added
Input schema / properties / creatureSizeAdded value: +{ + "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.", + "enum": [ + "medium", + "large" + ], + "type": "string" +} - added
Input schema / properties / footprintAdded value: +{ + "description": "Props only: grid cells wide x tall, e.g. \"2x1\" for a table (default \"1x1\"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: \"TC_Anvil 02_2x1.png\".", + "pattern": "^\\d{1,2}x\\d{1,2}$", + "type": "string" +} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "icon", - "token", - "portrait", - "illustration" -]New value: +[ + "icon", + "token", + "prop", + "portrait", + "illustration" +] - changed
Input schema / properties / references / items / properties / role / descriptionPrevious value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role." - changed
Input schema / properties / references / items / properties / role / enumPrevious value: -[ - "character", - "style" -]New value: +[ + "character", + "style", + "pose" +]
4 tool updates
v1.0.0- Added
cutout-image - Added
edit-image - Changed
generate-image12 fields changed- removed
Input schema / properties / batchRemoved value: -{ - "default": 6, - "description": "Draft mode only: images per batch.", - "maximum": 8, - "minimum": 1, - "type": "integer" -} - added
Input schema / properties / confirmProAdded value: +{ + "description": "Required true with tier: \"pro\". Offer pro to the owner as an option for portraits and illustrations (\"pro is available for a bit extra\"); never assume it.", + "type": "boolean" +} - removed
Input schema / properties / denoiseRemoved value: -{ - "default": 0.7, - "description": "Refine mode only. 0.7 (pinned by test) keeps the scene skeleton in dev style; ~0.55 clones composition but inherits the draft rendering style.", - "maximum": 0.95, - "minimum": 0.3, - "type": "number" -} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset — fixes generation and output resolution. No raw dimensions."New value: +"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "handout", - "scene-background", - "portrait", - "token" -]New value: +[ + "icon", + "token", + "portrait", + "illustration" +] - removed
Input schema / properties / modeRemoved value: -{ - "default": "draft", - "description": "draft: fast klein batch for curation. final: dev-quality render from the prompt alone, finished at output resolution. refine: dev img2img over sourceImage (a picked draft) — keeps its scene skeleton, re-renders in dev style, finished at output resolution.", - "enum": [ - "draft", - "final", - "refine" - ], - "type": "string" -} - changed
Input schema / properties / prompt / descriptionPrevious value: -"The full image prompt."New value: +"What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it." - added
Input schema / properties / referencesAdded value: +{ + "description": "Reference images, attached in this order. Bind each in the prompt by its label or 1-based index (\"Image 2 is Morgash\").", + "items": { + "properties": { + "label": { + "description": "Short name used to bind the reference in the prompt, e.g. \"Morgash\".", + "type": "string" + }, + "path": { + "description": "Absolute path of a PNG/JPEG on disk.", + "minLength": 1, + "type": "string" + }, + "role": { + "description": "character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice).", + "enum": [ + "character", + "style" + ], + "type": "string" + } + }, + "required": [ + "path", + "role" + ], + "type": "object" + }, + "maxItems": 14, + "type": "array" +} - removed
Input schema / properties / seedRemoved value: -{ - "description": "Fixed seed; random when omitted.", - "minimum": 0, - "type": "integer" -} - changed
Input schema / properties / slug / descriptionPrevious value: -"Short kebab-case subject name used in output filenames, e.g. \"smugglers-cove\"."New value: +"Kebab-cased into the filename: <kind>-<slug>-<id>.png." - removed
Input schema / properties / sourceImageRemoved value: -{ - "description": "Refine mode only (required there): absolute path of the picked draft PNG.", - "type": "string" -} - added
Input schema / properties / tierAdded value: +{ + "description": "Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. \"pro\" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.", + "enum": [ + "flash", + "pro" + ], + "type": "string" +}
- Removed
upscale-image
3 tool updates
v0.1.0- First observed
artificer-status - First observed
generate-image - First observed
upscale-image
TDQS
Scored across 4 tools
The tools are mostly distinct: generate-image creates new assets, edit-image modifies existing ones, cutout-image handles background removal, and artificer-status is a health check. However, edit-image and generate-image could be slightly confused since both can produce a final PNG, though the descriptions clarify the difference (new vs. existing).
All tool names follow a consistent verb_noun pattern (generate-image, edit-image, cutout-image, artificer-status). The pattern is uniform and predictable, making it easy for an agent to infer the action and target.
With only 4 tools, the server is tightly scoped to the image generation and processing workflow. Each tool serves a clear, necessary function, and the count is appropriate for the server's stated purpose—creating and preparing Foundry art assets.
The set covers the core lifecycle: generate, edit, cutout, and status. However, there are gaps—no explicit upload to Foundry, no batch processing, and no tool to list or delete existing assets. The documentation mentions previews and fallbacks, but the surface is missing endpoints for managing the output directory or integrating with Foundry beyond saving to disk.
Maintenance
Related MCP Connectors
Connect any AI to your Foundry VTT world: actors, combat, dice, journals, tokens, compendiums.
Generate logos, social posts, app screenshots, comic panels & visual-novel assets from prompts.
Remote MCP for RunComfy: ComfyUI deployments, hosted models, LoRA training. 31 tools.
Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceConnects Claude Desktop to Foundry VTT for AI-powered campaign management, enabling natural language interaction with game data including quest creation, character management, compendium searches, and dice rolling. Provides 20 MCP tools for seamless integration between Claude and your tabletop RPG sessions.72-
- FlicenseNot gradedqualityNot gradedmaintenanceIntegrates with FoundryVTT tabletop gaming sessions, allowing AI assistants to query game data, roll dice, generate content (NPCs, loot, encounters), manage combat, and provide tactical suggestions through natural language.7 npm-
- FlicenseNot gradedqualityFmaintenanceEnables AI-powered campaign management for Foundry Virtual Tabletop through natural language, supporting multiple RPG systems with tools for quest creation, character management, combat resolution, and more.-
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to interact with Foundry Virtual Tabletop, supporting reading world data, managing combat, rolling dice, and updating actor attributes via a sidecar architecture.67 npmMIT