fvtt-mcp-artificer
fvtt-mcp-artificer
一个专门面向 Foundry 的图像生成 Model Context Protocol 服务器,用于 D&D 桌面美术,由 Claude Code 驱动。它封装了本机上的无头 ComfyUI 实例,并暴露一组少量 Foundry 形态的工具,让 Claude 可以编写提示词、生成插画批次、通过实际查看结果来策选,并把入选结果交给其姊妹服务器 fvtt-mcp-molten5e 中的 Foundry 管线(upload-asset → set-actor-art / add-journal-image / 场景背景)。
整个闭环在目标硬件(RTX 5090)上足以进行对话式交互:6 张草稿批次约 10 秒完成,一张 2560×1600 成品渲染约 19 秒,而完整的 提示词 → 草稿 → 策选 → 成品 → 游戏内日志 循环已在 notes/m3-loop-proof.md 中端到端验证。
为什么采用这种形态
这不是通用的 ComfyUI 桥接器,这是硬性规定。工具使用 Foundry 的词汇——角色肖像、token、handout、场景背景——并且可以自由调整以适配 Foundry 工作。它与 fvtt-mcp-molten5e 保持分离:那个服务器只负责 Foundry 内容创作,绝不能与图像生成耦合。本服务器从不与 Foundry 桥接器通信;两者之间的交接是磁盘上的文件加上 molten5e 的上传工具。
与家族其他项目相同的核心理念:工具负责执行,技能负责决策。 正确性(工作流执行、尺寸、放大管线、文件约定)存在于这里经过测试的工具中;判断力(提示词功力、策选品味、哪个 actor 或 journal 获得美术、家室风格)则属于稍后的 illustration-builder 技能。
Claude ──MCP──> fvtt-mcp-artificer ──HTTP──> ComfyUI (headless, local)
│
└── pinned workflow JSONs (draft / final / final-refine / upscale)ComfyUI 以 API 模式无头运行;本服务器提交 workflows/ 下的固定工作流 JSON——绝不提交自由形式的工作流图——并通过节点 ID 替换 prompt/seed/batch/preset(以 class-type 漂移断言加以保护)。工具返回绝对文件路径;Claude 直接读取 PNG 进行策选。
Related MCP server: FoundryVTT MCP Server
用途预设,而非原始尺寸
generate-image 接受 kind,而非宽高:
kind | 生成分辨率 | 成品分辨率 |
| 1536×960 | 2560×1600 |
| 1024×1280 | 2048×2560 |
| 1024×1024 | 2048×2048 |
分辨率管线(锁定): 绝不在输出尺寸下生成——构图超过约 1.5 MP 后会退化。在预设的原始分辨率下生成,使用 4x-UltraSharp 模型放大 ×4,lanczos 降采样到成品尺寸,全部在一个固定工作流图中完成。
模型
FLUX.1-dev fp8(Comfy-Org all-in-one)——最终品质的渲染,约 15–19 秒完成。
FLUX.2-klein 4B(Apache 2.0)——草稿模型:4 步,约 1–2 秒/张,批量 6–8 张,挑选后用 dev 重新渲染。
4x-UltraSharp——流程尾部的模型放大器。
工具
工具 | 作用 |
| 接受 |
| 将现有图像(通常是未能直接通过策选的草稿)经过放大流程尾部处理到其 kind 对应的成品分辨率。 |
| 健康检查:ComfyUI 可达性/版本、显存、队列深度、固定工作流完整性、所需模型是否存在。 |
工具与固定工作流图之间的替换契约记录在 workflows/README.md 中。
要求
Windows + NVIDIA GPU。 在 RTX 5090 上开发并验证(Blackwell 需要 CUDA 12.8+ 的 PyTorch;当前 ComfyUI 便携版已附带该版本)。
ComfyUI(独立便携版,v0.34+)以无头模式运行——启动约定见
scripts/launch-comfyui.ps1(API 模式监听127.0.0.1:8188,固定--output-directory)。ComfyUI
models/目录中的模型文件(约 33 GB,全部无需门控):checkpoints/flux1-dev-fp8.safetensors、diffusion_models/flux-2-klein-4b.safetensors、text_encoders/qwen_3_4b.safetensors、vae/flux2-vae.safetensors、upscale_models/4x-UltraSharp.safetensors。artificer-status会报告缺失项。Node.js 22+ 用于运行 MCP 服务器本身。
构建
npm install
npm run build测试:npm test(离线单元测试套件;固定工作流 JSON 作为测试夹具)。实时套件——真实 ComfyUI、真实渲染——由 npm run test:integration 门控。质量关卡:npm run typecheck、npm run check(biome)、npm run knip。
接入 Claude Code
家族惯例是在用户作用域注册服务器(新增 MCP 工具后必须重启 Claude Code)。可以使用 CLI:
claude mcp add -s user artificer -- node D:/path/to/fvtt-mcp-artificer/dist/index.js或者将 .mcp.json.example 复制成 Claude Code 读取的 .mcp.json(或合并到 ~/.claude.json 的 mcpServers),路径使用绝对路径。在 Windows 上,如果 Node 不在 PATH 中,请将 command 指向完整的 node.exe 路径。
配置
将 .env.example 复制为 .env(已被 gitignore):
COMFY_URL——无头实例地址(默认http://127.0.0.1:8188)。COMFY_OUTPUT_DIR——必须与 ComfyUI 启动时使用的--output-directory一致;服务器直接从该路径读取生成的 PNG。ARTIFICER_TIMEOUT_MS——每个任务的等待上限(默认 300000)。
许可证
MIT License——详见 LICENSE。
Available Tools
4 toolsartificer-statusA
Health check: API key present, which image models the key can reach, the output directory, and estimated session spend by tier. Call this first on a cold start.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It effectively discloses that this is an informational read of several system states and gives useful specifics. However, it does not describe the return format, failure behavior, or whether the call itself has any side effects or costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences. The first packs the full scope of the health check into a comma-separated list, and the second delivers the call guidance. No filler words or repeated schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless health check with no annotations and no output schema, the description is nearly sufficient. It states what is checked and when to call it, which is enough to select and invoke the tool. The only gap is that it doesn't explicitly describe the shape of the returned status report, but that is minor for such a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so there is no parameter meaning to convey. The description adds nothing about parameters because none exist; the baseline of 4 for zero-parameter tools applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Health check') and enumerates the exact resources inspected: API key presence, image model reachability, output directory, and session spend. This clearly distinguishes it from the sibling image generation/editing tools without needing to inspect their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use the tool: 'Call this first on a cold start.' This gives clear usage context. It does not name alternatives or state when-not-to-use, but the distinct nature of the image-operation siblings makes the exclusion obvious enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cutout-imageA
Knock the background off a token image to real alpha and deliver it centred on a 512 square so Foundry scale 1.0 is right. Writes a magenta-composited *_preview.png beside it: READ THAT before trusting the edge. Returns coverage and residual-key numbers; a cut outside sane coverage falls back to the rembg AI matte automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Square canvas edge; default 512 (Foundry scale 1.0). 0 keeps the source canvas. | |
| trim | No | Tighten to the subject before fitting (default true); false letterboxes as-is. | |
| color | No | Chroma key colour: "green", "magenta", "blue", or #RRGGBB. Omit to sample the corners. | |
| erode | No | Shrink the matte N px to eat a fringe. | |
| method | No | auto (default): chroma if the plate is a flat key colour, with a rembg fallback when the cut fails verification. chroma: flat green/blue/magenta/solid plates, instant. rembg: AI matte for busy backgrounds, hair, and soft edges (first use downloads a ~176 MB model). | |
| output | No | Absolute output path (.png). Default: next to the source as <name>-cut.png. | |
| padPct | No | Transparent margin, % of the edge (4). | |
| dropShadow | No | Add the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24). | |
| keepShadow | No | chroma only: keep a cast shadow on the plate. | |
| sourceImage | Yes | Absolute path of the image to cut (PNG/JPEG/WebP). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and meets it well. It discloses that a *_preview.png file is written, that the tool returns coverage and residual-key numbers, that unsafe cuts automatically fall back to rembg, and the schema additionally notes the ~176 MB model download on first rembg use. Nothing about the tool's side effects is hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense sentences with no filler. The primary purpose is front-loaded, followed immediately by the most important safety caveat (verify with the preview), then fallback behavior. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter image-processing tool with no output schema and no annotations, the description plus rich schema covers nearly everything: purpose, output file, verification step, return metrics, and failure fallback. It is slightly jargon-heavy ('sane coverage', 'residual-key numbers') and does not define thresholds, but an agent can still select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The prose does reinforce some semantics like 'Foundry scale 1.0' and the preview-file caveat, but it does not add meaning to individual parameters beyond what the schema already provides. The schema's rich parameter descriptions carry the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Knock the background off a token image'), then states the exact output contract: real alpha, centred on a 512 square, Foundry scale 1.0. It clearly distinguishes this tool from siblings like generate-image and edit-image by focusing on background removal and alpha output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete operational guidance: read the magenta-composited preview before trusting edges, and be aware of automatic fallback to rembg when the cut fails verification. The method parameter further explains when to use chroma vs rembg. It does not explicitly compare against sibling tools, but those are not close alternatives, so this is a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit-imageA
Edit an existing image with one instruction while keeping identity, pose, angle, and style. Flash for every kind (pro was no better at fixes and re-cropped once). Tokens are prompted light (your instruction as you would type it in the Gemini app, plus a keep-face/hair/angle line and "remove any cast shadow"), put back on a chroma plate keyed to the token's own colours, and cut to alpha on the 512 square in the same call. "give this an updated painterly style" restyles a world token in place. Props (kind "prop") get object-only wording (no figures added) and come back cut at the source file's exact pixel size, ready for the same tile slot. Returns the new file path, dimensions, and estimated spend.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro. | |
| slug | Yes | Kebab-cased into the filename: <kind>-<slug>-<id>.png. | |
| tier | No | Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro. | |
| confirmPro | No | Required true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it. | |
| references | No | Optional extra references (attached after the source; indexes start at 2). | |
| instruction | Yes | The change, and only the change: "replace the greatsword with a war maul crackling with violet energy". For a flaw-fix pass, name every flaw precisely in one instruction ("the left peryton has four legs; give it two", "remove the second fireball") and end with "keep everything else identical". Everything else is kept by the tool's own wording. | |
| sourceImage | Yes | Absolute path of the image to edit (PNG/JPEG/WebP). | |
| creatureSize | No | Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses preservation of identity/pose/angle/style, default flash behavior, auto chroma-plating/cutting for tokens, prop pixel-size preservation, and returns new file path, dimensions, and estimated spend. It stops short of stating auth or side effects on the source file, but 'new file path' implies non-destructive output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a good front-loaded summary, but the rest is a rambling, run-on explanation mixing model behavior, usage tips, and process internals ('put back on a chroma plate keyed to the token's own colours'). Several phrases ('pro was no better at fixes and re-cropped once', 'give this an updated painterly style') are vague and could be cut or reorganized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 8 parameters, no output schema, and no annotations, yet the description covers return information and kind-specific behavior. It is not complete enough for an agent to fully predict behavior—e.g., it doesn't state whether the original is preserved, what happens on failure, or when pro is genuinely preferred despite the flash recommendation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and already very detailed, so the baseline is 3. The description adds value beyond the schema by giving instruction-phrasing guidance for tokens ('prompted light ... plus a keep-face/hair/angle line') and for props ('object-only wording'), which helps the agent construct a correct instruction value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Edit an existing image with one instruction while keeping identity, pose, angle, and style.' This clearly separates the tool from generate-image (creation) and cutout-image (isolation), and the rest of the description fills in kind-specific behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: it is for editing an existing image, flash is the default, tokens/props have specific wording guidance, and the returned payload is described. It does not explicitly name sibling tools or state when not to use this tool, but the 'existing image' framing makes the primary use case unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate-imageA
Generate one Foundry art asset from a prompt via the Gemini image API. kind picks the model tier, aspect, size, framing text, and post-processing; the result is a finished PNG on disk. READ IT before showing anyone: count limbs per creature, check for duplicated spell effects or props, stray signatures, and reference faces on the wrong figure; obvious flaws are one edit-image call away. Every kind runs on flash by default; tier: "pro" refuses without confirmPro: true and states the cost. Returns the file path, dimensions, and estimated spend.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: "large"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro. | |
| slug | Yes | Kebab-cased into the filename: <kind>-<slug>-<id>.png. | |
| tier | No | Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro. | |
| prompt | Yes | What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it. | |
| footprint | No | Props only: grid cells wide x tall, e.g. "2x1" for a table (default "1x1"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: "TC_Anvil 02_2x1.png". | |
| confirmPro | No | Required true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it. | |
| references | No | Reference images, attached in this order. Bind each in the prompt by its label or 1-based index ("Image 2 is Morgash"). | |
| creatureSize | No | Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool writes a PNG to disk, that pro tier refuses without confirmPro and states cost, that every kind defaults to flash, and that post-processing (alpha cut, framing, plate) is done automatically. It also warns to inspect the result before showing anyone. It does not mention rate limits or failure modes, but the behavioral traits that matter for invocation are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: a one-sentence purpose, a QA warning, a tier/cost note, and a return summary. It front-loads the core purpose and the most important behavioral caveat (read before showing). It is longer than ideal, but every sentence adds operational value; the only minor redundancy is repeating flash default and confirmPro, which also appear in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema and no annotations, the description is quite complete: it covers the output (file path, dimensions, estimated spend), the tier gating, the QA expectation, and the per-kind behavior. It does not describe error cases or what happens when references are invalid, but the schema already documents reference roles and limits. The description is sufficient for an agent to invoke the tool correctly in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema: it explains that kind picks model tier, aspect, size, framing text, and post-processing; it clarifies that icons/tokens get framing/background appended automatically; it explains the pro tier cost and confirmPro requirement; and it gives concrete output dimensions. This goes beyond the schema's field descriptions, so a 4 is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Generate one Foundry art asset from a prompt via the Gemini image API.' It names the output (finished PNG on disk) and distinguishes the tool from siblings by mentioning edit-image as a follow-up for fixing flaws. The kind parameter further clarifies the five asset types, so an agent can tell this apart from edit-image or cutout-image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool (generate a new asset) and when to use edit-image ('obvious flaws are one edit-image call away'). It also gives usage context for pro tier: 'refuses without confirmPro: true and states the cost,' and instructs to offer pro only as an option to the owner. This is strong routing guidance relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.1.0- Changed
cutout-image1 field changed- added
Input schema / properties / dropShadowAdded value: +{ + "description": "Add the world tokens' soft cast shadow (dark silhouette, ~38%, down-right) under the cut. Off by default; tokens from this server go out shadowless (owner rule 2026-09-24).", + "type": "boolean" +}
- Changed
edit-image5 fields changed- added
Input schema / properties / creatureSizeAdded value: +{ + "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.", + "enum": [ + "medium", + "large" + ], + "type": "string" +} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "icon", - "token", - "portrait", - "illustration" -]New value: +[ + "icon", + "token", + "prop", + "portrait", + "illustration" +] - changed
Input schema / properties / references / items / properties / role / descriptionPrevious value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role." - changed
Input schema / properties / references / items / properties / role / enumPrevious value: -[ - "character", - "style" -]New value: +[ + "character", + "style", + "pose" +]
- Changed
generate-image6 fields changed- added
Input schema / properties / creatureSizeAdded value: +{ + "description": "Tokens only. medium (default): one grid cell, Tiny through Medium, a 512 square. large: Large, Huge, or Gargantuan, a 1024 square, the native render edge.", + "enum": [ + "medium", + "large" + ], + "type": "string" +} - added
Input schema / properties / footprintAdded value: +{ + "description": "Props only: grid cells wide x tall, e.g. \"2x1\" for a table (default \"1x1\"). The prop is rendered at the nearest API aspect and delivered at 300 px per cell (600x300 here). Library files carry it in their name: \"TC_Anvil 02_2x1.png\".", + "pattern": "^\\d{1,2}x\\d{1,2}$", + "type": "string" +} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro."New value: +"Purpose preset. icon: 1:1 flash → 512 square. prop: a map prop (furniture, barrel, tree) seen straight down, object only, cut to alpha at its tile size (300 px per grid cell; an edit keeps the source size). token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (1024 for creatureSize: \"large\"), no shadow (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "icon", - "token", - "portrait", - "illustration" -]New value: +[ + "icon", + "token", + "prop", + "portrait", + "illustration" +] - changed
Input schema / properties / references / items / properties / role / descriptionPrevious value: -"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice)."New value: +"character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice). pose: match only its pose, head direction, camera angle, and silhouette, never its drawing; for replacing a weak token, attach the old one as the ONLY image with this role." - changed
Input schema / properties / references / items / properties / role / enumPrevious value: -[ - "character", - "style" -]New value: +[ + "character", + "style", + "pose" +]
4 tool updates
v1.0.0- Added
cutout-image - Added
edit-image - Changed
generate-image12 fields changed- removed
Input schema / properties / batchRemoved value: -{ - "default": 6, - "description": "Draft mode only: images per batch.", - "maximum": 8, - "minimum": 1, - "type": "integer" -} - added
Input schema / properties / confirmProAdded value: +{ + "description": "Required true with tier: \"pro\". Offer pro to the owner as an option for portraits and illustrations (\"pro is available for a bit extra\"); never assume it.", + "type": "boolean" +} - removed
Input schema / properties / denoiseRemoved value: -{ - "default": 0.7, - "description": "Refine mode only. 0.7 (pinned by test) keeps the scene skeleton in dev style; ~0.55 clones composition but inherits the draft rendering style.", - "maximum": 0.95, - "minimum": 0.3, - "type": "number" -} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset — fixes generation and output resolution. No raw dimensions."New value: +"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "handout", - "scene-background", - "portrait", - "token" -]New value: +[ + "icon", + "token", + "portrait", + "illustration" +] - removed
Input schema / properties / modeRemoved value: -{ - "default": "draft", - "description": "draft: fast klein batch for curation. final: dev-quality render from the prompt alone, finished at output resolution. refine: dev img2img over sourceImage (a picked draft) — keeps its scene skeleton, re-renders in dev style, finished at output resolution.", - "enum": [ - "draft", - "final", - "refine" - ], - "type": "string" -} - changed
Input schema / properties / prompt / descriptionPrevious value: -"The full image prompt."New value: +"What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it." - added
Input schema / properties / referencesAdded value: +{ + "description": "Reference images, attached in this order. Bind each in the prompt by its label or 1-based index (\"Image 2 is Morgash\").", + "items": { + "properties": { + "label": { + "description": "Short name used to bind the reference in the prompt, e.g. \"Morgash\".", + "type": "string" + }, + "path": { + "description": "Absolute path of a PNG/JPEG on disk.", + "minLength": 1, + "type": "string" + }, + "role": { + "description": "character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice).", + "enum": [ + "character", + "style" + ], + "type": "string" + } + }, + "required": [ + "path", + "role" + ], + "type": "object" + }, + "maxItems": 14, + "type": "array" +} - removed
Input schema / properties / seedRemoved value: -{ - "description": "Fixed seed; random when omitted.", - "minimum": 0, - "type": "integer" -} - changed
Input schema / properties / slug / descriptionPrevious value: -"Short kebab-case subject name used in output filenames, e.g. \"smugglers-cove\"."New value: +"Kebab-cased into the filename: <kind>-<slug>-<id>.png." - removed
Input schema / properties / sourceImageRemoved value: -{ - "description": "Refine mode only (required there): absolute path of the picked draft PNG.", - "type": "string" -} - added
Input schema / properties / tierAdded value: +{ + "description": "Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. \"pro\" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.", + "enum": [ + "flash", + "pro" + ], + "type": "string" +}
- Removed
upscale-image
3 tool updates
v0.1.0- First observed
artificer-status - First observed
generate-image - First observed
upscale-image
TDQS
Scored across 4 tools
The tools are mostly distinct: generate-image creates new assets, edit-image modifies existing ones, cutout-image handles background removal, and artificer-status is a health check. However, edit-image and generate-image could be slightly confused since both can produce a final PNG, though the descriptions clarify the difference (new vs. existing).
All tool names follow a consistent verb_noun pattern (generate-image, edit-image, cutout-image, artificer-status). The pattern is uniform and predictable, making it easy for an agent to infer the action and target.
With only 4 tools, the server is tightly scoped to the image generation and processing workflow. Each tool serves a clear, necessary function, and the count is appropriate for the server's stated purpose—creating and preparing Foundry art assets.
The set covers the core lifecycle: generate, edit, cutout, and status. However, there are gaps—no explicit upload to Foundry, no batch processing, and no tool to list or delete existing assets. The documentation mentions previews and fallbacks, but the surface is missing endpoints for managing the output directory or integrating with Foundry beyond saving to disk.
Maintenance
Related MCP Connectors
Connect any AI to your Foundry VTT world: actors, combat, dice, journals, tokens, compendiums.
Generate logos, social posts, app screenshots, comic panels & visual-novel assets from prompts.
Remote MCP for RunComfy: ComfyUI deployments, hosted models, LoRA training. 31 tools.
Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceConnects Claude Desktop to Foundry VTT for AI-powered campaign management, enabling natural language interaction with game data including quest creation, character management, compendium searches, and dice rolling. Provides 20 MCP tools for seamless integration between Claude and your tabletop RPG sessions.72-
- FlicenseNot gradedqualityNot gradedmaintenanceIntegrates with FoundryVTT tabletop gaming sessions, allowing AI assistants to query game data, roll dice, generate content (NPCs, loot, encounters), manage combat, and provide tactical suggestions through natural language.7 npm-
- FlicenseNot gradedqualityFmaintenanceEnables AI-powered campaign management for Foundry Virtual Tabletop through natural language, supporting multiple RPG systems with tools for quest creation, character management, combat resolution, and more.-
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to interact with Foundry Virtual Tabletop, supporting reading world data, managing combat, rolling dice, and updating actor attributes via a sidecar architecture.67 npmMIT