fvtt-mcp-artificer
fvtt-mcp-artificer
Ein Foundry-spezifischer Bildgenerierungs-Model Context Protocol-Server für D&D-Table-Art, angetrieben von Claude Code. Er kapselt eine headless ComfyUI-Instanz auf dem lokalen Rechner und stellt eine kleine Reihe auf Foundry zugeschnittener Werkzeuge bereit, damit Claude Prompts verfassen, Illustrations-Batches erzeugen, die Ergebnisse durch tatsächliches Ansehen kuratieren und die Gewinner an die Foundry-Pipeline in seinem Schwester-Server fvtt-mcp-molten5e übergeben kann (upload-asset → set-actor-art / add-journal-image / Szenen-Hintergründe).
Der gesamte Kreislauf ist auf der Zielhardware (RTX 5090) schnell genug für eine dialogische Nutzung: Ein 6-Bilder-Entwurfsbatch liegt nach ~10 s vor, ein fertiger 2560×1600-Render nach ~19 s, und der vollständige Zyklus Prompt → Entwurf → Kuratieren → Finale → In-World-Journal wurde Ende-zu-Ende in notes/m3-loop-proof.md nachgewiesen.
Warum diese Form
Dies ist keine generische ComfyUI-Brücke, so die Vorgabe. Die Werkzeuge sprechen Foundry-Vokabular — Heldenporträts, Token, Handouts, Szenen-Hintergründe — und dürfen für die Foundry-Arbeit frei angepasst werden. Und es bleibt getrennt von fvtt-mcp-molten5e: Dieser Server ist auf die Erstellung von Foundry-Inhalten ausgerichtet und darf nicht an die Bildgenerierung gekoppelt werden. Dieser Server spricht nie mit der Foundry-Brücke; die Übergabe zwischen ihnen erfolgt über Dateien auf der Festplatte plus die Molten5e-Upload-Werkzeuge.
Gleiche Hausphilosophie wie im Rest der Familie: Werkzeuge tun, Skills entscheiden. Korrektheit (Workflow-Ausführung, Abmessungen, die Upscale-Pipeline, Dateikonventionen) steckt hier in getesteten Werkzeugen; Urteilsvermögen (Prompt-Handwerk, Geschmack beim Kuratieren, welcher Actor oder welches Journal die Kunst bekommt, Hausstil) steckt in einem späteren illustration-builder-Skill.
Claude ──MCP──> fvtt-mcp-artificer ──HTTP──> ComfyUI (headless, local)
│
└── pinned workflow JSONs (draft / final / final-refine / upscale)ComfyUI läuft headless im API-Modus; dieser Server übermittelt die festgelegten Workflow-JSONs unter workflows/ — niemals frei formulierte Graphen — mit eingesetzten Prompt/Seed/Batch/Preset-Werten (per Knoten-ID, abgesichert durch Klassentyp-Drift-Assertions). Die Werkzeuge geben absolute Dateipfade zurück; Claude liest die PNGs direkt, um zu kuratieren.
Related MCP server: FoundryVTT MCP Server
Zweck-Presets statt roher Abmessungen
generate-image akzeptiert eine kind, nicht Breite/Höhe:
kind | erzeugt bei | fertige Ausgabe |
| 1536×960 | 2560×1600 |
| 1024×1280 | 2048×2560 |
| 1024×1024 | 2048×2048 |
Auflösungs-Pipeline (festgelegt): niemals in Ausgabegröße erzeugen — die Komposition verschlechtert sich jenseits von ~1,5 MP. In der nativen Auflösung des Presets erzeugen, per Modell ×4 mit 4x-UltraSharp hochskalieren, per Lanczos auf die Endgröße verkleinern, alles in einem festgelegten Graphen.
Modelle
FLUX.1-dev fp8 (Comfy-Org all-in-one) — Renderings in Endqualität, ~15–19 s fertig.
FLUX.2-klein 4B (Apache 2.0) — das Entwurfsmodell: 4 Schritte, ~1–2 s/Bild, Batch 6–8, auswählen, mit dev neu rendern.
4x-UltraSharp — der Modell-Upscaler am Ende der Pipeline.
Werkzeuge
Werkzeug | was es tut |
|
|
| Ein vorhandenes Bild (normalerweise ein Entwurf, der direkt die Kuratierung gewonnen hat) durch den Upscale-Teil am Ende auf die Ausgabeauflösung seiner Art bringen. |
| Gesundheitscheck: ComfyUI-Erreichbarkeit/-Version, VRAM, Warteschlangentiefe, Integrität der festgelegten Workflows, Vorhandensein der benötigten Modelle. |
Der Substitutionsvertrag zwischen den Werkzeugen und den festgelegten Graphen ist in workflows/README.md dokumentiert.
Anforderungen
Windows + NVIDIA-GPU. Entwickelt und nachgewiesen auf einer RTX 5090 (Blackwell benötigt PyTorch mit CUDA 12.8+; das aktuelle ComfyUI-Portable-Build wird damit ausgeliefert).
ComfyUI (Standalone-Portable-Build, v0.34+) headless laufend — siehe
scripts/launch-comfyui.ps1für die Startkonvention (API-Modus auf127.0.0.1:8188, festgelegtes--output-directory).Die Modelldateien im
models/-Verzeichnisbaum von ComfyUI (~33 GB, alle frei verfügbar):checkpoints/flux1-dev-fp8.safetensors,diffusion_models/flux-2-klein-4b.safetensors,text_encoders/qwen_3_4b.safetensors,vae/flux2-vae.safetensors,upscale_models/4x-UltraSharp.safetensors.artificer-statusmeldet alles Fehlende.Node.js 22+ für den MCP-Server selbst.
Build
npm install
npm run buildTests: npm test (Offline-Unit-Suite; die festgelegten Workflow-JSONs sind die Fixtures). Die Live-Suite — echtes ComfyUI, echte Renderings — ist nur über npm run test:integration erreichbar. Qualitäts-Gates: npm run typecheck, npm run check (biome), npm run knip.
Einbindung in Claude Code
Die Hauskonvention registriert den Server auf Benutzerebene (neue MCP-Werkzeuge ⇒ Claude Code neu starten). Entweder die CLI verwenden:
claude mcp add -s user artificer -- node D:/path/to/fvtt-mcp-artificer/dist/index.jsoder .mcp.json.example in eine .mcp.json kopieren, die Claude Code liest (oder in ~/.claude.json unter mcpServers einfügen), mit absoluten Pfaden. Unter Windows command auf den vollständigen node.exe-Pfad setzen, falls Node nicht im PATH ist.
Konfiguration
Kopieren Sie .env.example nach .env (gitignoriert):
COMFY_URL— die Headless-Instanz (Standardhttp://127.0.0.1:8188).COMFY_OUTPUT_DIR— muss mit dem--output-directoryübereinstimmen, mit dem ComfyUI gestartet wurde; der Server liest erzeugte PNGs direkt von diesem Pfad.ARTIFICER_TIMEOUT_MS— Wartezeit-Obergrenze pro Auftrag (Standard 300000).
Lizenz
MIT-Lizenz — siehe LICENSE für Details.
Available Tools
4 toolsartificer-statusA
Health check: API key present, which image models the key can reach, the output directory, and estimated session spend by tier. Call this first on a cold start.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It effectively discloses that this is an informational read of several system states and gives useful specifics. However, it does not describe the return format, failure behavior, or whether the call itself has any side effects or costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences. The first packs the full scope of the health check into a comma-separated list, and the second delivers the call guidance. No filler words or repeated schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless health check with no annotations and no output schema, the description is nearly sufficient. It states what is checked and when to call it, which is enough to select and invoke the tool. The only gap is that it doesn't explicitly describe the shape of the returned status report, but that is minor for such a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty schema, so there is no parameter meaning to convey. The description adds nothing about parameters because none exist; the baseline of 4 for zero-parameter tools applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Health check') and enumerates the exact resources inspected: API key presence, image model reachability, output directory, and session spend. This clearly distinguishes it from the sibling image generation/editing tools without needing to inspect their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to use the tool: 'Call this first on a cold start.' This gives clear usage context. It does not name alternatives or state when-not-to-use, but the distinct nature of the image-operation siblings makes the exclusion obvious enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cutout-imageA
Knock the background off a token image to real alpha and deliver it centred on a 512 square so Foundry scale 1.0 is right. Writes a magenta-composited *_preview.png beside it: READ THAT before trusting the edge. Returns coverage and residual-key numbers; a cut outside sane coverage falls back to the rembg AI matte automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Square canvas edge; default 512 (Foundry scale 1.0). 0 keeps the source canvas. | |
| trim | No | Tighten to the subject before fitting (default true); false letterboxes as-is. | |
| color | No | Chroma key colour: "green", "magenta", "blue", or #RRGGBB. Omit to sample the corners. | |
| erode | No | Shrink the matte N px to eat a fringe. | |
| method | No | auto (default): chroma if the plate is a flat key colour, with a rembg fallback when the cut fails verification. chroma: flat green/blue/magenta/solid plates, instant. rembg: AI matte for busy backgrounds, hair, and soft edges (first use downloads a ~176 MB model). | |
| output | No | Absolute output path (.png). Default: next to the source as <name>-cut.png. | |
| padPct | No | Transparent margin, % of the edge (4). | |
| keepShadow | No | chroma only: keep a cast shadow on the plate. | |
| sourceImage | Yes | Absolute path of the image to cut (PNG/JPEG/WebP). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does well: it discloses side-effect preview file creation, instructs the user to verify edges, mentions returned coverage/residual-key numbers, and reveals automatic fallback to rembg. It does not mention overwrite behavior or other filesystem side effects beyond the preview and output, but the core behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose and canvas behavior, preview verification warning, and output/fallback behavior. It is front-loaded and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description is unusually complete: it covers processing, output artifact, verification workflow, fallback behavior, and key return values. The remaining gaps are the lack of exact 'sane coverage' thresholds and a precise response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline of 3 applies. The description reinforces the purpose of size and output but adds no extra parameter-level detail beyond what the input schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it removes the background from a token image and produces a transparent cutout centered on a 512-square canvas. This clearly distinguishes it from image generation, general editing, and status-checking siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied clearly through 'token image' and 'Foundry scale 1.0', and the method parameter guidance explains when to prefer chroma versus rembg. However, it never names alternatives or says when not to use this tool versus generate-image or edit-image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit-imageA
Edit an existing image with one instruction while keeping identity, pose, angle, and style. Flash for every kind (pro was no better at fixes and re-cropped once). Tokens get the chroma plate re-applied so they can be cut again. Returns the new file path, dimensions, and estimated spend.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro. | |
| slug | Yes | Kebab-cased into the filename: <kind>-<slug>-<id>.png. | |
| tier | No | Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro. | |
| confirmPro | No | Required true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it. | |
| references | No | Optional extra references (attached after the source; indexes start at 2). | |
| instruction | Yes | The change, and only the change: "replace the greatsword with a war maul crackling with violet energy". For a flaw-fix pass, name every flaw precisely in one instruction ("the left peryton has four legs; give it two", "remove the second fireball") and end with "keep everything else identical". Everything else is kept by the tool's own wording. | |
| sourceImage | Yes | Absolute path of the image to edit (PNG/JPEG/WebP). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral burden and meets it: it discloses that flash is the default for every kind, that pro offered no benefit for fixes and re-cropped, that token edits re-apply the chroma plate for future cutting, and that the response includes file path, dimensions, and estimated spend. These are meaningful behavioral details beyond what the input schema alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four information-dense sentences with no filler. The primary verb and scope are front-loaded, followed by tier behavior, the token-specific quirk, and return values. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with 7 parameters, no annotations, and no output schema, the description covers the essential behavioral contract: return values, default tier behavior, the special token handling, and the 'keep everything else identical' philosophy. The schema covers parameter details, and the sibling context makes the tool's role clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters in detail. The description adds a little context around the instruction ('one instruction' and preserving attributes) but does not materially expand parameter semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pairing ('Edit an existing image') and adds the key constraints: one instruction, preserving identity, pose, angle, and style. This clearly distinguishes edit-image from generate-image and cutout-image by establishing it operates on an existing image rather than creating or extracting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when this tool applies: modifying an existing image while preserving its core attributes. Sibling names like generate-image imply the alternative of creating new images, but the description does not explicitly say 'use this instead of generate-image when the source already exists' or list exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate-imageA
Generate one Foundry art asset from a prompt via the Gemini image API. kind picks the model tier, aspect, size, framing text, and post-processing; the result is a finished PNG on disk. READ IT before showing anyone: count limbs per creature, check for duplicated spell effects or props, stray signatures, and reference faces on the wrong figure; obvious flaws are one edit-image call away. Every kind runs on flash by default; tier: "pro" refuses without confirmPro: true and states the cost. Returns the file path, dimensions, and estimated spend.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: "pro" is opt-in and needs confirmPro. | |
| slug | Yes | Kebab-cased into the filename: <kind>-<slug>-<id>.png. | |
| tier | No | Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. "pro" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro. | |
| prompt | Yes | What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it. | |
| confirmPro | No | Required true with tier: "pro". Offer pro to the owner as an option for portraits and illustrations ("pro is available for a bit extra"); never assume it. | |
| references | No | Reference images, attached in this order. Bind each in the prompt by its label or 1-based index ("Image 2 is Morgash"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden and does so thoroughly: it discloses default and pro model tiers with cost ranges, that pro refuses unless confirmPro is true, that framing/plate/cut post-processing is applied automatically, and the return payload (path, dimensions, estimated spend). It even includes a quality-control caveat about inspecting for artifacts before sharing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact for a 6-parameter tool with no output schema and front-loads purpose plus the most consequential behavior: model tier, confirmPro refusal, and cost. Every sentence earns its place, including the post-generation checklist and return-value note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, the description explicitly states what the tool returns: file path, dimensions, and estimated spend. It covers cost behavior, confirmPro requirements, cross-tool routing to edit-image, and automatic post-processing, leaving the agent with the information needed to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents kind, slug, tier, prompt, confirmPro, and references in detail. The description adds a useful overview of what kind controls and the default-vs-pro behavior, but it does not materially exceed the parameter-level explanations already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: Generate a Foundry art asset from a prompt via the Gemini image API, and states the deliverable (finished PNG on disk). It also separates generation from the edit-image sibling by explicitly routing post-generation flaws to edit-image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use flash vs. pro tiers, when confirmPro is mandatory, and points flawed outputs to edit-image. It does not explicitly contrast generate-image with artificer-status or cutout-image, but the generation-vs-post-processing distinction is largely evident from the sibling names and the finished-PNG framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- Added
cutout-image - Added
edit-image - Changed
generate-image12 fields changed- removed
Input schema / properties / batchRemoved value: -{ - "default": 6, - "description": "Draft mode only: images per batch.", - "maximum": 8, - "minimum": 1, - "type": "integer" -} - added
Input schema / properties / confirmProAdded value: +{ + "description": "Required true with tier: \"pro\". Offer pro to the owner as an option for portraits and illustrations (\"pro is available for a bit extra\"); never assume it.", + "type": "boolean" +} - removed
Input schema / properties / denoiseRemoved value: -{ - "default": 0.7, - "description": "Refine mode only. 0.7 (pinned by test) keeps the scene skeleton in dev style; ~0.55 clones composition but inherits the draft rendering style.", - "maximum": 0.95, - "minimum": 0.3, - "type": "number" -} - changed
Input schema / properties / kind / descriptionPrevious value: -"Purpose preset — fixes generation and output resolution. No raw dimensions."New value: +"Purpose preset. icon: 1:1 flash → 512 square. token: 1:1 flash, top-down full body on a chroma plate, cut to alpha on a 512 square (framing, plate, and cut are done for you). portrait: 3:4 at 2K. illustration: 16:9 at 4K → 2560×1600 (16:10 crop). Every kind defaults to flash; tier: \"pro\" is opt-in and needs confirmPro." - changed
Input schema / properties / kind / enumPrevious value: -[ - "handout", - "scene-background", - "portrait", - "token" -]New value: +[ + "icon", + "token", + "portrait", + "illustration" +] - removed
Input schema / properties / modeRemoved value: -{ - "default": "draft", - "description": "draft: fast klein batch for curation. final: dev-quality render from the prompt alone, finished at output resolution. refine: dev img2img over sourceImage (a picked draft) — keeps its scene skeleton, re-renders in dev style, finished at output resolution.", - "enum": [ - "draft", - "final", - "refine" - ], - "type": "string" -} - changed
Input schema / properties / prompt / descriptionPrevious value: -"The full image prompt."New value: +"What a camera would see, in illustrator terms. Do not add framing or background text for icons and tokens; the preset appends it." - added
Input schema / properties / referencesAdded value: +{ + "description": "Reference images, attached in this order. Bind each in the prompt by its label or 1-based index (\"Image 2 is Morgash\").", + "items": { + "properties": { + "label": { + "description": "Short name used to bind the reference in the prompt, e.g. \"Morgash\".", + "type": "string" + }, + "path": { + "description": "Absolute path of a PNG/JPEG on disk.", + "minLength": 1, + "type": "string" + }, + "role": { + "description": "character: hold this face/figure (up to 4 on flash, 5 on pro). style: match palette, brushwork, light, camera angle; never copy the subject (works on both tiers in practice).", + "enum": [ + "character", + "style" + ], + "type": "string" + } + }, + "required": [ + "path", + "role" + ], + "type": "object" + }, + "maxItems": 14, + "type": "array" +} - removed
Input schema / properties / seedRemoved value: -{ - "description": "Fixed seed; random when omitted.", - "minimum": 0, - "type": "integer" -} - changed
Input schema / properties / slug / descriptionPrevious value: -"Short kebab-case subject name used in output filenames, e.g. \"smugglers-cove\"."New value: +"Kebab-cased into the filename: <kind>-<slug>-<id>.png." - removed
Input schema / properties / sourceImageRemoved value: -{ - "description": "Refine mode only (required there): absolute path of the picked draft PNG.", - "type": "string" -} - added
Input schema / properties / tierAdded value: +{ + "description": "Default flash (Nano Banana 2, ~7-15¢), never needs a confirm. \"pro\" (Nano Banana Pro, ~13-24¢, style-reference slots, stronger multi-figure scenes) needs confirmPro.", + "enum": [ + "flash", + "pro" + ], + "type": "string" +}
- Removed
upscale-image
3 tool updates
v0.1.0- First observed
artificer-status - First observed
generate-image - First observed
upscale-image
TDQS
Scored across 4 tools
Each tool targets a distinct action: generate, edit, cutout, and status. There is no meaningful overlap, and an agent can confidently select the right tool for creating new art, modifying existing art, preparing tokens, or checking system health.
Three tools follow a clear verb-noun pattern: generate-image, edit-image, cutout-image. artificer-status is a noun-noun outlier, but it still fits the server's naming style and is not confusing.
Four tools is well-scoped for a focused Foundry VTT art asset pipeline. Each tool represents a necessary step—create, edit, cut out, and check status—without unnecessary bloat.
The toolset covers the core lifecycle of generating, editing, and preparing token art, plus a health/status check. Minor gaps exist around listing or deleting assets, but for the stated purpose it is functionally complete and has no dead ends.
Maintenance
Related MCP Connectors
Connect any AI to your Foundry VTT world: actors, combat, dice, journals, tokens, compendiums.
Generate logos, social posts, app screenshots, comic panels & visual-novel assets from prompts.
Remote MCP for RunComfy: ComfyUI deployments, hosted models, LoRA training. 31 tools.
Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceConnects Claude Desktop to Foundry VTT for AI-powered campaign management, enabling natural language interaction with game data including quest creation, character management, compendium searches, and dice rolling. Provides 20 MCP tools for seamless integration between Claude and your tabletop RPG sessions.70-
- FlicenseNot gradedqualityNot gradedmaintenanceIntegrates with FoundryVTT tabletop gaming sessions, allowing AI assistants to query game data, roll dice, generate content (NPCs, loot, encounters), manage combat, and provide tactical suggestions through natural language.1 npm-
- FlicenseNot gradedqualityCmaintenanceEnables AI-powered campaign management for Foundry Virtual Tabletop through natural language, supporting multiple RPG systems with tools for quest creation, character management, combat resolution, and more.-
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to interact with Foundry Virtual Tabletop, supporting reading world data, managing combat, rolling dice, and updating actor attributes via a sidecar architecture.42 npmMIT