forge-mcp
Summary: forge-mcp lets an MCP client generate, edit, and inspect images through a local Forge/AUTOMATIC1111 SDXL backend.
Generate txt2img images via
generate_imagewith prompt, shot/aspect presets, model, seed, steps, CFG, sampler/scheduler, negative prompts/profiles, style, and raw width/height.Edit or vary images via
edit_imageusing a base64 init image, denoising strength, and the same generation parameters; intended for holistic variants, not single-feature fixes.List available checkpoints and the currently loaded model via
list_models; README also says refresh rescans checkpoints/LoRAs without restart, though the shown schema has norefreshparameter.Get deterministic results: resolved seed and parameters are returned, prompts are sent verbatim, and SFW blocking is enforced server-side.
Work safely on a shared GPU: GPU work is serialized, checkpoint reloads are avoided unless needed, and VRAM-safe aspect buckets are enforced.
Receive inline previews and, per README, short-lived full-resolution PNG links; schema descriptions mention full-res base64 when
include_full=True.Use out-of-band upload/init-image refs per README for edit sources, though the shown schema only exposes base64
init_image.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@forge-mcpgenerate a portrait of a cyberpunk samurai in the rain"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
forge-mcp
An MCP server that exposes a local Forge / AUTOMATIC1111 SDXL backend as image-generation tools — so any MCP client (Claude Code, Claude Desktop, Open WebUI, …) can generate and edit images through a clean, typed interface instead of hand-rolled API calls.
It is built to be frictionless, deterministic, and a good citizen on a shared GPU — the common case where the same box also runs a local LLM or other workloads.
Features
generate_image(txt2img) andedit_image(img2img) with sane, overridable defaults.Deterministic by default — the resolved seed and every parameter are returned with each result, so a keeper is its own recipe. Prompts are sent verbatim; the server never silently rewrites them.
Good shared-GPU citizen — all GPU work is serialized; the checkpoint is switched only when it differs from the one already loaded (the reload is the VRAM spike that OOMs); generations are clamped to VRAM-safe aspect buckets; out-of-memory is surfaced as a clean, typed error.
Stateless — it owns no database and no stored state; the calling client/repo stays the source of truth for any character anchors or profiles.
SFW by default — a server-side NSFW block is applied unconditionally (it does not trust the prompt, the profile, or the calling model).
Returns a viewable preview inline plus a short-lived link to the full-resolution PNG (a URL, not base64 — base64 would flood the model's context), so the caller downloads the keeper out-of-band. This is topology-independent: the server and the saving machine can be different hosts.
Init images go in the same way, by reference —
edit_imagetakes aninit_image_ref: upload the keeper out-of-band (curl -H "Authorization: Bearer $FORGE_MCP_TOKEN" --data-binary @keeper.png <server>/upload→{"ref": ...}), or reuse a recentfull_res_url. Raw base64init_imagestill works for small images.list_models(refresh=True)rescans Forge's checkpoint + LoRA folders (new files show up without a restart) and lists LoRA names for<lora:NAME:weight>prompts.Named negative/style profiles and shot shortcuts (
portrait,establishing,square) over raw width/height.
Related MCP server: SD + TTS MCP Server
Requirements
A running Forge or AUTOMATIC1111 instance with the API enabled (
--api, and--listenif remote).Python 3.12+ and uv.
Install & run
uv sync
# stdio (for a local MCP client that launches the server):
FORGE_URL=http://127.0.0.1:7860 uv run forge-mcp serve
# streamable-http (for a networked client / gateway), with a bearer token:
FORGE_URL=http://127.0.0.1:7860 FORGE_MCP_TOKEN=$(openssl rand -hex 32) \
uv run forge-mcp serve --transport http --host 0.0.0.0 --port 8000Configuration (environment variables)
Var | Default | Purpose |
|
| Base URL of the Forge/A1111 |
| (unset) | If set, require |
|
| Checkpoint used when a request doesn't name one |
|
| Seconds to wait on Forge (covers a cold model-load + generation) |
|
| Longest edge of the inline JPEG preview |
| (unset) | Base URL clients use to reach this server (e.g. |
|
| Seconds a served full-res link (or uploaded init-image ref) stays valid (TTL GC, not delete-on-first-GET) |
|
| Size cap for |
TLS note: behind a TLS-inspecting proxy,
uvmay reportUnknownIssuer; pass--system-certs(or setUV_NATIVE_TLS=1) so it trusts the OS certificate store. The server itself usestruststoreat runtime for the same reason.
Docker
docker build -t forge-mcp:latest .
docker run --rm -p 8000:8000 -e FORGE_URL=http://host.docker.internal:7860 \
-e FORGE_MCP_TOKEN=... forge-mcp:latestThe server holds no GPU and no model — it brokers to your Forge instance and enforces the VRAM-safe envelope and the NSFW block.
Design principles
Deterministic over magic — reproducibility (seed + params returned, verbatim prompts) is a feature, not a nicety. No silent "enhance."
The caller owns state — no character database inside the server; it stays a stateless proxy.
A good tenant — never thrash a GPU that something else is using.
SFW is enforced server-side — because an MCP can't assume a well-behaved caller.
See CLAUDE.md for the fuller design rationale and the roadmap (upscale, async batch
variations, and a deferred inpaint).
Style
Python 3.12, uv, standard library + httpx / mcp / Pillow / truststore — few dependencies by design.
Keep it stateless and deterministic; do not add anything that rewrites prompts silently or relaxes the NSFW
floor via a tool flag.
AI assistance
forge-mcp is developed openly with the help of Claude (Anthropic). We state this plainly: commits
Claude helped write carry a Co-Authored-By: Claude trailer. The code and design are open source so the
work can be inspected, reused, and given back.
License
MPL-2.0 © Daniel Núñez.
Available Tools
3 toolsedit_imageA
Edit/vary an image (img2img). ⚠ HOLISTIC — re-derives the WHOLE frame from the init, so use it for
variants that should re-settle (wardrobe, pose, lighting, hair-mass), NOT single-feature fixes (those
belong in a PSD — pushing just the eyes also moves skin/age/expression). For character consistency,
pass the character's canonical keeper as init_image and its seed/prompt, then change only the delta.
Args:
init_image: the source image as base64 (your client reads the keeper PNG from the repo and passes
it; also load its sidecar params so this render builds on the keeper's recipe).
denoising_strength: 0.3-0.45 = keep the subject, fix details; 0.5-0.65 = restyle. Default 0.45.
(all other args: see generate_image.)
| Name | Required | Description | Default |
|---|---|---|---|
| cfg | No | ||
| seed | No | ||
| shot | No | portrait | |
| model | No | ||
| steps | No | ||
| style | No | ||
| width | No | ||
| height | No | ||
| prompt | Yes | ||
| sampler | No | DPM++ SDE | |
| negative | No | ||
| scheduler | No | Karras | |
| init_image | Yes | ||
| include_full | No | ||
| negative_profile | No | sfw-strict | |
| denoising_strength | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=false, so the description carries most of the behavioral burden and does it well: it warns that the render is HOLISTIC and will also move skin/age/expression, maps denoising_strength bands to outcomes, and instructs the client to load the keeper's sidecar params. It omits cost/latency, determinism of seed, and return-format behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the critical ⚠ HOLISTIC caveat before the Args block, and nearly every clause carries actionable information. The Google-style Args formatting is slightly awkward for a tool description and the deferred-args sentence is a bit hand-wavy, but there is little filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool with no output schema and thin annotations, the description covers the distinctive behavior and the two key parameters but leaves the bulk of the argument surface to be resolved via generate_image. An agent can call it correctly for the common case but must cross-reference another tool for anything beyond init_image/denoising_strength.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 16 parameters at 0% schema description coverage, the description documents only init_image and denoising_strength (with concrete 0.3-0.45 / 0.5-0.65 bands and the 0.45 default). The remaining 14 parameters, including the required prompt, are deferred wholesale to generate_image, which is a useful pointer but leaves most of the surface undocumented in place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Edit/vary an image (img2img)") and scopes it precisely: whole-frame re-derivation, not single-feature fixes. It implicitly separates itself from generate_image by noting that tool covers "all other args" and by requiring an init_image, so an agent can route correctly between the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and when-not-to-use: use it for variants that should re-settle (wardrobe, pose, lighting, hair-mass), and NOT for single-feature fixes, which "belong in a PSD." It also gives a concrete workflow for character consistency (pass the canonical keeper as init_image plus its seed/prompt, change only the delta).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate an image (txt2img) on the local SDXL box. Returns an inline preview + (in the text block) the exact params and the full-res PNG base64 to save in your repo.
Args:
prompt: the positive prompt, sent VERBATIM (attention weights like `(grey eyes:1.3)` work).
shot: aspect/framing shortcut — `portrait` (832x1216, default, face-focus), `full-figure`,
`establishing` (1216x832 scene), `square` (1024). Raw width/height override this.
model: checkpoint (default DreamShaperXL_Turbo_SFW). Switched only if different from loaded.
seed: -1 = random (the resolved seed is returned so you can lock it); set it to reproduce.
steps: default 8 (Turbo). cfg: default 2 (Turbo). sampler/scheduler: DPM++ SDE / Karras (Turbo).
negative: extra negative text (combined with the profile; the NSFW block is always appended).
negative_profile: named profile(s), e.g. "sfw-strict" (default), "mature", "candid+sfw-strict".
style: optional positive preset, e.g. "photoreal".
width/height: raw size, must be a VRAM-safe bucket (832x1216, 1216x832, 1024x1024).
include_full: include the full-res PNG base64 in the result (default True; False = preview only,
lighter for rapid iteration — re-run with the echoed seed to get the full-res of a keeper).
| Name | Required | Description | Default |
|---|---|---|---|
| cfg | No | ||
| seed | No | ||
| shot | No | portrait | |
| model | No | ||
| steps | No | ||
| style | No | ||
| width | No | ||
| height | No | ||
| prompt | Yes | ||
| sampler | No | DPM++ SDE | |
| negative | No | ||
| scheduler | No | Karras | |
| include_full | No | ||
| negative_profile | No | sfw-strict |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing return format (inline preview + params + full-res base64 PNG), model-switch behavior ('Switched only if different from loaded'), that the NSFW block is always appended, the VRAM-safe bucket constraint, and that the resolved seed is echoed back. These are genuinely useful behavioral facts an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and return info, then a well-structured Args block. It is long, but the length is justified by 14 parameters; a few entries (cfg/steps/sampler/scheduler) are thinner than others, keeping it just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter local-generation tool with only minimal annotations and no output schema, the description covers purpose, all key parameter semantics, return values, and behavioral constraints. Nothing an agent needs to call it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and 14 parameters, the description carries the full burden and documents essentially all of them: prompt verbatim semantics with attention-weight syntax, shot enum-like values with dimensions, seed=-1 meaning and return, profile names, style presets, and the allowed width/height buckets. This compensates thoroughly for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Generate an image (txt2img) on the local SDXL box'), which clearly separates it from the edit_image and list_models siblings. It stops short of explicitly naming and contrasting those siblings, so it lands just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context for parameters — include_full=False for rapid iteration, seed=-1 for random vs. locking the echoed seed to reproduce, and the shot shortcuts. However, it never explicitly states when to use this tool versus the alternatives (edit_image, list_models), so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsARead-only
List the image models (checkpoints) available on the Forge box, and which one is loaded.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful context beyond that: it discloses what comes back (available checkpoints) and the loaded-state information, which matters with no output schema present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource and scope come first and the extra return detail follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with no output schema, the description adequately covers what is returned and the state it reflects. It could add a note on cost/caching or output shape, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No parameter detail is required here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the image models (checkpoints)') plus the environment ('on the Forge box') and an extra facet ('which one is loaded'). It is unambiguous against the generate_image/edit_image siblings, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent would infer this is a discovery step before generating or editing images, but the description never says when to call it, nor names any alternative or prerequisite.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
edit_image - First observed
generate_image - First observed
list_models
TDQS
Scored across 3 tools
generate_image (txt2img), edit_image (img2img), and list_models each have a clearly distinct purpose, and the descriptions explicitly explain the txt2img vs img2img boundary (including when to use PSD instead). No meaningful overlap remains.
All three tools follow a clean verb_noun snake_case pattern: generate_image, edit_image, list_models. Fully predictable and consistent.
Three tools is a coherent, well-scoped set for a focused SDXL render server, but it sits at the thin end—common needs like upscaling or inpainting have no dedicated tool.
Core lifecycle (text-to-image, image-to-image, model discovery) is covered and the model arg handles checkpoint switching. Minor gaps like upscaling/detail-fix or masked inpainting are acknowledged as out of scope but would otherwise cause dead ends.
Maintenance
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
Remote MCP for RunComfy: ComfyUI deployments, hosted models, LoRA training. 31 tools.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceExposes multiple image generation backends as independent MCP tools for generating images with configurable models and parameters.MIT
- FlicenseNot gradedqualityDmaintenanceExposes Stable Diffusion for text-to-image generation and GPT-SoVITS for text-to-speech synthesis as MCP tools, enabling image and audio generation via natural language.-
- AlicenseNot gradedqualityBmaintenanceMCP server that wraps ComfyUI for SDXL image generation. It exposes tools for generating images, listing models, and checking ComfyUI health, with presets and GPU resource coordination.MIT
- AlicenseNot gradedqualityBmaintenanceTurns curated ComfyUI workflows into typed MCP tools and REST endpoints, enabling image generation and job management through natural language or HTTP calls.1MIT