Skip to main content
Glama

forge-mcp

An MCP server that exposes a local Forge / AUTOMATIC1111 SDXL backend as image-generation tools — so any MCP client (Claude Code, Claude Desktop, Open WebUI, …) can generate and edit images through a clean, typed interface instead of hand-rolled API calls.

It is built to be frictionless, deterministic, and a good citizen on a shared GPU — the common case where the same box also runs a local LLM or other workloads.

Features

  • generate_image (txt2img) and edit_image (img2img) with sane, overridable defaults.

  • Deterministic by default — the resolved seed and every parameter are returned with each result, so a keeper is its own recipe. Prompts are sent verbatim; the server never silently rewrites them.

  • Good shared-GPU citizen — all GPU work is serialized; the checkpoint is switched only when it differs from the one already loaded (the reload is the VRAM spike that OOMs); generations are clamped to VRAM-safe aspect buckets; out-of-memory is surfaced as a clean, typed error.

  • Stateless — it owns no database and no stored state; the calling client/repo stays the source of truth for any character anchors or profiles.

  • SFW by default — a server-side NSFW block is applied unconditionally (it does not trust the prompt, the profile, or the calling model).

  • Returns a viewable preview inline plus a short-lived link to the full-resolution PNG (a URL, not base64 — base64 would flood the model's context), so the caller downloads the keeper out-of-band. This is topology-independent: the server and the saving machine can be different hosts.

  • Init images go in the same way, by reference — edit_image takes an init_image_ref: upload the keeper out-of-band (curl -H "Authorization: Bearer $FORGE_MCP_TOKEN" --data-binary @keeper.png <server>/upload → {"ref": ...}), or reuse a recent full_res_url. Raw base64 init_image still works for small images.

  • list_models(refresh=True) rescans Forge's checkpoint + LoRA folders (new files show up without a restart) and lists LoRA names for <lora:NAME:weight> prompts.

  • Named negative/style profiles and shot shortcuts (portrait, establishing, square) over raw width/height.

Related MCP server: SD + TTS MCP Server

Requirements

  • A running Forge or AUTOMATIC1111 instance with the API enabled (--api, and --listen if remote).

  • Python 3.12+ and uv.

Install & run

uv sync

# stdio (for a local MCP client that launches the server):
FORGE_URL=http://127.0.0.1:7860 uv run forge-mcp serve

# streamable-http (for a networked client / gateway), with a bearer token:
FORGE_URL=http://127.0.0.1:7860 FORGE_MCP_TOKEN=$(openssl rand -hex 32) \
  uv run forge-mcp serve --transport http --host 0.0.0.0 --port 8000

Configuration (environment variables)

Var

Default

Purpose

FORGE_URL

http://127.0.0.1:7860

Base URL of the Forge/A1111 --api server

FORGE_MCP_TOKEN

(unset)

If set, require Authorization: Bearer <token> on the HTTP transport

FORGE_DEFAULT_MODEL

DreamShaperXL_Turbo_SFW

Checkpoint used when a request doesn't name one

FORGE_GEN_TIMEOUT

180

Seconds to wait on Forge (covers a cold model-load + generation)

FORGE_PREVIEW_MAX_PX

768

Longest edge of the inline JPEG preview

FORGE_PUBLIC_URL

(unset)

Base URL clients use to reach this server (e.g. http://host:8000); required so include_full=True returns a downloadable full_res_url. Must be reachable from the saving machine.

FORGE_IMG_TTL

600

Seconds a served full-res link (or uploaded init-image ref) stays valid (TTL GC, not delete-on-first-GET)

FORGE_UPLOAD_MAX_BYTES

20971520

Size cap for POST /upload (init images for edit_image)

TLS note: behind a TLS-inspecting proxy, uv may report UnknownIssuer; pass --system-certs (or set UV_NATIVE_TLS=1) so it trusts the OS certificate store. The server itself uses truststore at runtime for the same reason.

Docker

docker build -t forge-mcp:latest .
docker run --rm -p 8000:8000 -e FORGE_URL=http://host.docker.internal:7860 \
  -e FORGE_MCP_TOKEN=... forge-mcp:latest

The server holds no GPU and no model — it brokers to your Forge instance and enforces the VRAM-safe envelope and the NSFW block.

Design principles

  • Deterministic over magic — reproducibility (seed + params returned, verbatim prompts) is a feature, not a nicety. No silent "enhance."

  • The caller owns state — no character database inside the server; it stays a stateless proxy.

  • A good tenant — never thrash a GPU that something else is using.

  • SFW is enforced server-side — because an MCP can't assume a well-behaved caller.

See CLAUDE.md for the fuller design rationale and the roadmap (upscale, async batch variations, and a deferred inpaint).

Style

Python 3.12, uv, standard library + httpx / mcp / Pillow / truststore — few dependencies by design. Keep it stateless and deterministic; do not add anything that rewrites prompts silently or relaxes the NSFW floor via a tool flag.

AI assistance

forge-mcp is developed openly with the help of Claude (Anthropic). We state this plainly: commits Claude helped write carry a Co-Authored-By: Claude trailer. The code and design are open source so the work can be inspected, reused, and given back.

License

MPL-2.0 © Daniel Núñez.

Available Tools

3 tools
edit_imageA

Edit/vary an image (img2img). ⚠ HOLISTIC — re-derives the WHOLE frame from the init, so use it for variants that should re-settle (wardrobe, pose, lighting, hair-mass), NOT single-feature fixes (those belong in a PSD — pushing just the eyes also moves skin/age/expression). For character consistency, pass the character's canonical keeper as init_image and its seed/prompt, then change only the delta.

    Args:
        init_image: the source image as base64 (your client reads the keeper PNG from the repo and passes
            it; also load its sidecar params so this render builds on the keeper's recipe).
        denoising_strength: 0.3-0.45 = keep the subject, fix details; 0.5-0.65 = restyle. Default 0.45.
        (all other args: see generate_image.)
    
ParametersJSON Schema
NameRequiredDescriptionDefault
cfgNo
seedNo
shotNoportrait
modelNo
stepsNo
styleNo
widthNo
heightNo
promptYes
samplerNoDPM++ SDE
negativeNo
schedulerNoKarras
init_imageYes
include_fullNo
negative_profileNosfw-strict
denoising_strengthNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=false, so the description carries most of the behavioral burden and does it well: it warns that the render is HOLISTIC and will also move skin/age/expression, maps denoising_strength bands to outcomes, and instructs the client to load the keeper's sidecar params. It omits cost/latency, determinism of seed, and return-format behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the critical ⚠ HOLISTIC caveat before the Args block, and nearly every clause carries actionable information. The Google-style Args formatting is slightly awkward for a tool description and the deferred-args sentence is a bit hand-wavy, but there is little filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter tool with no output schema and thin annotations, the description covers the distinctive behavior and the two key parameters but leaves the bulk of the argument surface to be resolved via generate_image. An agent can call it correctly for the common case but must cross-reference another tool for anything beyond init_image/denoising_strength.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 16 parameters at 0% schema description coverage, the description documents only init_image and denoising_strength (with concrete 0.3-0.45 / 0.5-0.65 bands and the 0.45 default). The remaining 14 parameters, including the required prompt, are deferred wholesale to generate_image, which is a useful pointer but leaves most of the surface undocumented in place.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Edit/vary an image (img2img)") and scopes it precisely: whole-frame re-derivation, not single-feature fixes. It implicitly separates itself from generate_image by noting that tool covers "all other args" and by requiring an init_image, so an agent can route correctly between the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use and when-not-to-use: use it for variants that should re-settle (wardrobe, pose, lighting, hair-mass), and NOT for single-feature fixes, which "belong in a PSD." It also gives a concrete workflow for character consistency (pass the canonical keeper as init_image plus its seed/prompt, change only the delta).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generate an image (txt2img) on the local SDXL box. Returns an inline preview + (in the text block) the exact params and the full-res PNG base64 to save in your repo.

    Args:
        prompt: the positive prompt, sent VERBATIM (attention weights like `(grey eyes:1.3)` work).
        shot: aspect/framing shortcut — `portrait` (832x1216, default, face-focus), `full-figure`,
            `establishing` (1216x832 scene), `square` (1024). Raw width/height override this.
        model: checkpoint (default DreamShaperXL_Turbo_SFW). Switched only if different from loaded.
        seed: -1 = random (the resolved seed is returned so you can lock it); set it to reproduce.
        steps: default 8 (Turbo). cfg: default 2 (Turbo). sampler/scheduler: DPM++ SDE / Karras (Turbo).
        negative: extra negative text (combined with the profile; the NSFW block is always appended).
        negative_profile: named profile(s), e.g. "sfw-strict" (default), "mature", "candid+sfw-strict".
        style: optional positive preset, e.g. "photoreal".
        width/height: raw size, must be a VRAM-safe bucket (832x1216, 1216x832, 1024x1024).
        include_full: include the full-res PNG base64 in the result (default True; False = preview only,
            lighter for rapid iteration — re-run with the echoed seed to get the full-res of a keeper).
    
ParametersJSON Schema
NameRequiredDescriptionDefault
cfgNo
seedNo
shotNoportrait
modelNo
stepsNo
styleNo
widthNo
heightNo
promptYes
samplerNoDPM++ SDE
negativeNo
schedulerNoKarras
include_fullNo
negative_profileNosfw-strict

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing return format (inline preview + params + full-res base64 PNG), model-switch behavior ('Switched only if different from loaded'), that the NSFW block is always appended, the VRAM-safe bucket constraint, and that the resolved seed is echoed back. These are genuinely useful behavioral facts an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and return info, then a well-structured Args block. It is long, but the length is justified by 14 parameters; a few entries (cfg/steps/sampler/scheduler) are thinner than others, keeping it just under a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter local-generation tool with only minimal annotations and no output schema, the description covers purpose, all key parameter semantics, return values, and behavioral constraints. Nothing an agent needs to call it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage and 14 parameters, the description carries the full burden and documents essentially all of them: prompt verbatim semantics with attention-weight syntax, shot enum-like values with dimensions, seed=-1 meaning and return, profile names, style presets, and the allowed width/height buckets. This compensates thoroughly for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Generate an image (txt2img) on the local SDXL box'), which clearly separates it from the edit_image and list_models siblings. It stops short of explicitly naming and contrasting those siblings, so it lands just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context for parameters — include_full=False for rapid iteration, seed=-1 for random vs. locking the echoed seed to reproduce, and the shot shortcuts. However, it never explicitly states when to use this tool versus the alternatives (edit_image, list_models), so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA
Read-only

List the image models (checkpoints) available on the Forge box, and which one is loaded.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful context beyond that: it discloses what comes back (available checkpoints) and the loaded-state information, which matters with no output schema present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the resource and scope come first and the extra return detail follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with no output schema, the description adequately covers what is returned and the state it reflects. It could add a note on cost/caching or output shape, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No parameter detail is required here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the image models (checkpoints)') plus the environment ('on the Forge box') and an extra facet ('which one is loaded'). It is unambiguous against the generate_image/edit_image siblings, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent would infer this is a discovery step before generating or editing images, but the description never says when to call it, nor names any alternative or prerequisite.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observededit_image
    • First observedgenerate_image
    • First observedlist_models

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

generate_image (txt2img), edit_image (img2img), and list_models each have a clearly distinct purpose, and the descriptions explicitly explain the txt2img vs img2img boundary (including when to use PSD instead). No meaningful overlap remains.

Naming Consistency5/5

All three tools follow a clean verb_noun snake_case pattern: generate_image, edit_image, list_models. Fully predictable and consistent.

Tool Count4/5

Three tools is a coherent, well-scoped set for a focused SDXL render server, but it sits at the thin end—common needs like upscaling or inpainting have no dedicated tool.

Completeness4/5

Core lifecycle (text-to-image, image-to-image, model discovery) is covered and the model arg handles checkpoint switching. Minor gaps like upscaling/detail-fix or masked inpainting are acknowledged as out of scope but would otherwise cause dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers