MCP ComfyUI Flux
Enables AI image generation using FLUX models (schnell and dev) via ComfyUI, with support for fp8 quantization, batch processing, 4x upscaling, and background removal.
Downloads FLUX models and related AI models from Hugging Face repositories for image generation workflows.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP ComfyUI Fluxgenerate a cyberpunk cityscape with neon lights"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP ComfyUI Flux - Optimized Docker Solution
A fully containerized MCP (Model Context Protocol) server for generating images with FLUX models via ComfyUI. Features optimized Docker builds, PyTorch 2.5.1, automatic GPU acceleration, and Claude Desktop integration.
π Features
π Optimized Performance: PyTorch 2.5.1 with native RMSNorm support
π¦ Efficient Images: 25% smaller Docker images (10.9GB vs 14.6GB)
β‘ Fast Rebuilds: BuildKit cache mounts for rapid iterations
π¨ FLUX Models: Supports schnell (4-step) and dev models with fp8 quantization
π€ MCP Integration: Works seamlessly with Claude Desktop
πͺ GPU Acceleration: Automatic NVIDIA GPU detection and CUDA 12.1
π Background Removal: Built-in RMBG-2.0 for transparent backgrounds
π Image Upscaling: 4x upscaling with UltraSharp/AnimeSharp models
π‘οΈ Production Ready: Health checks, auto-recovery, extensive logging
Related MCP server: ComfyUI MCP Server
π Table of Contents
π Quick Start
# Clone the repository
git clone <repository-url> mcp-comfyui-flux
cd mcp-comfyui-flux
# Run the automated installer
./install.sh
# Or build manually with the optimized build script
./build.sh --start
# That's it! The installer will:
# - Check prerequisites
# - Configure environment
# - Download FLUX models
# - Build optimized Docker containers
# - Start all servicesπ» System Requirements
Minimum Requirements
OS: Linux, macOS, Windows 10+ (WSL2)
CPU: 4 cores
RAM: 16GB (20GB for WSL2)
Storage: 50GB free space
Docker: 20.10+
Docker Compose: 2.0+ or 1.29+ (legacy)
Recommended Requirements
CPU: 8+ cores
RAM: 32GB
GPU: NVIDIA RTX 3090/4090 (12GB+ VRAM)
Storage: 100GB free space
CUDA: 12.1+ with NVIDIA Container Toolkit
WSL2 Specific (Windows)
# .wslconfig in Windows user directory
[wsl2]
memory=20GB
processors=8
localhostForwarding=trueπ¦ Installation
Prerequisites
Install Docker:
# Ubuntu/Debian curl -fsSL https://get.docker.com | bash # macOS brew install docker docker-compose # Windows - Install Docker DesktopInstall NVIDIA Container Toolkit (for GPU):
# Ubuntu/Debian distribution=$(. /etc/os-release;echo $ID$VERSION_ID) curl -s -L https://nvidia.github.io/nvidia-docker/gpgkey | sudo apt-key add - curl -s -L https://nvidia.github.io/nvidia-docker/$distribution/nvidia-docker.list | \ sudo tee /etc/apt/sources.list.d/nvidia-docker.list sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit sudo systemctl restart docker
Automated Installation
# Standard installation
./install.sh
# Non-interactive installation
./install.sh --yes
# CPU-only mode
./install.sh --cpu-only
# With specific models
./install.sh --models minimal # or all/none/auto
# Debug mode
./install.sh --debugBuild Script Options
# Build only
./build.sh
# Build and start
./build.sh --start
# Build with cleanup
./build.sh --start --cleanup
# Rebuild without cache
./build.sh --no-cacheπ¨ MCP Tools
Available Tools in Claude Desktop
1. generate_image
Generate images using FLUX schnell fp8 model (optimized defaults).
// Parameters
{
"prompt": "a majestic mountain landscape, golden hour", // Required
"negative_prompt": "blurry, low quality", // Optional
"width": 1024, // Default: 1024
"height": 1024, // Default: 1024
"steps": 4, // Default: 4 (schnell optimized)
"cfg_scale": 1.0, // Default: 1.0 (schnell optimized)
"seed": -1, // Default: -1 (random)
"batch_size": 1 // Default: 1 (max: 8)
}
// Example usage
generate_image({
prompt: "cyberpunk city at night, neon lights, detailed",
steps: 4,
seed: 42
})2. upscale_image
Upscale images to 4x resolution using AI models.
// Parameters
{
"image_path": "flux_output_00001_.png", // Required
"model": "ultrasharp", // Options: "ultrasharp", "animesharp"
"scale_factor": 1.0, // Additional scaling (0.5-2.0)
"content_type": "general" // Auto-select model based on content
}
// Example usage
upscale_image({
image_path: "output/my_image.png",
model: "ultrasharp"
})3. remove_background
Remove background using RMBG-2.0 AI model.
// Parameters
{
"image_path": "output/image.png", // Required
"alpha_matting": true, // Better edge quality (default: true)
"output_format": "png" // Options: "png", "webp"
}
// Example usage
remove_background({
image_path: "flux_output_00001_.png"
})4. check_models
Verify available models in ComfyUI.
// No parameters required
check_models()5. connect_comfyui / disconnect_comfyui
Manage ComfyUI connection (usually auto-connects).
MCP Configuration
Add to Claude Desktop config (%APPDATA%\Claude\claude_desktop_config.json on Windows):
{
"mcpServers": {
"comfyui-flux": {
"command": "wsl.exe",
"args": [
"bash", "-c",
"cd /path/to/mcp-comfyui-flux && docker exec -i mcp-comfyui-flux-mcp-server-1 node /app/src/index.js"
]
}
}
}For macOS/Linux:
{
"mcpServers": {
"comfyui-flux": {
"command": "docker",
"args": [
"exec", "-i", "mcp-comfyui-flux-mcp-server-1",
"node", "/app/src/index.js"
]
}
}
}π³ Docker Management
Service Commands
# Start services
docker-compose -p mcp-comfyui-flux up -d
# Stop services
docker-compose -p mcp-comfyui-flux down
# View logs
docker-compose -p mcp-comfyui-flux logs -f
docker-compose -p mcp-comfyui-flux logs -f comfyui
# Check status
docker-compose -p mcp-comfyui-flux ps
# Restart services
docker-compose -p mcp-comfyui-flux restartContainer Access
# Access ComfyUI container
docker exec -it mcp-comfyui-flux-comfyui-1 bash
# Access MCP server
docker exec -it mcp-comfyui-flux-mcp-server-1 sh
# Check GPU status
docker exec mcp-comfyui-flux-comfyui-1 nvidia-smi
# Test PyTorch
docker exec mcp-comfyui-flux-comfyui-1 python3.11 -c "import torch; print(f'PyTorch {torch.__version__}')"Health Monitoring
# Full health check
./scripts/health-check.sh
# Check ComfyUI API
curl http://localhost:8188/system_stats
# Container health status
docker inspect mcp-comfyui-flux-comfyui-1 --format='{{.State.Health.Status}}'π Advanced Features
Performance Optimizations
The optimized build includes:
PyTorch 2.5.1: Latest stable with native RMSNorm support
BuildKit Cache Mounts: Reduces I/O operations in WSL2
FP8 Quantization: FLUX schnell fp8 uses ~10GB VRAM (vs 24GB fp16)
Multi-stage Builds: Separates build and runtime dependencies
Compiled Python: Pre-compiled bytecode for faster startup
FLUX Model Configurations
Schnell (Default - Fast)
Steps: 4 (optimized for schnell)
CFG Scale: 1.0 (works best with low guidance)
Scheduler: simple
Generation Time: ~2-4 seconds per image
VRAM Usage: ~10GB base + 1GB per batch
Dev (High Quality)
Steps: 20-50
CFG Scale: 7.0
Scheduler: normal/karras
Requires: Hugging Face authentication
VRAM Usage: ~12-16GB
Batch Generation
Generate multiple images efficiently:
generate_image({
prompt: "fantasy landscape",
batch_size: 4 // Generates 4 variations in parallel
})Custom Nodes
Included custom nodes:
ComfyUI-Manager: Node management and updates
ComfyUI-KJNodes: Advanced processing nodes
ComfyUI-RMBG: Background removal (31 nodes)
π§ Troubleshooting
Common Issues
GPU Not Detected
# Verify NVIDIA driver
nvidia-smi
# Check Docker GPU support
docker run --rm --gpus all nvidia/cuda:12.1.0-base-ubuntu22.04 nvidia-smi
# Ensure NVIDIA Container Toolkit is installed
sudo apt-get install -y nvidia-container-toolkit
sudo systemctl restart dockerOut of Memory
# Reduce batch size
batch_size: 1
# Use CPU mode (in .env)
CUDA_VISIBLE_DEVICES=-1
# Adjust PyTorch memory
PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:256WSL2 Specific Issues
# If Docker/WSL2 crashes with I/O errors
# Avoid recursive chown on large directories
# Use the optimized Dockerfile which handles this
# Increase WSL2 memory in .wslconfig
memory=20GB
# Reset WSL2 if needed
wsl --shutdownPort Conflicts
# Check what's using port 8188
lsof -i :8188 # macOS/Linux
netstat -ano | findstr :8188 # Windows
# Use different port
PORT=8189 docker-compose -p mcp-comfyui-flux up -dLog Locations
Installation:
install.logDocker builds:
docker-compose logsComfyUI: Inside container at
/app/ComfyUI/user/comfyui.logMCP Server:
docker logs mcp-comfyui-flux-mcp-server-1
ποΈ Architecture
System Overview
βββββββββββββββββββββββββββββββββββββββββββ
β Claude Desktop (MCP Client) β
ββββββββββββββ¬βββββββββββββββββββββββββββββ
β docker exec stdio
ββββββββββββββΌβββββββββββββββββββββββββββββ
β MCP Server Container β
β β’ Node.js 20 Alpine (581MB) β
β β’ MCP Protocol Implementation β
β β’ Auto-connects to ComfyUI β
ββββββββββββββ¬βββββββββββββββββββββββββββββ
β WebSocket (port 8188)
ββββββββββββββΌβββββββββββββββββββββββββββββ
β ComfyUI Container β
β β’ Ubuntu 22.04 + CUDA 12.1 β
β β’ Python 3.11 + PyTorch 2.5.1 β
β β’ FLUX schnell fp8 (4.5GB) β
β β’ Custom nodes (KJNodes, RMBG) β
β β’ Optimized image size: 10.9GB β
βββββββββββββββββββββββββββββββββββββββββββKey Improvements
Docker Optimization
Multi-stage builds reduce image size by 25%
BuildKit cache mounts speed up rebuilds
No Python venv (Docker IS the isolation)
Model Configuration
FLUX schnell fp8: 4.5GB (vs 11GB fp16)
T5-XXL fp8: 4.9GB text encoder
CLIP-L: 235MB text encoder
VAE: 320MB decoder
Performance
4-step generation in 2-4 seconds
Batch processing up to 8 images
Native RMSNorm in PyTorch 2.5.1
High VRAM mode for 24GB+ GPUs
Directory Structure
mcp-comfyui-flux/
βββ src/ # MCP server source
β βββ index.js # MCP protocol handler
β βββ comfyui-client.js # WebSocket client
β βββ workflows/ # ComfyUI workflows
βββ models/ # Model storage
β βββ unet/ # FLUX models (fp8)
β βββ clip/ # Text encoders
β βββ vae/ # VAE models
β βββ upscale_models/ # Upscaling models
βββ output/ # Generated images
βββ scripts/ # Utility scripts
βββ docker-compose.yml # Service orchestration
βββ Dockerfile.comfyui # Optimized ComfyUI
βββ Dockerfile.mcp # MCP server
βββ requirements.txt # Python dependencies
βββ build.sh # Build script
βββ install.sh # Automated installerπ Security
Local Execution: All processing happens locally
No External APIs: Except model downloads from Hugging Face
Container Isolation: Services run in isolated containers
Non-root Execution: Containers run as non-root user
Token Security: Stored in
.env(gitignored)
π Additional Documentation
CLAUDE.md - Claude Code development guide
ARCHITECTURE.md - Technical architecture details
API.md - Complete MCP API reference
TROUBLESHOOTING.md - Detailed troubleshooting
π€ Contributing
Contributions welcome! Please:
Fork the repository
Create a feature branch
Make your changes
Submit a pull request
π License
MIT License - see LICENSE file for details.
π Acknowledgments
ComfyUI - The workflow engine
Black Forest Labs - FLUX model creators
Anthropic - MCP protocol and Claude
NVIDIA - CUDA and GPU support
Made with β€οΈ for efficient AI image generation
Available Tools
6 toolscheck_modelsC
Check if Flux dev model is available in ComfyUI
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description lacks behavioral details beyond the basic function. It does not state whether the tool requires an active connection, what it returns (e.g., boolean, message), or any side effects. With no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It efficiently communicates the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimal and does not explain what the tool returns or how to interpret the result. Given no output schema, a bit more context about return values or examples would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so schema coverage is trivially 100%. The description does not add anything about parameters beyond the schema, but this is acceptable given no parameters exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'check' and the specific resource 'Flux dev model availability in ComfyUI'. It distinguishes this tool from siblings like 'generate_image' and 'connect_comfyui'. However, it could be slightly more specific about the output (e.g., boolean), so not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. While it is implied that one might check availability before generating, the description does not explicitly state usage context or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connect_comfyuiC
Connect to ComfyUI server
| Name | Required | Description | Default |
|---|---|---|---|
| server_address | No | ComfyUI server address (default: 127.0.0.1:8188) | 127.0.0.1:8188 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavior. It only states 'Connect' without explaining side effects, state changes, authentication, or failure modes. This leaves the agent uncertain about what the tool does beyond establishing a connection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words. It is appropriately short for a simple connection tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having only one optional parameter, the description lacks important context such as whether the connection is persistent, returns a handle, or requires disconnection. Without an output schema, more detail is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with the parameter already having a description. The tool description adds no additional meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Connect' and the resource 'ComfyUI server', indicating a specific action. However, it does not distinguish this tool from siblings like 'disconnect_comfyui' or 'check_models'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool, such as prerequisites or that it should be called before other tools. There is no mention of alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
disconnect_comfyuiB
Disconnect from ComfyUI server
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description does not disclose any behavioral traits such as side effects (e.g., losing connection state), error conditions, or safety implications. It only states the action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that efficiently conveys the tool's purpose with no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no annotations), the description is mostly complete. However, it could mention potential error states or prerequisites for a fully complete context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, and schema coverage is 100%. The description adds no additional context about parameters, but the lack of parameters makes this adequate. A score of 3 is baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Disconnect from ComfyUI server' clearly specifies the action (disconnect) and the resource (ComfyUI server), distinguishing it from sibling tools like connect_comfyui.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool or when not to. There is no mention of prerequisites (e.g., being connected) or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageC
Generate an image using Flux dev model in ComfyUI
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for reproducibility (-1 for random) | |
| steps | No | Number of sampling steps | |
| width | No | Width of the generated image | |
| height | No | Height of the generated image | |
| prompt | Yes | The text prompt to generate an image from | |
| cfg_scale | No | Classifier-free guidance scale | |
| scheduler | No | Scheduler to use | normal |
| batch_size | No | Number of images to generate in parallel (1-4 recommended) | |
| sampler_name | No | Sampling method to use | euler |
| negative_prompt | No | Negative prompt to avoid certain features |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only says 'generate an image' without mentioning computational cost, time, or error handling. Minimal behavioral info.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no wasted words. It efficiently conveys the core purpose, though it sacrifices detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters and no output schema or annotations, the description is insufficient. It does not explain return values, limitations, or how results are delivered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds no additional meaning beyond what the schema provides; it is a single generic sentence.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it generates an image using the Flux dev model in ComfyUI, distinguishing it from sibling tools like 'check_models' or 'upscale_image'. However, it lacks specificity about the model's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description only states the action without context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_backgroundC
Remove background from an image using AI-powered segmentation
| Name | Required | Description | Default |
|---|---|---|---|
| image_path | Yes | Path to the input image file | |
| alpha_matting | No | Use alpha matting for better edge quality (especially for hair/fur) | |
| output_format | No | Output image format | png |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description should disclose behavioral aspects such as whether the input file is modified, where the output is saved, and any side effects. It only states the method without clarifying the outcome or operation side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with clear action and method. No unnecessary words. Efficiently communicates the core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fails to cover output details (e.g., where the result is saved, return value) and any constraints (e.g., file size limits, supported input formats). With three parameters and no annotations, more context is needed for full understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with adequate descriptions for each parameter (image_path, alpha_matting, output_format). The tool description adds no extra meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair ('Remove background from an image') and mentions the AI method ('AI-powered segmentation'). It is clearly distinct from sibling tools like generate_image or upscale_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage context or alternatives are mentioned. The agent is not told when to prefer this tool over other image tools or any limitations (e.g., image size, supported formats).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upscale_imageC
Upscale an image using AI models
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Upscaling model to use | ultrasharp |
| image_path | Yes | Path to the image file to upscale | |
| content_type | No | Content type for auto model selection | general |
| scale_factor | No | Additional scaling factor (1.0 = model native, usually 4x) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is minimal and does not disclose any behavioral traits beyond the basic operation. There are no annotations, so the agent has no information about side effects (e.g., whether the original image is modified, what happens on failure). This is inadequate for a tool with 4 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at 5 words, with no wasted sentences. However, it may be too minimal; a bit more detail could improve utility without sacrificing conciseness. Structure is front-loaded but sparse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of 4 parameters (including enum and scale factor) and no output schema, the description is incomplete. It does not explain what the tool returns (e.g., path to upscaled image) or any important context like supported image formats or resolution limits. The agent lacks key information to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The description adds no extra parameter-specific information beyond the schema. Baseline score of 3 is appropriate; the description provides no additional value to parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: upscaling an image using AI models. However, it does not differentiate from sibling tools like generate_image or remove_background, which could be confused with upscaling. The purpose is clear but lacks distinctiveness.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like check_models or generate_image. There is no mention of prerequisites (e.g., file format, size limits) or when not to use it. The usage context is entirely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
check_models - First observed
connect_comfyui - First observed
disconnect_comfyui - First observed
generate_image - First observed
remove_background - First observed
upscale_image
TDQS
Scored across 6 tools
Each tool has a clearly distinct purpose: model checking, connection management, image generation, and two distinct post-processing operations. No overlap or ambiguity.
All tool names follow a consistent verb_noun pattern in snake_case, making them predictable and easy to understand.
With 6 tools covering connection, model check, generation, and post-processing, the number is well-scoped for the server's purpose.
The tool set covers core workflows (connect, check, generate) and adds useful post-processing. Minor gap: only checks for one specific model; a general model listing could be useful.
Maintenance
Related MCP Connectors
Remote MCP for RunComfy: ComfyUI deployments, hosted models, LoRA training. 31 tools.
AI image, video & music generation. Flux, Veo 3.1, Suno V5. Free tier included.
Best Image and video generation: 20+ models (Kling, Seedance, Veo, NB, FLUX.2), OAuth, pay-per-use.
Generate, edit, and explore AI images. Flux, Imagen, LoRA identity swap, upscale, and more.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenancePowerful image generation system leveraging multiple Stable Diffusion models (flux-schnell, flux-dev, sdxl, sd3, sd15) for creating high-quality AI-generated images with precise customization.19MIT
- AlicenseAqualityDmaintenanceEnables comprehensive ComfyUI workflow automation including image generation, workflow management, node discovery, and system monitoring through natural language interactions with local or remote ComfyUI servers.3114MIT
- AlicenseNot gradedqualityNot gradedmaintenanceEnables high-quality image generation with advanced text rendering and contextual understanding using FLUX.1 Kontext Max through the Replicate API. Supports text-to-image generation, image editing, multiple aspect ratios, and automatically downloads generated images locally.1-
- AlicenseAqualityBmaintenanceAI image generation with 6 Flux models (flux-dev, flux-pro, flux-kontext) including context-aware image editing, async task management, and built-in model guide.6128 PyPI3MIT