MCP Flux Studio
MCP 플럭스 스튜디오
Flux의 고급 이미지 생성 기능을 AI 코딩 어시스턴트에 제공하는 강력한 모델 컨텍스트 프로토콜(MCP) 서버입니다. 이 서버를 통해 Flux의 이미지 생성, 조작 및 제어 기능을 Cursor 및 Windsurf(Codeium) IDE에 직접 통합할 수 있습니다.
개요
MCP Flux Studio는 AI 코딩 어시스턴트와 Flux의 강력한 이미지 생성 API 간의 격차를 해소하여 이미지 생성 기능을 개발 워크플로에 원활하게 통합할 수 있도록 해줍니다.
특징
이미지 생성
정밀한 제어를 통한 텍스트-이미지 생성
다양한 모델 지원(flux.1.1-pro, flux.1-pro, flux.1-dev, flux.1.1-ultra)
사용자 정의 가능한 종횡비 및 치수
이미지 조작
이미지-이미지 변환
사용자 정의 가능한 마스크로 페인팅
해상도 확대 및 향상
고급 컨트롤
엣지 기반 생성(canny)
깊이 인식 세대
포즈 가이드 생성
IDE 통합
커서에 대한 전체 지원(v0.45.7+)
Windsurf/Codeium Cascade(Wave 3+)와 호환 가능
AI 어시스턴트를 통한 원활한 도구 호출
Related MCP server: Flux Schnell MCP Server
빠른 시작
필수 조건
노드.js 18+
파이썬 3.12+
플럭스 API 키
호환 IDE(Cursor 또는 Windsurf)
설치
Smithery를 통해 설치
Smithery를 통해 Claude Desktop용 Flux Studio를 자동으로 설치하려면:
지엑스피1
수동 설치
git clone https://github.com/jmanhype/mcp-flux-studio.git
cd mcp-flux-studio
npm install
npm run build기본 구성
BFL_API_KEY=your_flux_api_key FLUX_PATH=/path/to/flux/installation
IDE별 구성 및 문제 해결을 포함한 자세한 설정 지침은 설치 가이드를 참조하세요.
선적 서류 비치
IDE 통합
커서(v0.45.7+)
MCP Flux Studio는 Cursor의 AI 어시스턴트와 완벽하게 통합됩니다.
구성
설정 > 기능 > MCP를 통해 구성하세요.
stdio와 SSE 연결을 모두 지원합니다
환경 변수는 래퍼 스크립트를 통해 설정할 수 있습니다.
용법
Cursor의 AI 어시스턴트가 자동으로 사용할 수 있는 도구
도구 호출에는 사용자 승인이 필요합니다.
세대 진행 상황에 대한 실시간 피드백
윈드서핑/코디움(파도 3+)
Windsurf의 Cascade AI와 통합:
구성
~/.codeium/windsurf/mcp_config.json편집하세요프로세스 기반 도구 실행을 지원합니다
JSON으로 구성된 환경 변수
용법
Cascade의 MCP 도구 모음을 통해 도구에 액세스하세요
자동 도구 검색 및 로딩
Cascade의 AI 기능과 통합
IDE별 자세한 설정 지침은 설치 가이드를 참조하세요.
용법
서버는 다음과 같은 도구를 제공합니다.
생성하다
텍스트 프롬프트에서 이미지를 생성합니다.
{
"prompt": "A photorealistic cat",
"model": "flux.1.1-pro",
"aspect_ratio": "1:1",
"output": "generated.jpg"
}이미지2이미지
다른 이미지를 참조하여 이미지를 생성합니다.
{
"image": "input.jpg",
"prompt": "Convert to oil painting",
"model": "flux.1.1-pro",
"strength": 0.85,
"output": "output.jpg",
"name": "oil_painting"
}인페인트
마스크를 사용하여 이미지를 칠합니다.
{
"image": "input.jpg",
"prompt": "Add flowers",
"mask_shape": "circle",
"position": "center",
"output": "inpainted.jpg"
}제어
구조적 제어를 사용하여 이미지를 생성합니다.
{
"type": "canny",
"image": "control.jpg",
"prompt": "A realistic photo",
"output": "controlled.jpg"
}개발
프로젝트 구조
flux-mcp-server/
├── src/
│ ├── index.ts # Main server implementation
│ └── types.ts # TypeScript type definitions
├── tests/
│ └── server.test.ts # Server tests
├── docs/
│ ├── API.md # API documentation
│ └── CONTRIBUTING.md # Contribution guidelines
├── examples/
│ ├── generate.json # Example tool usage
│ └── config.json # Example configuration
├── package.json
├── tsconfig.json
└── README.md테스트 실행
npm test건물
npm run build기여하다
행동 강령과 풀 리퀘스트 제출 프로세스에 대한 자세한 내용은 CONTRIBUTING.md를 읽어보세요.
특허
이 프로젝트는 MIT 라이선스에 따라 라이선스가 부여되었습니다. 자세한 내용은 라이선스 파일을 참조하세요.
감사의 말
모델 컨텍스트 프로토콜 - 프로토콜 사양
Flux API - 기본 이미지 생성 API
Available Tools
4 toolscontrolC
Generate an image using structural control
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | Type of control to use | |
| image | Yes | Input control image path | |
| prompt | Yes | Text prompt for generation | |
| steps | No | Number of inference steps | |
| guidance | No | Guidance scale | |
| output | No | Output filename |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions 'structural control' but doesn't explain what this entails operationally—such as how control affects generation, whether it modifies existing images or creates new ones, potential side effects, or performance characteristics. This leaves significant gaps for a tool with 6 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero wasted words. It's appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (6 parameters, no annotations, no output schema), the description is incomplete. It doesn't explain what 'structural control' means, how it interacts with parameters, or what the tool returns. For a generation tool with multiple controls, more context is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 6 parameters. The description adds no additional meaning about parameters beyond implying 'structural control' relates to the 'type' parameter. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Generate an image using structural control' states a clear purpose (generating images with control mechanisms) but is vague about what 'structural control' means and doesn't distinguish from sibling tools like 'generate', 'img2img', or 'inpaint'. It doesn't specify what makes this tool unique compared to those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling tools ('generate', 'img2img', 'inpaint'). There's no mention of appropriate contexts, prerequisites, or exclusions. The agent must infer usage from the tool name and parameters alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generateC
Generate an image from a text prompt
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text prompt for image generation | |
| model | No | Model to use for generation | flux.1.1-pro |
| aspect_ratio | No | Aspect ratio of the output image | |
| width | No | Image width (ignored if aspect-ratio is set) | |
| height | No | Image height (ignored if aspect-ratio is set) | |
| output | No | Output filename | generated.jpg |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure but offers minimal information. It mentions generation but doesn't cover critical aspects like whether this is a read-only or destructive operation, potential rate limits, authentication needs, or what the output entails (e.g., image format, storage location). This leaves significant gaps in understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded with a single, clear sentence that directly states the tool's core function. There is no wasted language or redundancy, making it efficient and easy to parse, though this brevity contributes to gaps in other dimensions like guidelines and transparency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a 6-parameter image generation tool with no annotations and no output schema, the description is incomplete. It fails to address behavioral traits, usage context, or output details (e.g., what is returned, error handling), leaving the agent under-informed for effective tool invocation in a real-world scenario.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no parameter semantics beyond what the input schema already provides, as schema description coverage is 100%. The schema thoroughly documents all 6 parameters, including enums for 'model' and 'aspect_ratio', defaults, and dependencies (e.g., 'width'/'height' ignored if 'aspect-ratio' set). Thus, the description meets the baseline but doesn't enhance parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('generate') and resource ('image from a text prompt'), making it immediately understandable. However, it doesn't differentiate from sibling tools like 'img2img' or 'inpaint' which likely also generate images but from different inputs, leaving room for potential confusion about when to choose this specific tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'control', 'img2img', or 'inpaint'. It lacks context about prerequisites, such as needing a text prompt as input, or exclusions, like not being suitable for image-to-image transformations. This absence leaves the agent without clear direction for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
img2imgC
Generate an image using another image as reference
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | Input image path | |
| prompt | Yes | Text prompt for generation | |
| model | No | Model to use for generation | flux.1.1-pro |
| strength | No | Generation strength | |
| width | No | Output image width | |
| height | No | Output image height | |
| output | No | Output filename | outputs/generated.jpg |
| name | Yes | Name for the generation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool generates an image but doesn't disclose behavioral traits such as whether it overwrites files, requires specific permissions, has rate limits, or what the output format/behavior is (e.g., file creation, error handling). This is a significant gap for a tool with 8 parameters and no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without any wasted words. It's appropriately sized for the tool's complexity, making it easy to scan and understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, no annotations, no output schema), the description is insufficient. It doesn't explain the tool's behavior, output (e.g., file saved to disk), or usage context relative to siblings. For an image generation tool with multiple parameters, more detail is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all parameters well (e.g., 'image' as input path, 'prompt' for text, 'strength' for generation intensity). The description adds no additional meaning beyond implying the 'image' parameter is used as a reference, which is somewhat redundant with the schema. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate an image using another image as reference.' It specifies both the action ('generate') and the resource ('image'), though it doesn't explicitly differentiate from sibling tools like 'generate' or 'inpaint' beyond the reference image aspect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'generate' (which likely generates from text only) or 'inpaint' (which might modify parts of an image). It mentions using an image as reference but doesn't clarify scenarios or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inpaintC
Inpaint an image using a mask
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | Input image path | |
| prompt | Yes | Text prompt for inpainting | |
| mask_shape | No | Shape of the mask | circle |
| position | No | Position of the mask | center |
| output | No | Output filename | inpainted.jpg |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action ('inpaint') but doesn't explain what inpainting entails (e.g., filling masked areas based on a prompt), potential side effects, permissions needed, or output behavior. This leaves significant gaps for a tool that modifies images.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste—'Inpaint an image using a mask'—making it highly concise and front-loaded. Every word earns its place by conveying the core action and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of an image inpainting tool with no annotations and no output schema, the description is insufficient. It doesn't explain what inpainting does, how the output is handled, or any behavioral traits, leaving the agent with incomplete context for proper tool selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (image, prompt, mask_shape, position, output) with descriptions and enums. The description adds no additional meaning beyond what the schema provides, such as explaining how the prompt influences inpainting or how mask shape/position interact. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Inpaint an image using a mask' clearly states the action (inpaint) and resource (image with mask), making the purpose understandable. However, it doesn't differentiate from sibling tools like 'img2img' or 'generate', which might also involve image manipulation, so it's not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'img2img' or 'generate'. It lacks context about specific use cases, prerequisites, or exclusions, leaving the agent to infer usage based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
control - First observed
generate - First observed
img2img - First observed
inpaint
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose with no overlap: 'control' uses structural guidance, 'generate' creates from text, 'img2img' references an image, and 'inpaint' modifies with a mask. The descriptions make it easy to differentiate between structural generation, text-to-image, image-to-image, and inpainting workflows.
The naming is mixed: 'control' and 'generate' are verbs only, while 'img2img' and 'inpaint' are compound terms. There's no consistent pattern like verb_noun, but the names are still readable and descriptive of their functions, avoiding chaotic conventions.
With 4 tools, this is well-scoped for an image generation server. Each tool earns its place by covering distinct aspects of image creation and manipulation, providing a focused set without being too thin or overwhelming for the domain.
The toolset covers core image generation workflows: text-to-image, image-to-image, inpainting, and controlled generation. A minor gap might be the lack of tools for post-processing or batch operations, but the essential CRUD-like operations for image creation are well-represented.
Maintenance
Related MCP Connectors
MCP server for Flux AI image generation
Official FLUX MCP server. Generate, edit, vary, and browse images from Black Forest Labs.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
The Canva MCP server connects AI assistants (like Claude, ChatGPT, and Cursor) to Canva's API, enabling them to create and manage designs directly within chat conversations. Key capabilities include generating new designs from prompts, autofilling templates, searching and resizing existing designs, importing files from URLs, exporting designs as PDFs or images, and managing folders and comments without switching between tools.
Related MCP Servers
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables high-quality image generation using the Flux.1 Schnell model via Together AI with customizable parameters.131 npm10MIT
- AlicenseNot gradedqualityDmaintenanceA server that enables generating images through the Replicate API by calling the Flux Schnell model via the Model Context Protocol (MCP).3MIT
- AlicenseBqualityDmaintenanceAn MCP server that enables AI assistants to generate images using Black Forest Labs' Flux model via Cloudflare Workers.1MIT
- AlicenseAqualityBmaintenanceAI image generation with 6 Flux models (flux-dev, flux-pro, flux-kontext) including context-aware image editing, async task management, and built-in model guide.683 PyPI3MIT