Skip to main content
Glama

MCP フラックススタジオ

鍛冶屋のバッジ

Flux の高度な画像生成機能を AI コーディングアシスタントに提供する、強力なモデルコンテキストプロトコル (MCP) サーバーです。このサーバーにより、Flux の画像生成、操作、制御機能を Cursor および Windsurf (Codeium) IDE に直接統合できます。

概要

MCP Flux Studio は、AI コーディング アシスタントと Flux の強力な画像生成 API 間のギャップを埋め、画像生成機能を開発ワークフローに直接シームレスに統合できるようにします。

特徴

  • 画像生成

    • 正確な制御によるテキストから画像への生成

    • 複数のモデルのサポート (flux.1.1-pro、flux.1-pro、flux.1-dev、flux.1.1-ultra)

    • カスタマイズ可能なアスペクト比と寸法

  • 画像操作

    • 画像から画像への変換

    • カスタマイズ可能なマスクによるインペインティング

    • 解像度のアップスケーリングと強化

  • 高度なコントロール

    • エッジベースの生成(Canny)

    • 深度を考慮した生成

    • ポーズガイド生成

  • IDE統合

    • カーソルの完全サポート (v0.45.7+)

    • Windsurf/Codeium Cascade (Wave 3+) と互換性あり

    • AIアシスタントによるシームレスなツール呼び出し

Related MCP server: Flux Schnell MCP Server

クイックスタート

  1. 前提条件

    • Node.js 18歳以上

    • Python 3.12以上

    • Flux APIキー

    • 互換性のある IDE (Cursor または Windsurf)

  2. インストール

Smithery経由でインストール

Smithery経由で Flux Studio for Claude Desktop を自動的にインストールするには:

npx -y @smithery/cli install @jmanhype/mcp-flux-studio --client claude

手動インストール

git clone https://github.com/jmanhype/mcp-flux-studio.git
cd mcp-flux-studio
npm install
npm run build
  1. 基本構成

    BFL_API_KEY=your_flux_api_key
    FLUX_PATH=/path/to/flux/installation

IDE 固有の構成やトラブルシューティングを含む詳細なセットアップ手順については、インストール ガイドを参照してください。

ドキュメント

IDE統合

カーソル (v0.45.7+)

MCP Flux Studio は、Cursor の AI アシスタントとシームレスに統合されます。

  1. 構成

    • 設定 > 機能 > MCP から設定します。

    • stdioとSSE接続の両方をサポート

    • 環境変数はラッパースクリプト経由で設定できます

  2. 使用法

    • カーソルのAIアシスタントに自動的に利用できるツール

    • ツールの呼び出しにはユーザーの承認が必要です

    • 生成の進行状況に関するリアルタイムフィードバック

ウィンドサーフィン/コディウム(ウェーブ3以上)

Windsurf の Cascade AI との統合:

  1. 構成

    • ~/.codeium/windsurf/mcp_config.jsonを編集します。

    • プロセスベースのツール実行をサポート

    • JSONで設定された環境変数

  2. 使用法

    • Cascade の MCP ツールバーからツールにアクセスします

    • 自動ツール検出と読み込み

    • CascadeのAI機能と統合

IDE 固有の詳細なセットアップ手順については、インストール ガイドを参照してください。

使用法

サーバーは次のツールを提供します。

生成する

テキストプロンプトから画像を生成します。

{
  "prompt": "A photorealistic cat",
  "model": "flux.1.1-pro",
  "aspect_ratio": "1:1",
  "output": "generated.jpg"
}

画像2画像

別の画像を参照として使用して画像を生成します。

{
  "image": "input.jpg",
  "prompt": "Convert to oil painting",
  "model": "flux.1.1-pro",
  "strength": 0.85,
  "output": "output.jpg",
  "name": "oil_painting"
}

インペイント

マスクを使用して画像を修復します。

{
  "image": "input.jpg",
  "prompt": "Add flowers",
  "mask_shape": "circle",
  "position": "center",
  "output": "inpainted.jpg"
}

コントロール

構造制御を使用して画像を生成します。

{
  "type": "canny",
  "image": "control.jpg",
  "prompt": "A realistic photo",
  "output": "controlled.jpg"
}

発達

プロジェクト構造

flux-mcp-server/
├── src/
│   ├── index.ts          # Main server implementation
│   └── types.ts          # TypeScript type definitions
├── tests/
│   └── server.test.ts    # Server tests
├── docs/
│   ├── API.md           # API documentation
│   └── CONTRIBUTING.md  # Contribution guidelines
├── examples/
│   ├── generate.json    # Example tool usage
│   └── config.json      # Example configuration
├── package.json
├── tsconfig.json
└── README.md

テストの実行

npm test

建物

npm run build

貢献

行動規範とプル リクエストの送信プロセスの詳細については、 CONTRIBUTING.md をお読みください。

ライセンス

このプロジェクトは MIT ライセンスに基づいてライセンスされています - 詳細についてはLICENSEファイルを参照してください。

謝辞

Available Tools

4 tools
controlC

Generate an image using structural control

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYesType of control to use
imageYesInput control image path
promptYesText prompt for generation
stepsNoNumber of inference steps
guidanceNoGuidance scale
outputNoOutput filename

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions 'structural control' but doesn't explain what this entails operationally—such as how control affects generation, whether it modifies existing images or creates new ones, potential side effects, or performance characteristics. This leaves significant gaps for a tool with 6 parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words. It's appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, no annotations, no output schema), the description is incomplete. It doesn't explain what 'structural control' means, how it interacts with parameters, or what the tool returns. For a generation tool with multiple controls, more context is needed to guide effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all 6 parameters. The description adds no additional meaning about parameters beyond implying 'structural control' relates to the 'type' parameter. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Generate an image using structural control' states a clear purpose (generating images with control mechanisms) but is vague about what 'structural control' means and doesn't distinguish from sibling tools like 'generate', 'img2img', or 'inpaint'. It doesn't specify what makes this tool unique compared to those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling tools ('generate', 'img2img', 'inpaint'). There's no mention of appropriate contexts, prerequisites, or exclusions. The agent must infer usage from the tool name and parameters alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateC

Generate an image from a text prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText prompt for image generation
modelNoModel to use for generationflux.1.1-pro
aspect_ratioNoAspect ratio of the output image
widthNoImage width (ignored if aspect-ratio is set)
heightNoImage height (ignored if aspect-ratio is set)
outputNoOutput filenamegenerated.jpg

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure but offers minimal information. It mentions generation but doesn't cover critical aspects like whether this is a read-only or destructive operation, potential rate limits, authentication needs, or what the output entails (e.g., image format, storage location). This leaves significant gaps in understanding the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded with a single, clear sentence that directly states the tool's core function. There is no wasted language or redundancy, making it efficient and easy to parse, though this brevity contributes to gaps in other dimensions like guidelines and transparency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a 6-parameter image generation tool with no annotations and no output schema, the description is incomplete. It fails to address behavioral traits, usage context, or output details (e.g., what is returned, error handling), leaving the agent under-informed for effective tool invocation in a real-world scenario.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter semantics beyond what the input schema already provides, as schema description coverage is 100%. The schema thoroughly documents all 6 parameters, including enums for 'model' and 'aspect_ratio', defaults, and dependencies (e.g., 'width'/'height' ignored if 'aspect-ratio' set). Thus, the description meets the baseline but doesn't enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('generate') and resource ('image from a text prompt'), making it immediately understandable. However, it doesn't differentiate from sibling tools like 'img2img' or 'inpaint' which likely also generate images but from different inputs, leaving room for potential confusion about when to choose this specific tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'control', 'img2img', or 'inpaint'. It lacks context about prerequisites, such as needing a text prompt as input, or exclusions, like not being suitable for image-to-image transformations. This absence leaves the agent without clear direction for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

img2imgC

Generate an image using another image as reference

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesInput image path
promptYesText prompt for generation
modelNoModel to use for generationflux.1.1-pro
strengthNoGeneration strength
widthNoOutput image width
heightNoOutput image height
outputNoOutput filenameoutputs/generated.jpg
nameYesName for the generation

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool generates an image but doesn't disclose behavioral traits such as whether it overwrites files, requires specific permissions, has rate limits, or what the output format/behavior is (e.g., file creation, error handling). This is a significant gap for a tool with 8 parameters and no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose without any wasted words. It's appropriately sized for the tool's complexity, making it easy to scan and understand quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, no annotations, no output schema), the description is insufficient. It doesn't explain the tool's behavior, output (e.g., file saved to disk), or usage context relative to siblings. For an image generation tool with multiple parameters, more detail is needed to guide effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all parameters well (e.g., 'image' as input path, 'prompt' for text, 'strength' for generation intensity). The description adds no additional meaning beyond implying the 'image' parameter is used as a reference, which is somewhat redundant with the schema. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate an image using another image as reference.' It specifies both the action ('generate') and the resource ('image'), though it doesn't explicitly differentiate from sibling tools like 'generate' or 'inpaint' beyond the reference image aspect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'generate' (which likely generates from text only) or 'inpaint' (which might modify parts of an image). It mentions using an image as reference but doesn't clarify scenarios or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inpaintC

Inpaint an image using a mask

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesInput image path
promptYesText prompt for inpainting
mask_shapeNoShape of the maskcircle
positionNoPosition of the maskcenter
outputNoOutput filenameinpainted.jpg

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action ('inpaint') but doesn't explain what inpainting entails (e.g., filling masked areas based on a prompt), potential side effects, permissions needed, or output behavior. This leaves significant gaps for a tool that modifies images.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste—'Inpaint an image using a mask'—making it highly concise and front-loaded. Every word earns its place by conveying the core action and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of an image inpainting tool with no annotations and no output schema, the description is insufficient. It doesn't explain what inpainting does, how the output is handled, or any behavioral traits, leaving the agent with incomplete context for proper tool selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (image, prompt, mask_shape, position, output) with descriptions and enums. The description adds no additional meaning beyond what the schema provides, such as explaining how the prompt influences inpainting or how mask shape/position interact. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Inpaint an image using a mask' clearly states the action (inpaint) and resource (image with mask), making the purpose understandable. However, it doesn't differentiate from sibling tools like 'img2img' or 'generate', which might also involve image manipulation, so it's not a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'img2img' or 'generate'. It lacks context about specific use cases, prerequisites, or exclusions, leaving the agent to infer usage based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updates
    • First observedcontrol
    • First observedgenerate
    • First observedimg2img
    • First observedinpaint

TDQS

B3.2/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: 'control' uses structural guidance, 'generate' creates from text, 'img2img' references an image, and 'inpaint' modifies with a mask. The descriptions make it easy to differentiate between structural generation, text-to-image, image-to-image, and inpainting workflows.

Naming Consistency3/5

The naming is mixed: 'control' and 'generate' are verbs only, while 'img2img' and 'inpaint' are compound terms. There's no consistent pattern like verb_noun, but the names are still readable and descriptive of their functions, avoiding chaotic conventions.

Tool Count5/5

With 4 tools, this is well-scoped for an image generation server. Each tool earns its place by covering distinct aspects of image creation and manipulation, providing a focused set without being too thin or overwhelming for the domain.

Completeness4/5

The toolset covers core image generation workflows: text-to-image, image-to-image, inpainting, and controlled generation. A minor gap might be the lack of tools for post-processing or batch operations, but the essential CRUD-like operations for image creation are well-represented.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers