MiniMax Vision MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MiniMax Vision MCP ServerDescribe what's happening in this image: /tmp/screenshot.png"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MiniMax Vision MCP Server
MCP server for MiniMax vision models â analyze images through the
OpenAI-compatible POST /v1/chat/completions endpoint.
Features
đŧī¸ Analyze images â local files (png/jpg/jpeg/gif/webp/bmp) and remote HTTP(S) URLs
đ§ MiniMax-M3: image + video understanding, 1M context, adaptive thinking
đī¸
temperaturefully configurable[0, 2](unlike some providers that lock it)đ
thinkingflag â enables adaptive thinking +reasoning_spliton M3⥠Zero non-MCP dependencies, one-line
npxdeploy
Related MCP server: z_ai_vision_mcp_server_clone
Requirements
Node.js >= 18
A MiniMax platform API key (č´ĻæˇįŽĄį â æĨåŖå¯éĨ)
Install & run
cd minimax-vision-mcp-server
npm install
npm startEnvironment variables
Variable | Required | Default | Description |
| â | â | Your MiniMax API key. |
|
| Model: | |
|
| Override endpoint (for proxies). | |
|
| Default max output tokens. M3 supports up to 524288. |
Note: International users may use
https://api.minimaxi.com/v1as the base URL.
Claude Code / CC-Switch config
{
"mcpServers": {
"minimax-vision": {
"type": "stdio",
"command": "npx",
"args": ["-y", "minimax-vision-mcp-server"],
"env": {
"MINIMAX_API_KEY": "your-minimax-api-key",
"MINIMAX_MODEL": "MiniMax-M3"
}
}
}
}Tool: minimax_vision_understand
Parameter | Type | Required | Description |
| string | â | Local image path or remote HTTP(S) URL. |
| string | â | What to ask about the image. |
| number | Max output tokens. Default 8192. | |
| number | 0-2, default 1. | |
| bool | Enable adaptive thinking + reasoning split (M3). M2.x always think; ignored there. |
Why MiniMax for vision?
1M context on M3 â analyze long documents alongside images
Native image + video understanding
Configurable temperature â fine-grained control over determinism
Adaptive thinking â reasoning on by default, splittable via
reasoning_split
Related projects
kimi-vision-mcp-server â Moonshot Kimi
doubao-vision-mcp-server â ByteDance Doubao
glm-vision-mcp-server â Zhipu GLM
qwen-vision-mcp-server â Alibaba Qwen
@kira4094/agnes-image-mcp-server â Agnes Image (text-to-image)
License
MIT
Available Tools
1 toolminimax_vision_understandA
Analyze an image using MiniMax vision models. Supports local image files (png/jpg/jpeg/gif/webp/bmp) AND remote HTTP(S) URLs. Default model: MiniMax-M3. MiniMax-M3 supports image + video understanding, 1M context, adaptive thinking.
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | Local image file path OR remote HTTP(S) URL. Both are supported by MiniMax. | |
| prompt | Yes | What to ask about the image. Be specific. | |
| thinking | No | Enable adaptive thinking + reasoning split (M3). M2.x models always think; param ignored there. | |
| max_tokens | No | Maximum output tokens. | |
| temperature | No | Sampling temperature (0-2). Default 1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It discloses the default model, support for video understanding, 1M context, and adaptive thinking, and adds a note about the 'thinking' parameter being ignored on M2.x models. This is transparent about operational nuances, though it doesn't explicitly state the tool is read-only or describe the output format, which are minor gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a tight three sentences: it states the core purpose, lists supported formats, and highlights model capabilities. Every sentence provides useful information without fluff, and it is front-loaded with the primary action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters and no output schema, the description covers a lot: input formats, default model, model capabilities, and parameter nuance. It doesn't explicitly state the return type (likely text analysis), which would be helpful, but it is reasonably complete for an AI-driven vision tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers all 5 parameters, so the baseline is 3. The description adds value by explaining the image parameter supports both local files and URLs (reinforcing schema), and crucially notes that the 'thinking' parameter is ignored on M2.x models. This extra context goes beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Analyze an image using MiniMax vision models' with a specific verb and resource, and distinguishes the tool's capabilities (local/remote image support, model defaults). This is a clear, specific purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (for image analysis) and even clarifies supported input formats and URL vs local files. However, with no sibling tools to differentiate from, explicit alternatives or exclusions are absent. The guidance is adequate but not explicit about when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
minimax_vision_understand
TDQS
Scored across 1 tool
Only one tool exists, so there is no possibility of confusing it with another. The tool's purpose is clearly defined for image understanding.
The single tool name 'minimax_vision_understand' follows a clear verb_noun pattern, and with only one tool, naming consistency is trivially perfect.
A single tool feels thin for a server named 'Vision MCP Server', but it covers the core image understanding use case. It is borderline but not severely under-scoped.
The tool covers basic image understanding and supports multiple input formats, but lacks options for model selection or video understanding, despite the underlying model supporting video. Some common vision tasks like OCR or object detection are not present, but that may be out of scope.
Maintenance
Related MCP Connectors
MCP server for MiniMax H3 multimodal video generation
MCP server for Hailuo (MiniMax) AI video generation
MCP server for Qwen Image 3 AI image generation
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
Related MCP Servers
- AlicenseAqualityCmaintenanceAn MCP server for analyzing images using ModelScope's vision models. Supports both local files and URLs, enabling image content description and question answering.1166 npm11MIT
- FlicenseBqualityBmaintenanceOpenAI-compatible MCP server for running image analysis tools against your own vision model endpoint.74 npm-
- AlicenseNot gradedqualityCmaintenanceMCP server for analyzing images using multiple vision LLM providers (OpenCode, OpenAI, Anthropic, Google, and custom OpenAI-compatible endpoints). Provides tools to analyze single or multiple images, list providers, and test vision capabilities.MIT
- AlicenseNot gradedqualityAmaintenanceEnables image analysis via OpenAI-compatible vision APIs, supporting local files, URLs, and base64 inputs with intelligent tiling for high-resolution images. Provides a secure, configurable MCP stdio server for structured vision analysis.151 npm2MIT