Skip to main content
Glama
kira4094

MiniMax Vision MCP Server

by kira4094

MiniMax Vision MCP Server

MCP server for MiniMax vision models — analyze images through the OpenAI-compatible POST /v1/chat/completions endpoint.

Features

  • đŸ–ŧī¸ Analyze images — local files (png/jpg/jpeg/gif/webp/bmp) and remote HTTP(S) URLs

  • 🧠 MiniMax-M3: image + video understanding, 1M context, adaptive thinking

  • đŸŽ›ī¸ temperature fully configurable [0, 2] (unlike some providers that lock it)

  • 💭 thinking flag → enables adaptive thinking + reasoning_split on M3

  • ⚡ Zero non-MCP dependencies, one-line npx deploy

Related MCP server: z_ai_vision_mcp_server_clone

Requirements

  • Node.js >= 18

  • A MiniMax platform API key (č´ĻæˆˇįŽĄį† → æŽĨåŖå¯†é’Ĩ)

Install & run

cd minimax-vision-mcp-server
npm install
npm start

Environment variables

Variable

Required

Default

Description

MINIMAX_API_KEY

✅

—

Your MiniMax API key.

MINIMAX_MODEL

MiniMax-M3

Model: MiniMax-M3, MiniMax-M2.7, MiniMax-M2.7-highspeed, MiniMax-M2.5, MiniMax-M2.5-highspeed, MiniMax-M2.1, MiniMax-M2.1-highspeed, MiniMax-M2.

MINIMAX_BASE_URL

https://api.minimax.chat/v1

Override endpoint (for proxies).

MINIMAX_MAX_TOKENS

8192

Default max output tokens. M3 supports up to 524288.

Note: International users may use https://api.minimaxi.com/v1 as the base URL.

Claude Code / CC-Switch config

{
  "mcpServers": {
    "minimax-vision": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "minimax-vision-mcp-server"],
      "env": {
        "MINIMAX_API_KEY": "your-minimax-api-key",
        "MINIMAX_MODEL": "MiniMax-M3"
      }
    }
  }
}

Tool: minimax_vision_understand

Parameter

Type

Required

Description

image

string

✅

Local image path or remote HTTP(S) URL.

prompt

string

✅

What to ask about the image.

max_tokens

number

Max output tokens. Default 8192.

temperature

number

0-2, default 1.

thinking

bool

Enable adaptive thinking + reasoning split (M3). M2.x always think; ignored there.

Why MiniMax for vision?

  • 1M context on M3 — analyze long documents alongside images

  • Native image + video understanding

  • Configurable temperature — fine-grained control over determinism

  • Adaptive thinking — reasoning on by default, splittable via reasoning_split

License

MIT

Available Tools

1 tool
minimax_vision_understandA

Analyze an image using MiniMax vision models. Supports local image files (png/jpg/jpeg/gif/webp/bmp) AND remote HTTP(S) URLs. Default model: MiniMax-M3. MiniMax-M3 supports image + video understanding, 1M context, adaptive thinking.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesLocal image file path OR remote HTTP(S) URL. Both are supported by MiniMax.
promptYesWhat to ask about the image. Be specific.
thinkingNoEnable adaptive thinking + reasoning split (M3). M2.x models always think; param ignored there.
max_tokensNoMaximum output tokens.
temperatureNoSampling temperature (0-2). Default 1.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It discloses the default model, support for video understanding, 1M context, and adaptive thinking, and adds a note about the 'thinking' parameter being ignored on M2.x models. This is transparent about operational nuances, though it doesn't explicitly state the tool is read-only or describe the output format, which are minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a tight three sentences: it states the core purpose, lists supported formats, and highlights model capabilities. Every sentence provides useful information without fluff, and it is front-loaded with the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters and no output schema, the description covers a lot: input formats, default model, model capabilities, and parameter nuance. It doesn't explicitly state the return type (likely text analysis), which would be helpful, but it is reasonably complete for an AI-driven vision tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all 5 parameters, so the baseline is 3. The description adds value by explaining the image parameter supports both local files and URLs (reinforcing schema), and crucially notes that the 'thinking' parameter is ignored on M2.x models. This extra context goes beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Analyze an image using MiniMax vision models' with a specific verb and resource, and distinguishes the tool's capabilities (local/remote image support, model defaults). This is a clear, specific purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (for image analysis) and even clarifies supported input formats and URL vs local files. However, with no sibling tools to differentiate from, explicit alternatives or exclusions are absent. The guidance is adequate but not explicit about when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedminimax_vision_understand

TDQS

A4.1/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no possibility of confusing it with another. The tool's purpose is clearly defined for image understanding.

Naming Consistency5/5

The single tool name 'minimax_vision_understand' follows a clear verb_noun pattern, and with only one tool, naming consistency is trivially perfect.

Tool Count3/5

A single tool feels thin for a server named 'Vision MCP Server', but it covers the core image understanding use case. It is borderline but not severely under-scoped.

Completeness3/5

The tool covers basic image understanding and supports multiple input formats, but lacks options for model selection or video understanding, despite the underlying model supporting video. Some common vision tasks like OCR or object detection are not present, but that may be out of scope.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers