Skip to main content
Glama
kira4094

Qwen Vision MCP Server

by kira4094

Qwen Vision MCP Server

MCP server for Alibaba Qwen3.7-plus (multimodal vision model), connected via DashScope's Anthropic-compatible endpoint.

Features

  • Single tool: qwen_vision_understand — analyze an image with Qwen3.7-plus

  • Supports local image files (png/jpg/jpeg/gif/webp/bmp) and remote URLs

  • Same Anthropic Messages API format as Claude

  • Token usage reported back in each response

Related MCP server: Image Parse MCP

Requirements

Install

cd qwen-vision-mcp-server
npm install

Configuration

Environment variables:

Variable

Required

Default

Description

DASHSCOPE_API_KEY

Your DashScope API key

QWEN_MODEL

qwen3.7-plus

Model name (e.g. qwen-vl-plus, qwen3-vl-plus)

QWEN_BASE_URL

https://dashscope.aliyuncs.com/apps/anthropic/v1/messages

Override endpoint

QWEN_MAX_TOKENS

4096

Default max output tokens

Run locally

DASHSCOPE_API_KEY=sk-xxx npm start

CC-Switch / Claude Code config

{
  "mcpServers": {
    "qwen-vision": {
      "type": "stdio",
      "command": "cmd",
      "args": ["/c", "npx", "-y", "qwen-vision-mcp-server"],
      "env": {
        "DASHSCOPE_API_KEY": "sk-your-dashscope-key",
        "QWEN_MODEL": "qwen3.7-plus"
      }
    }
  }
}

Or if installed locally:

{
  "mcpServers": {
    "qwen-vision": {
      "type": "stdio",
      "command": "node",
      "args": ["E:/Projects/Claude/MCP/qwen/qwen-vision-mcp-server/src/index.js"],
      "env": {
        "DASHSCOPE_API_KEY": "sk-your-dashscope-key"
      }
    }
  }
}

Tool: qwen_vision_understand

Parameter

Type

Required

Description

image

string

Local file path or HTTP(S) URL

prompt

string

What to ask about the image

max_tokens

number

Max output tokens (default 4096)

temperature

number

0-2, default 1

Example:

Use qwen_vision_understand to analyze this UI mockup: 
/Users/me/Downloads/login-page.png
"Recreate this as HTML with Tailwind CSS"

License

MIT

Available Tools

1 tool
qwen_vision_understandA

Analyze an image using Alibaba Qwen3.7-plus (multimodal vision model). Supports local image files and remote URLs. Reaches DashScope via the Anthropic-compatible endpoint (https://dashscope.aliyuncs.com/apps/anthropic). Excels at: UI screenshot→code, design mockup analysis, visual debugging, chart/document understanding.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYesImage source: local file path (e.g. C:/path/to/screenshot.png) or URL (https://...)
promptYesWhat to ask about the image. Be specific for best results. E.g.: 'Recreate this UI as HTML with Tailwind CSS'
max_tokensNoMaximum output tokens
temperatureNoSampling temperature (0-2)

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses the model, endpoint, and supported image sources, but lacks details on output format, latency, cost, or potential failure modes. A read from the agent perspective would want to know what the tool returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences plus a compact bullet-like list. Front-loaded with the core purpose, each sentence adds unique value: model, supported sources, endpoint, and use cases. Zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, description is the sole source of context. It covers input and use cases adequately, but omits output format, error behavior, and performance characteristics. Adequate for a straightforward vision analysis tool, but not comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so schema already documents all parameters. Description adds examples and context (e.g., 'be specific for best results') but does not significantly enhance meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it analyzes images using a specific model (Qwen3.7-plus), with a list of concrete use cases (UI to code, design mockup analysis, etc.). Verb+resource combination is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implicit guidance through the list of use cases (e.g., 'UI screenshot→code'), but no explicit when-to-use, when-not-to-use, or alternatives. Since there are no sibling tools, the miss is less severe, but still room for improvement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedqwen_vision_understand

TDQS

A3.6/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no possibility of confusion or overlap.

Naming Consistency5/5

The single tool follows a clear verb_noun pattern (qwen_vision_understand), consistent with common MCP naming conventions.

Tool Count2/5

A vision MCP server typically requires multiple specialized tools (e.g., describe, analyze, compare) rather than a single monolithic tool. One tool feels under-scoped for the domain.

Completeness3/5

The single tool covers a broad range of vision tasks (UI analysis, charts, documents), but lacks structured decomposition into separate operations, which limits granularity and may cause agent confusion.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers