Qwen Vision MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Qwen Vision MCP ServerWhat's in this photo? https://example.com/photo.jpg"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Qwen Vision MCP Server
MCP server for Alibaba Qwen3.7-plus (multimodal vision model), connected via DashScope's Anthropic-compatible endpoint.
Features
Single tool:
qwen_vision_understand— analyze an image with Qwen3.7-plusSupports local image files (png/jpg/jpeg/gif/webp/bmp) and remote URLs
Same Anthropic Messages API format as Claude
Token usage reported back in each response
Related MCP server: Image Parse MCP
Requirements
Node.js >= 18
A DashScope API key — get one at https://dashscope.console.aliyun.com/apiKey
Install
cd qwen-vision-mcp-server
npm installConfiguration
Environment variables:
Variable | Required | Default | Description |
| ✅ | — | Your DashScope API key |
|
| Model name (e.g. | |
|
| Override endpoint | |
|
| Default max output tokens |
Run locally
DASHSCOPE_API_KEY=sk-xxx npm startCC-Switch / Claude Code config
{
"mcpServers": {
"qwen-vision": {
"type": "stdio",
"command": "cmd",
"args": ["/c", "npx", "-y", "qwen-vision-mcp-server"],
"env": {
"DASHSCOPE_API_KEY": "sk-your-dashscope-key",
"QWEN_MODEL": "qwen3.7-plus"
}
}
}
}Or if installed locally:
{
"mcpServers": {
"qwen-vision": {
"type": "stdio",
"command": "node",
"args": ["E:/Projects/Claude/MCP/qwen/qwen-vision-mcp-server/src/index.js"],
"env": {
"DASHSCOPE_API_KEY": "sk-your-dashscope-key"
}
}
}
}Tool: qwen_vision_understand
Parameter | Type | Required | Description |
| string | ✅ | Local file path or HTTP(S) URL |
| string | ✅ | What to ask about the image |
| number | Max output tokens (default 4096) | |
| number | 0-2, default 1 |
Example:
Use qwen_vision_understand to analyze this UI mockup:
/Users/me/Downloads/login-page.png
"Recreate this as HTML with Tailwind CSS"License
MIT
Available Tools
1 toolqwen_vision_understandA
Analyze an image using Alibaba Qwen3.7-plus (multimodal vision model). Supports local image files and remote URLs. Reaches DashScope via the Anthropic-compatible endpoint (https://dashscope.aliyuncs.com/apps/anthropic). Excels at: UI screenshot→code, design mockup analysis, visual debugging, chart/document understanding.
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | Image source: local file path (e.g. C:/path/to/screenshot.png) or URL (https://...) | |
| prompt | Yes | What to ask about the image. Be specific for best results. E.g.: 'Recreate this UI as HTML with Tailwind CSS' | |
| max_tokens | No | Maximum output tokens | |
| temperature | No | Sampling temperature (0-2) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses the model, endpoint, and supported image sources, but lacks details on output format, latency, cost, or potential failure modes. A read from the agent perspective would want to know what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus a compact bullet-like list. Front-loaded with the core purpose, each sentence adds unique value: model, supported sources, endpoint, and use cases. Zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, description is the sole source of context. It covers input and use cases adequately, but omits output format, error behavior, and performance characteristics. Adequate for a straightforward vision analysis tool, but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so schema already documents all parameters. Description adds examples and context (e.g., 'be specific for best results') but does not significantly enhance meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it analyzes images using a specific model (Qwen3.7-plus), with a list of concrete use cases (UI to code, design mockup analysis, etc.). Verb+resource combination is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicit guidance through the list of use cases (e.g., 'UI screenshot→code'), but no explicit when-to-use, when-not-to-use, or alternatives. Since there are no sibling tools, the miss is less severe, but still room for improvement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
qwen_vision_understand
TDQS
Scored across 1 tool
Only one tool exists, so there is no possibility of confusion or overlap.
The single tool follows a clear verb_noun pattern (qwen_vision_understand), consistent with common MCP naming conventions.
A vision MCP server typically requires multiple specialized tools (e.g., describe, analyze, compare) rather than a single monolithic tool. One tool feels under-scoped for the domain.
The single tool covers a broad range of vision tasks (UI analysis, charts, documents), but lacks structured decomposition into separate operations, which limits granularity and may cause agent confusion.
Maintenance
Related MCP Connectors
Analyze images and videos with Gemini to get fast, reliable visual insights. Handle content from U…
Generate images with your own ChatGPT subscription (Plus, Pro or Team), without spending API credits
Analyze images from multiple angles to extract detailed insights or quick summaries. Describe visu…
- lightgenOAuthapp.lightgen
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables image analysis and understanding using Vision Language Models through OpenAI-compatible APIs. Supports analyzing images from URLs or local files with custom prompts.12MIT
- AlicenseAqualityDmaintenanceEnables image analysis using any OpenAI-compatible vision API, supporting URLs, local files, or base64 input with custom prompts.1MIT
- FlicenseNot gradedqualityCmaintenanceEnables Claude Code to analyze images using Qwen vision models (via DashScope) when the main model is text-only.-
- AlicenseNot gradedqualityCmaintenanceEnables text-only models like Claude Code to recognize images by calling Qwen vision models from DashScope, returning text descriptions for continued reasoning.MIT