MCP Vision Relay
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| QWEN_CLI_COMMAND | No | Path to the Qwen CLI executable | |
| GEMINI_CLI_COMMAND | No | Path to the Gemini CLI executable | |
| MCP_IMAGE_TEMP_DIR | No | Directory for storing downloaded/decoded temporary image files | |
| QWEN_DEFAULT_MODEL | No | Default model name for Qwen (e.g., qwen2.5-omni-medium) | |
| MCP_MAX_IMAGE_BYTES | No | Maximum allowed image size in bytes | |
| QWEN_DEFAULT_PROMPT | No | Default prompt for Qwen image analysis | |
| GEMINI_DEFAULT_MODEL | No | Default model name for Gemini (e.g., gemini-2.0-flash) | |
| GEMINI_OUTPUT_FORMAT | No | Controls Gemini output format (text or json) | |
| GEMINI_DEFAULT_PROMPT | No | Default prompt for Gemini image analysis | |
| MCP_COMMAND_TIMEOUT_MS | No | Global timeout in milliseconds for CLI commands | |
| MCP_ALLOWED_IMAGE_EXTENSIONS | No | List of allowed image file extensions |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| gemini_analyze_imageC | Use Google Gemini CLI to describe or analyze an image using multimodal capabilities. |
| qwen_analyze_imageC | Use Qwen CLI to describe or analyze an image with its multimodal capabilities. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools are essentially identical in purpose—both describe or analyze images using multimodal capabilities, differing only in the underlying model (Gemini vs. Qwen). An agent would have no clear basis to choose one over the other based on their descriptions, leading to confusion and misselection.
The tool names follow a perfectly consistent pattern: both use a clear 'model_verb_noun' structure (gemini_analyze_image, qwen_analyze_image). This consistency makes it easy to understand what each tool does at a glance.
With only two tools, the server feels thin for a vision-related domain, as it lacks coverage for common operations like image generation, editing, or filtering. The tools are redundant in functionality, making the count seem artificially low for the apparent scope.
The tool surface is severely incomplete for a vision server; it only offers image analysis via two similar models, with no support for tasks like image creation, transformation, or retrieval. This creates significant gaps that will limit agent capabilities in handling broader vision workflows.