GLM Vision Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@GLM Vision Serveranalyze this screenshot and explain what the error message says"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Server GLM Vision
A Model Context Protocol (MCP) server that integrates GLM-4.5V from Z.AI with Claude Code.
Features
Image Analysis: Analyze images using GLM-4.5V's vision capabilities
Local File Support: Analyze local image files or URLs
Configurable: Easy setup with environment variables
Related MCP server: glm-vision-mcp-server
Installation
Prerequisites
Python 3.10 or higher
GLM API key from Z.AI
Claude Code installed
Setup
Clone or create the project directory:
cd /path/to/your/projectCreate and activate virtual environment:
python3 -m venv env source env/bin/activate # On Windows: env\Scripts\activateInstall dependencies:
pip install -r requirements.txt # or with uv (recommended) uv pip install -r requirements.txtSet up environment variables:
cp .env.example .env # Edit .env with your GLM API key from Z.AIAdd the server to Claude Code:
# Using uv (recommended) uv run mcp install -e . --name "GLM Vision Server" # Or manually add to Claude Desktop configuration: claude mcp add-json --scope user glm-vision '{ "type": "stdio", "command": "/path/to/your/project/env/bin/python", "args": ["/path/to/your/project/glm-vision.py"], "env": {"GLM_API_KEY": "your_api_key_here"} }'
Configuration
Set these environment variables in your .env file:
Variable | Description | Default |
| Your GLM API key from Z.AI | (required) |
| GLM API base URL |
|
| Model name to use |
|
Usage
Available Tools
glm-vision
Analyze an image file using GLM-4.5V's vision capabilities. Supports both local files and URLs.
Parameters:
image_path(required): Local file path or URL of the image to analyzeprompt(required): What to ask about the imagetemperature(optional): Response randomness (0.0-1.0, default: 0.7)thinking(optional): Enable thinking mode to see model's reasoning process (default: false)max_tokens(optional): Maximum tokens in response (max 64K, default: 2048)
Example:
Use the glm-vison tool with:
- image_path: "/path/to/your/image.jpg"
- prompt: "Describe what you see in this image"Testing
Test the server using the MCP Inspector:
# With uv
uv run python glm-vision.py
# Or with python
python glm-vision.pyDevelopment
Running Tests
# Install development dependencies
pip install -e ".[dev]"
# Run tests
pytest
# Format code
black .
isort .
# Type checking
mypy glm-vision.pyTroubleshooting
API Key Issues: Make sure your
GLM_API_KEYis correctly set in the environmentConnection Problems: Check your internet connection and API endpoint
Model Errors: Verify that the model name (
GLM_MODEL) is correct and available
License
MIT License - see LICENSE file for details.
Contributing
Fork the repository
Create a feature branch
Make your changes
Add tests if applicable
Submit a pull request
Support
For issues related to the GLM API, contact Z.AI support. For MCP server issues, please create an issue in the repository.
Available Tools
1 toolglm_visionC
Analyze an image file using GLM-4.5V's vision capabilities. Supports both local files and URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| image_path | Yes | ||
| prompt | Yes | ||
| temperature | No | ||
| thinking | No | ||
| max_tokens | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must reveal behavioral traits. Only states basic capability; omits details like permissions, network usage, latency, or size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with purpose. No wasted words, but slightly too brief given parameter count.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description lacks essential context for a 5-parameter vision tool. Agent needs parameter semantics, usage scenarios, and constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. Description does not explain any of the 5 parameters (image_path, prompt, temperature, thinking, max_tokens). Agent cannot infer how to format inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear statement of tool purpose: analyze images using GLM-4.5V vision capabilities, supporting local files and URLs. No sibling tools exist, so no differentiation needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when or when not to use this tool. Lacks context about alternative tools or edge cases (e.g., unsupported image formats).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only one tool, there is no possibility of confusion between tools, making disambiguation perfect.
A single tool trivially follows a consistent naming pattern; consistency is not a concern here.
The single tool is appropriate for a focused vision analysis server, though it borders on being too minimal for broader usage.
The tool covers image analysis but lacks supporting tools (e.g., listing models, checking capabilities), leaving potential gaps for agents.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Analyze images from multiple angles to extract detailed insights or quick summaries. Describe visu…
Analyze images and videos with Gemini to get fast, reliable visual insights. Handle content from U…
Create images and videos from prompts, with options for image mixing, reference images, and start/…
Generate images, videos, voiceovers, and captions from a chat prompt.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables image analysis and understanding using Vision Language Models through OpenAI-compatible APIs. Supports analyzing images from URLs or local files with custom prompts.12MIT
- AlicenseAqualityCmaintenanceEnables Claude Code to analyze images using Zhipu AI's GLM-5V-Turbo vision model, supporting local files and URLs with customizable prompts.162MIT
- AlicenseAqualityCmaintenanceEnables image analysis using any OpenAI-compatible vision API, supporting URLs, local files, or base64 input with custom prompts.1MIT
- FlicenseNot gradedqualityBmaintenanceEnables image analysis using GLM-4V multimodal model, supporting local files and base64 images with optional custom prompts.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/danilofalcao/mcp-server-glm-vision'
If you have feedback or need assistance with the MCP directory API, please join our Discord server