vision-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MODEL | No | Vision model name | gpt-4o |
| VISION_API_KEY | Yes | API key for the gateway | |
| VISION_BASE_URL | No | OpenAI-compatible API base URL | https://api.openai.com/v1 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageA | Analyze an image with a vision LLM (OpenAI-compatible chat/completions). Provide exactly one of |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 1 tool
Only one tool exists, so there is no ambiguity between tools. An agent cannot misselect among tools.
With a single tool, naming consistency is not applicable but there is no inconsistency to flag.
A single tool covering all vision analysis tasks feels too thin for the scope. While the tool is versatile via prompts, it lacks separate endpoints for different operations, making the surface sparse.
The tool covers the core vision analysis functionality with options for local/remote/base64 input and custom prompts. However, there are no tools for managing image resources or handling results, leaving minor gaps for complex workflows.