vision-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| MODEL | No | Alternative to VISION_MODEL. | |
| API_KEY | No | Alternative to OPENAI_API_KEY. | |
| API_BASE_URL | No | Alternative to OPENAI_BASE_URL. | |
| VISION_MODEL | No | Model name to use. Defaults to gpt-4o. | gpt-4o |
| OPENAI_API_KEY | No | API key for the vision provider. Required for most providers. | |
| OPENAI_BASE_URL | No | Custom API endpoint URL. Defaults to https://api.openai.com/v1. | https://api.openai.com/v1 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
| resources | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageA | Analyze an image using vision AI. Provide either a URL or base64-encoded image data. Returns detailed analysis including objects, text, colors, scene description, and more. |
| compare_imagesA | Compare two or more images and describe their differences, similarities, or relationships. |
| extract_textA | Extract and transcribe text from an image (OCR). Useful for documents, screenshots, signs, etc. |
| describe_sceneB | Get a detailed description of a scene, including spatial relationships, atmosphere, and context. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
The tools have distinct purposes, but analyze_image is a catch-all that includes text extraction and scene description, overlapping with extract_text and describe_scene. An agent may struggle to choose between the broad tool and the specialized ones.
All tools follow a consistent verb_noun pattern with lowercase and underscores: analyze_image, compare_images, extract_text, describe_scene. The naming is uniform and predictable.
Four tools is well-scoped for a vision analysis server, providing a focused set of capabilities without unnecessary bloat or thinness.
The tools cover core vision tasks: analysis, comparison, OCR, and scene description. Some potential gaps include dedicated object detection or face recognition, but analyze_image likely covers these generically, so the surface is nearly complete.