visionMCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MODEL | No | Model name for the vision provider (overrides model in config.json; empty = provider default) | |
| OPENAI_API_KEY | No | OpenAI API key, used when the provider is 'openai' and the api_key field in config.json is empty | |
| VISION_API_KEY | No | API key for the vision provider (overrides api_key in config.json and provider-specific env vars) | |
| VISION_API_URL | No | API URL for the vision provider (overrides api_url in config.json; empty = provider default) | |
| VISION_PROVIDER | No | Vision provider: 'ollama', 'openai', or 'anthropic' (overrides provider in config.json) | ollama |
| ANTHROPIC_API_KEY | No | Anthropic API key, used when the provider is 'anthropic' and the api_key field in config.json is empty |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| logging | {} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| extensions | {
"io.modelcontextprotocol/ui": {}
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| describe_imageA | Describe an image. |
| ask_about_imageC | Answer a specific question about an image. |
| extract_textB | Read all text in an image (OCR). |
| compare_imagesB | Compare two images. Optional |
| server_statusA | Show the active provider, model, and transport. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool targets a clearly distinct task: general description, specific Q&A, OCR, image comparison, and server status. An agent can easily select the right one based on the user's intent, with no meaningful overlap in scope.
Four tools follow a verb_noun pattern (describe_image, ask_about_image, extract_text, compare_images) in snake_case. However, server_status deviates as a noun_noun, and compare_images uses plural while others are singular, slightly breaking the pattern.
With only 5 tools, the server is well-scoped and each tool serves a necessary function. This is within the ideal range of 3-15 for a focused vision understanding server.
Core vision understanding workflows are covered: describing, asking questions, OCR, comparison, and status checks. Minor missing features like explicit object detection or metadata retrieval are easily approximated via ask_about_image, so no critical dead ends.