agent-vision
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OPENAI_MODEL | No | Compatibility fallback for AGENT_VISION_MODEL. | |
| VISION_MODEL | No | Compatibility fallback for AGENT_VISION_MODEL. | |
| OPENAI_API_KEY | No | Compatibility fallback for AGENT_VISION_API_KEY. | |
| VISION_API_KEY | No | Compatibility fallback for AGENT_VISION_API_KEY. | |
| OPENAI_BASE_URL | No | Compatibility fallback for AGENT_VISION_BASE_URL. | |
| VISION_BASE_URL | No | Compatibility fallback for AGENT_VISION_BASE_URL. | |
| AGENT_VISION_MODEL | No | The model name to use for vision analysis. | |
| AGENT_VISION_API_KEY | No | Your API key for the vision API. Optional if the endpoint has no authentication. | |
| AGENT_VISION_HEADERS | No | Additional request headers as a JSON string. | {} |
| AGENT_VISION_BASE_URL | No | The base URL of your OpenAI-compatible vision API endpoint. | |
| AGENT_VISION_MAX_TOKENS | No | Maximum output tokens. | 4096 |
| AGENT_VISION_TIMEOUT_MS | No | Download and API timeout in milliseconds. | 120000 |
| AGENT_VISION_FFMPEG_PATH | No | Path to ffmpeg executable. If not set, the server uses PATH then the bundled ffmpeg. | |
| AGENT_VISION_MAX_IMAGE_MB | No | Maximum image size in megabytes. | 20 |
| AGENT_VISION_MAX_VIDEO_MB | No | Maximum video size in megabytes. | 200 |
| AGENT_VISION_ALLOW_PRIVATE_URLS | No | Whether to allow private network URLs for images/videos. | false |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageA | Analyze an image from a local path, HTTP(S) URL, or base64 data URI. Use this for screenshots, UI, OCR, diagrams, photos, and visual debugging. |
| analyze_videoA | Analyze a local or remote video by extracting evenly spaced keyframes. Use this for screen recordings, UI flows, demos, and event summaries. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools are perfectly distinct: analyze_image handles static images, while analyze_video handles moving media via keyframes. There is no overlap in input types or use cases, so an agent can easily choose the correct tool.
Both tools follow the exact same verb_noun pattern: analyze_image and analyze_video. This consistent naming makes the tool set predictable and easy to navigate.
With only two tools, the server is minimal but well-scoped for its stated purpose of visual media analysis. While it's on the low end, the narrow domain justifies the count, and each tool covers a major media type.
The server covers the two essential types of visual input—images and videos. Both tools are generic enough to handle a wide range of analysis tasks (screenshots, OCR, UI flows, etc.), leaving no obvious gaps in the covered domain.