Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
RUN_VISION_MODELNoVision model name.qwen3-vl-flash
RUN_VISION_API_KEYYesAPI key for the configured vision provider.
RUN_VISION_BASE_URLNoBase URL without /chat/completions.DashScope compatible endpoint
RUN_VISION_MAX_TOKENSNoMax response tokens.2048
RUN_VISION_TIMEOUT_MSNoUpstream API timeout.60000
RUN_VISION_ALLOWED_DIRSNoComma-separated directories local image_path may read from.unrestricted
RUN_VISION_MAX_IMAGE_BYTESNoMax local image size in bytes.20971520

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
describe_imageA

Image-understanding tool for non-multimodal (text-only) models. Uses an external vision model to read text (OCR), describe scenes, interpret charts and diagrams, compare images, and return the results as text. Models that can understand images directly must not call this tool; use their own vision capability instead. Accepts image_path, image_url, image_base64, image_ref, or images[]. For faster, more useful answers, ask the specific question you need answered (e.g. "read the error text", "what does this chart show") instead of an open-ended "describe everything".

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.2/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no risk of an agent confusing it with another server tool. The description also clearly scopes who should use it, removing ambiguity about when it should be called.

Naming Consistency5/5

The single tool name follows a clear verb_noun convention (describe_image), so there is no mixed naming style or inconsistency. A larger set would provide more evidence of a pattern, but nothing here violates a predictable convention.

Tool Count3/5

One tool feels thin for a server named VisionPower, even though the single tool is substantial and consolidates several vision tasks. The count is borderline rather than clearly inappropriate.

Completeness5/5

The tool covers the major image-understanding needs: OCR, scene description, chart/diagram interpretation, and image comparison, with multiple input formats supported. For its stated purpose of serving vision to text-only models, there are no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessWithin a week