Vision Toolkit
by leiming2333
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| QWEN_VL_MODEL | No | Qwen vision model | qwen-vl-plus |
| GEMINI_API_KEY | No | Google AI API key | |
| OPENAI_API_KEY | No | OpenAI API key | |
| OPENAI_BASE_URL | No | OpenAI-compatible endpoint (proxy supported) | https://api.openai.com/v1 |
| QWEN_IMAGE_MODEL | No | Wanxiang image model | wanx2.1-t2i-turbo |
| DASHSCOPE_API_KEY | No | Alibaba DashScope API key | |
| GEMINI_IMAGE_MODEL | No | Imagen image model | imagen-3.0-generate-002 |
| OPENAI_IMAGE_MODEL | No | Image generation model | dall-e-3 |
| GEMINI_VISION_MODEL | No | Gemini vision model | gemini-2.0-flash |
| OPENAI_VISION_MODEL | No | Vision model | gpt-4o |
| QWEN_EMBEDDING_MODEL | No | Multimodal embedding model | multimodal-embedding-one-peace-v1 |
| VISION_TOOLKIT_PYTHON | No | Python interpreter path to use | |
| GEMINI_EMBEDDING_MODEL | No | Text embedding model (for image similarity) | text-embedding-004 |
| OPENAI_EMBEDDING_MODEL | No | Text embedding model (for image similarity) | text-embedding-3-small |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
No tools | |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessUnresponsive