vision_kit
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| describe_imageA | 通用视觉描述:调用视觉模型观察图片并返回中文描述文本。 Args: image_path: 本地图片文件路径(支持 jpg/png/bmp 等,超长图自动分块)。 prompt: 可选的自定义描述要求;缺省使用通用描述提示词。 |
| describe_image_structuredA | 结构化识别:提取图片中带数字的标注(向量/矩阵/坐标/角度/未知量)。 Args: image_path: 本地图片文件路径。 Returns: dict,包含 type(图形类型)、vectors(向量)、matrices(矩阵)、 note(说明)与 text(渲染文本)。识别失败时返回 {"error": "..."}(避免返回 None 与声明类型 dict 不符)。 |
| describe_image_statsA | 统计图 → 数据表:提取柱状/折线/饼图中的类别与数值序列(4.3)。 Args: image_path: 本地图片文件路径。 Returns: dict,包含 type(图表类型)、categories(类别)、series(系列数值)、 note(说明)与 checks(自洽校验:长度对齐 / 百分比求和 / 非负)。 识别失败时返回 {"error": "..."}。 |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
describe_image is a generic catch-all, while the other two tools target specific image types with distinct output schemas. The overlap with the generic tool is acceptable because the specialized descriptions clearly define when to use them.
All tools share the describe_image_ prefix followed by a clear modifier: structured and stats. This creates a predictable and consistent naming convention across the entire server.
Three tools is within the well-scoped range, and each tool serves a distinct, meaningful vision task without redundancy. The small count feels deliberate rather than incomplete.
The set covers generic visual description, structured diagram extraction, and statistical chart parsing, which are the main advertised capabilities. Some advanced vision tasks like object detection or OCR are absent, but the generic describe_image tool can work around many gaps via custom prompts.