npu-vision-fallback
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| health_checkA | Check server health |
| list_backendsA | List available vision backends |
| ocr_regionB | OCR a screen region. region=[x1,y1,x2,y2] in screen coords; omit for full screen. |
| detect_uiA | Detect objects / UI elements in a screen region using YOLOv8n on OpenVINO (NPU or CPU). Returns bounding boxes with labels and confidence scores. region=[x1,y1,x2,y2] in screen coords; omit for full screen. |
| analyze_screenA | Capture a screen region, run NPU YOLO UI detection and system OCR in parallel, then spatially fuse the results. Returns an ordered list of interactive elements (buttons, fields, headings, …) each annotated with the visible text inside them — ideal for agents that need to understand and act on the current screen. region=[x1,y1,x2,y2] in screen coords; omit for full screen. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool has a distinct purpose: analyze_screen combines detection and OCR, detect_ui does detection only, ocr_region does OCR only, while health_check and list_backends are utility tools. No overlap or ambiguity.
All tool names follow a consistent verb_noun snake_case pattern (analyze_screen, detect_ui, ocr_region, list_backends, health_check), making them predictable and easy to understand.
Five tools is well-scoped for a vision fallback server: core detection, OCR, combined analysis, health check, and backend listing. Each tool earns its place without being overwhelming or insufficient.
The tool surface covers the main vision operations (detection, OCR, combined) plus utility. A minor gap might be adjustable OCR language or detection parameters, but overall the set is functional and avoids dead ends.