Skip to main content
Glama
Byte-Naut

npu-vision-fallback

by Byte-Naut
README.md
<div align="center">

# πŸ”‹ npu-vision-fallback

**⚠️ Archived experiment β€” no longer actively developed**

An Intel NPU screen-vision + MCP integration experiment. Backend routing and
Windows-native OCR integration worked; the NPU-first product hypothesis did
not find a stable user need. Code preserved for reference.

[![PyPI](https://img.shields.io/pypi/v/npu-vision-fallback.svg)](https://pypi.org/project/npu-vision-fallback/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue.svg)](https://www.python.org/downloads/)

**English** | [δΈ­ζ–‡](README.zh-CN.md)

</div>

---

## πŸ—„οΈ Why This Is Archived

This project tested whether desktop AI agents would benefit from running
low-level screen perception β€” "find the button", "read this region" β€”
**locally on an Intel NPU**, exposed as an MCP server. That was a hypothesis,
not a verified need; it never crystallized into stable demand, so development
stopped.

Two things the experiment did prove:

- **Backend routing with graceful fallback** β€” NPU β†’ CPU β†’ system OCR β†’
  cross-platform OCR, driven by a power policy, with heavy dependencies kept
  optional.
- **Windows-native OCR over MCP** β€” WinRT OCR wired into the same tool surface
  as the OpenVINO NPU UI detector.

**What I still use it for:** batch OCR of long screenshots. Note that the
implementation has **no complete tiling guarantee for very long images** β€” it
may drop content on extreme captures. Treat the code as a reference, not a
maintained dependency.

---

## 🧱 Original Design

Three constraints that shaped the project:

1. **Cheapest path first.** OS-native OCR, then a local detector. A cloud
   multimodal model was the *last* resort β€” for reasoning, not for finding
   buttons.
2. **Isolate compute.** Inference on the NPU so the GPU stays free for the
   app the agent is watching. The service requests **NPU or CPU only β€” never
   any GPU**.
3. **Local, period.** Screenshots stay in memory and are **never written to
   disk**; OCR text is never logged; nothing leaves the machine.

**Pipeline:** `mcp_server.py` β†’ `core/orchestrator.py` β†’
`core/backend_selector.py` β†’ selected backend β†’ screen-space remap. Full
detail: [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md).

---

## πŸ“‹ Reference

### MCP Tools

| Tool | Purpose | Key Arguments |
|---|---|---|
| `health_check` | Server status | β€” |
| `list_backends` | Available backends | β€” |
| `ocr_region` | Extract text from region | `region=[x1,y1,x2,y2]` |
| `detect_ui` | Find UI elements | `region=[x1,y1,x2,y2]` |
| `analyze_screen` | Combined OCR + detection | `region=[x1,y1,x2,y2]` |
| `analyze_image_file` | Analyze an image file | `path`, `mode=ocr\|ui\|all` |
| `analyze_image_directory` | Batch-analyze images in a directory | `input_dir`, `output_dir`, `mode`, `recursive`, `overwrite`, `max_workers` |

`analyze_screen` fuses detection + OCR into spatially-sorted elements with
text annotations. `analyze_image_file` and `analyze_image_directory` reuse the
same pipeline on image files.

### CLI

```bash
# Single image
uv run npu-vision-fallback-cli analyze-image path/to/image.png --mode all

# Directory batch
uv run npu-vision-fallback-cli analyze-dir path/to/dir --output-dir out/ --recursive
```

### Installation (reference only β€” not recommended for new deployments)

```bash
pip install "npu-vision-fallback[ocr-win,detect]"
python scripts/download_ui_model.py  # one-time model export to OpenVINO IR
```

Further extras are listed in `pyproject.toml`.

### Supported Backends

| Backend | Type | Device | Platform | Status |
|---|---|---|---|---|
| `winocr` | System OCR | CPU/NPU | Windows | βœ… Primary |
| `openvino_npu` | UI Detection | NPU | Win/Linux + Intel NPU | βœ… Primary |
| `openvino_cpu` | UI Detection | CPU | Win/Linux/macOS | βœ… Fallback |
| `rapid_ocr` | OCR | CPU | All | βœ… Cross-platform |
| `pytesseract` | OCR | CPU | All | βœ… Last-resort |
| `vision` | System OCR | ANE | macOS | 🚧 Never implemented |

### Measured Performance

**Intel Core Ultra 9 275HX**, 2560Γ—1600, on battery:

| Task | Backend | Latency | Energy | Notes |
|---|---|---|---|---|
| **OCR** | WinOCR | ~1100ms | 2.5J | Native Windows API |
| **OCR** | RapidOCR | ~6300ms | 14.5J | Cross-platform ONNX CPU |
| **UI Detection** | OpenVINO NPU | ~80ms | 0.3J | YOLOv8n on Intel AI Boost |
| **UI Detection** | OpenVINO CPU | ~120ms | β€” | No-NPU fallback |

Full details: [`outputs/power_report.md`](outputs/power_report.md)

### Examples

| Example | Description |
|---|---|
| [`basic_ocr.py`](examples/basic_ocr.py) | OCR a screen region |
| [`agent_ui_navigation.py`](examples/agent_ui_navigation.py) | Find and click UI elements |
| [`desktop_remote_vnc.py`](examples/desktop_remote_vnc.py) | Vision fallback in remote desktop |

```bash
uv run python examples/basic_ocr.py --region 0 0 1280 800
```

### Docs

- [Architecture](docs/ARCHITECTURE.md) Β· [Backends](docs/BACKENDS.md) Β· [FAQ](docs/FAQ.md)
- [Changelog](docs/CHANGELOG.md) Β· [Code Guide](CLAUDE.md)

---

## πŸ“„ License

[MIT](LICENSE) Β© npu-vision-fallback contributors

### πŸ™ Acknowledgments

Built with [MCP](https://modelcontextprotocol.io/) (Anthropic),
[OpenVINO](https://github.com/openvinotoolkit/openvino),
[Ultralytics YOLO](https://github.com/ultralytics/ultralytics),
[RapidOCR](https://github.com/RapidAI/RapidOCR),
[Tesseract](https://github.com/tesseract-ocr/tesseract), and
[python-mss](https://github.com/BoboTiG/python-mss).

Development assisted by [Claude Code](https://claude.ai/claude-code) (Anthropic).

TDQS

A4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a distinct purpose: analyze_screen combines detection and OCR, detect_ui does detection only, ocr_region does OCR only, while health_check and list_backends are utility tools. No overlap or ambiguity.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern (analyze_screen, detect_ui, ocr_region, list_backends, health_check), making them predictable and easy to understand.

Tool Count5/5

Five tools is well-scoped for a vision fallback server: core detection, OCR, combined analysis, health check, and backend listing. Each tool earns its place without being overwhelming or insufficient.

Completeness4/5

The tool surface covers the main vision operations (detection, OCR, combined) plus utility. A minor gap might be adjustable OCR language or detection parameters, but overall the set is functional and avoids dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues