Locate element (fusion)
computer_locateLocate GUI elements by natural-language instruction or structured descriptor, using platform APIs first and vision fallback, then return screenshot-pixel coordinates for computer actions.
Instructions
Find an element by natural language or structured descriptor. Structured (UIA/AX/AT-SPI/DOM) first, UI-Venus vision grounding as fallback; vision points are snapped back to structured elements when possible. Coordinates in results are screenshot pixel space; use with computer_action.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Unified device target (spec §14). Defaults to this machine with auto-detected platform. | |
| descriptor | No | ||
| instruction | Yes | Natural language element description, e.g. "关闭按钮" or "the login button" | |
| preferStructured | No |