vision_find
Locate a UI element by description and return its coordinates and confidence score without clicking. Helps agents identify elements for subsequent actions.
Instructions
Return coords + confidence for a described element (no click).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| lease | No | Optional lease token to present if the target session is leased (0.7.0). Threaded per-call; never read from the server's env. | |
| intent | Yes | Description. | |
| session | No | Optional session name to target (omit for the shared 'default'). On a daemon shared with other agents, pass a UNIQUE name for stateful multi-step work (go→click→fill) so you don't collide on 'default'. | |
| min_confidence | No | Minimum confidence. |