Perceive Screen
perceive_screenCaptures annotated screenshots with element tap points and flags, enabling AI agents to perceive and interact with mobile device screens.
Instructions
LOOK at the screen: an annotated screenshot plus every element and its tap point.
e is an array whose INDEX is the som_id (first entry = som_id 1): [center_x, center_y,
name, flags] (name/flags omitted when empty). elements is the same list as text.
Box colours: BLUE tappable, GREEN text input (type_text), MAGENTA scrollable, AMBER
toggle, GREY nothing declared, RED on-host vision (detector/OCR; a good guess, not a fact).
Flags (only when true): e editable, c checked, o unchecked, d disabled, f focused,
l long-pressable, ? low-confidence vision box, w scroll host (aim inside it).
offscreen: text that exists but is not on screen (cannot be tapped).
detail="full" adds the visual pass (OmniParser YOLOv8 icon detector + OCR) for icons the
tree does not describe; perception_tier reports tree_only or full. description is
logged only. ids go stale after any action.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| ocr | No | auto | |
| lang | No | eng | |
| detail | No | ||
| device | No | ||
| max_marks | No | ||
| description | No | ||
| include_image | No |