screenshot_annotate
Captures a screenshot with numbered bounding-box overlays, enabling vision-language models to reference specific elements spatially.
Instructions
Screenshot with @eN bounding-box overlays (vision-LLM friendly).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Output PNG path. | |
| lease | No | Optional lease token to present if the target session is leased (0.7.0). Threaded per-call; never read from the server's env. | |
| session | No | Optional session name to target (omit for the shared 'default'). On a daemon shared with other agents, pass a UNIQUE name for stateful multi-step work (go→click→fill) so you don't collide on 'default'. | |
| full_page | No | Capture full page. |