vision
Computer vision: describe or analyze an image with a LOCAL multimodal model (llava). For agents that need to 'see' (describe scenes, read diagrams, classify images). Upload the image via multipart or as 'archivo_b64'; 'input' = the question or instruction abou [x402: 0.008 USDC on Base, pay-per-use]
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | Qué quieres saber de la imagen | |
| archivo_b64 | Yes | Imagen en base64 (o subir 'archivo' por multipart) |