Sube y analiza un archivo mediante EnriProxy (extracción del lado servidor + análisis con modelo).
/ Upload and analyze a media file via EnriProxy (server-side extraction + model analysis).
IMPORTANTE para modelos sin visión: la respuesta es SIEMPRE TEXTO (descripción visual generada del lado del servidor y/o transcripción del audio); nunca se devuelven bloques de imagen, así que CUALQUIER modelo puede consumirla — si no puedes ver imágenes, esta herramienta es tu vía para 'ver' archivos multimedia pidiendo la descripción en question.
/ IMPORTANT for models without vision: the response is ALWAYS TEXT (server-side visual description and/or audio transcription); image blocks are never returned, so ANY model can consume it — if you cannot see images, this tool is how you 'see' media by asking for the description in question.
Formatos aceptados — Imágenes estáticas: PNG, JPEG, WebP, TIFF, BMP, AVIF, HEIC/HEIF, SVG. Imágenes animadas: GIF animado, WebP animado, APNG, SVG animado (se extraen fotogramas clave y se describen sus cambios). Video: MP4, WebM, MKV, MOV, AVI y cualquier contenedor/códec decodificable, CON o SIN audio (en video con audio se procesan juntos: fotogramas + transcripción sobre la misma línea de tiempo). Audio: WAV, MP3, OGG (Vorbis/Opus), M4A/AAC, FLAC, ALAC, Opus y cualquier formato común (todo se normaliza a WAV 16 kHz mono antes de transcribir). Documentos: PDF (de texto, escaneado o mixto), Word (.docx), Excel (.xlsx), PowerPoint (.pptx) — con texto e imágenes embebidas — y JSONL. Conjuntos de varias imágenes: use paths.
/ Accepted formats — Static images: PNG, JPEG, WebP, TIFF, BMP, AVIF, HEIC/HEIF, SVG. Animated images: animated GIF, animated WebP, APNG, animated SVG (key frames are extracted and their changes described). Video: MP4, WebM, MKV, MOV, AVI and any decodable container/codec, WITH or WITHOUT audio (video with audio processes both together: frames + transcription on the same timeline). Audio: WAV, MP3, OGG (Vorbis/Opus), M4A/AAC, FLAC, ALAC, Opus and any common format (everything is normalized to 16 kHz mono WAV before transcription). Documents: PDF (text, scanned, or mixed), Word (.docx), Excel (.xlsx), PowerPoint (.pptx) — with embedded text and images — and JSONL. Multiple-image sets: use paths.
Cuándo usarla: PDFs grandes o escaneados donde el Read puede truncar; video/audio u otros binarios que el cliente no puede leer; HEIC/AVIF/TIFF/BMP/APNG/SVG/Office cuando el Read no es confiable; archivos muy grandes con subidas reanudables (hasta 4 GiB).
/ When to use: large or scanned PDFs where client Read may truncate; video/audio or binary media the client cannot Read; HEIC/AVIF/TIFF/BMP/APNG/SVG/Office docs when client Read is unreliable; very large files needing resumable uploads (up to 4 GiB).
Reglas: use path para un archivo, paths para varias imágenes. Cuando paths trae al menos una entrada válida, path se ignora (contrato explícito: mandar ambos se permite, path se ignora en silencio — prefiera semántica oneOf y mande solo uno). Las entradas en blanco se descartan; claves desconocidas en video/audio/document/images se rechazan. question es opcional aquí (obligatoria en EnriCode); attachmentIndex/attachmentId no existen aquí (sólo EnriCode).
/ Rules: use path for one file, paths for several images (UI screenshots/photo sets). When paths carries at least one valid entry, path is ignored (explicit ignore-path contract: sending both is allowed, path is silently ignored — prefer oneOf semantics and send only one). Blank paths entries are discarded; unknown keys inside video/audio/document/images are rejected (check typos like max_pages_totall). question is optional here (required in EnriCode vision.analyze_media); attachmentIndex/attachmentId do not exist here (EnriCode-only).
Presupuestos (timeout = min(ENRIVISION_TIMEOUT_MS del operador, presupuesto del modo)): single = 10 min (una pasada, rápida y barata); multipass = 20 min (por segmentos/lotes + reducción); auto = 20 min (el servidor elige y puede escalar a multipass). Si no sabe cuál usar, omita el afinado (auto).
/ Analysis budgets (client analyze timeout = min(operator ENRIVISION_TIMEOUT_MS, mode budget)): single = 10 min (one pass, fast and cheap, 1 image or simple questions); multipass = 20 min (per-segment/batch map + reduce; PDFs over ~20 pages, long videos, image sets); auto = 20 min (the server picks and may escalate to multipass). If unsure, omit tuning (auto).
Clip de video: para preguntas en un tiempo específico ("¿qué pasa en 12:34?") use video.clip_start_seconds + video.clip_duration_seconds: convierta a segundos (12:34 = 1260+34 = 754), por ejemplo clip_start_seconds=754 y clip_duration_seconds=30. O dé video.clip_end_seconds (fin = inicio + duración, 0-86400 s).
/ Video clip targeting: for time-specific questions ("what happens at 12:34?") use video.clip_start_seconds + video.clip_duration_seconds: convert to seconds (12:34 = 1260+34 = 754) and request a window, e.g. clip_start_seconds=754 and clip_duration_seconds=30. Or give video.clip_end_seconds instead (end = start + duration, 0-86400 s).
Enteros estrictos: los knobs enteros aceptan números o strings enteras completas ("60" vale; "8.0", "8abc" y 7.9 fallan). Los flotantes aceptan decimales ("12.5" vale). transcribe vale true por defecto y no tiene efecto en imágenes/documentos (se declara en warnings, se ignora). Requiere API key válida de EnriProxy (env ENRIPROXY_API_KEY).
/ Strict integers: integer knobs accept numbers or complete integer strings ("60" works; "8.0", "8abc", 7.9 fail). Floats accept decimals ("12.5" works). transcribe defaults to true and has no effect on images/documents (declared in warnings, ignored). Requires a valid EnriProxy API key (env ENRIPROXY_API_KEY, sent as Authorization: Bearer ...).
Errores: las fallas devuelven isError con texto bilingüe más structuredContent {code, retryable, httpStatus?} con el vocabulario EnriCode; retryable marca 429/5xx/timeouts. Si el mensaje trae Detalle del servidor: en el idioma del proxy, repórtelo tal cual. Fotogramas y transcripción comparten la MISMA línea de tiempo. model es el id del modelo para afinidad de dispatch (máximo 128 caracteres o env ENRIVISION_MODEL; omita para auto-dispatch). language controla el idioma de la RESPUESTA; transcription_language aparte el idioma que Whisper espera al TRANSCRIBIR ("auto" = detectar solo).
/ Errors: failures return isError with bilingual text plus structuredContent {code, retryable, httpStatus?} reusing the EnriCode vocabulary (ENRICODE_ERR_TOOL_INPUT_INVALID / EXECUTION_FAILED / EXECUTION_TIMEOUT / EXECUTION_ABORTED); retryable marks 429/5xx/timeouts. If the message carries a Detalle del servidor: fragment in the proxy language, report it verbatim. Video frames and transcription share the SAME video timeline. Animated GIF/WebP/APNG/SVG become representative key frames. model is the active model id for server-side dispatch affinity (max 128 chars, or env ENRIVISION_MODEL; omit for auto-dispatch). Set language (e.g. "es") to match the user language and avoid drift: language controls the analysis RESPONSE language; transcription_language separately controls the language Whisper expects when TRANSCRIBING audio ("auto" = detect only).
Ejemplos mínimos: (1) una imagen: {"path": "/tmp/foto.png", "question": "..."}. (2) clip de video 12:34->754s: {"path": "/tmp/charla.mp4", "question": "...", "video": {"clip_start_seconds": 754, "clip_duration_seconds": 30}}. (3) PDF largo multipass: {"path": "/tmp/manual.pdf", "question": "...", "analysis_mode": "multipass"}. Rutas absolutas del host MCP (en Windows valen C:/...; en POSIX lanzarían error).
/ Minimal examples: (1) single image: {"path": "/tmp/shot.png", "question": "What does each capture show?"}. (2) video clip 12:34->754s: {"path": "/tmp/talk.mp4", "question": "What happens at 12:34?", "video": {"clip_start_seconds": 754, "clip_duration_seconds": 30}}. (3) long PDF multipass: {"path": "/tmp/manual.pdf", "question": "Summarize each chapter.", "analysis_mode": "multipass"}. Absolute MCP-host paths (C:/... drive paths only work on a Windows host; POSIX rejects them).
Continuación: si la respuesta trae has_more con cursor (segment_summaries_cursor o transcription_segments_cursor), pida el resto con solo cursor (+ offset opcional, por defecto next_offset; también limit opcional 1-100 para acotar la ventana). Con cursor no mande path/paths. / Continuation: when the response carries has_more with a cursor (segment_summaries_cursor or transcription_segments_cursor), ask for the rest with only cursor (+ optional offset, defaults to next_offset; optional limit 1-100 bounds the window); never send path/paths with cursor.
Depuración de capturas de UI: abra con un veredicto de una línea; describa zona por zona; aproxime colores como hex; cuantifique defectos de layout; transcriba etiquetas, botones y errores visibles; compare observado vs esperado cuando aplique. / UI-screenshot debugging (when the media are app screenshots): open with a one-line plain verdict; describe zone by zone (header, sidebar, main content, modals, notifications), not as a general scene; approximate colors as hex values (e.g. #1F6FEB) and name them; quantify layout defects (overflows, clipping, overlaps, misalignments, missing spacing, cut text) estimating pixel magnitudes when possible; transcribe labels, buttons, and any visible error/status text; when the request states what was expected, compare observed vs expected explicitly.