EnriVision
EnriVision is an MCP server that lets any model analyze local or remote media by uploading it to EnriProxy and returning text-based server-side extraction and model analysis.
Analyze images, videos, audio, documents (PDF, DOCX, XLSX, PPTX), animated images, and image sets via the
analyze_mediatool.Upload local files (resumable, up to 4 GiB) or download http(s) URLs (up to 64 MiB each; larger solitary URLs escalate to server-side ingestion).
Ask custom questions, provide context hints, and choose response language (Spanish/English and others).
Target video clips by start/duration/end seconds and control frame counts.
Enable or disable audio transcription, with optional Whisper language hints.
Choose analysis modes:
single(fast/cheap),multipass(segmented/batched for long PDFs, videos, image sets), orauto.Continue truncated results with opaque cursors, offsets, and limits.
Zoom into image regions at native resolution using element boxes returned by previous analyses.
Tune multipass behavior: segments, max pages, batch sizes, image dimensions, and scanned-text thresholds.
Receive structured output:
analysis,media_type,extraction, plus optional element boxes and warnings.Handle errors with bilingual messages and machine-readable
structuredContentcodes for retryable failures.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@EnriVisionsummarize the video at /Users/me/demo.mp4 and transcribe the audio"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
EnriVision
EnriVision is a Model Context Protocol (MCP) server over stdio that uploads local media to EnriProxy and returns server-side extraction + model analysis.
This is useful for media types that many MCP clients cannot read reliably (videos, audio, scanned PDFs, HEIC/AVIF, large files), while keeping the MCP server itself lightweight.
What this project is
An MCP server process your MCP host launches (OpenCode, Claude Code, Codex, etc.)
A thin client for EnriProxy (resumable upload + structured output)
Related MCP server: multimodal-reader-mcp
Requirements
Node.js
>= 24A reachable EnriProxy server with these endpoints enabled:
POST /v1/uploadsHEAD /v1/uploads/:idPATCH /v1/uploads/:idDELETE /v1/uploads/:id(best-effort cleanup of orphaned upload sessions)POST /v1/vision/analyzePOST /v1/vision/segments(cursor continuation for long analyses)GET /v1/account/models(fail-open vision-capability probe before upload)
An EnriProxy API key (configured on the EnriProxy side)
Install
# Global install
npm install -g @bedolla/enrivision
# Or run without installing
npx -y @bedolla/enrivision@latest --helpBuild
npm install
npm run typecheck
npm run buildUsage
1) Configure your MCP host
EnriVision runs as an MCP server over stdio. Your MCP host is responsible for launching the process.
Example: global install
{
"EnriVision": {
"type": "stdio",
"command": "enrivision",
"args": [],
"env": {
"ENRIPROXY_URL": "http://127.0.0.1:8787",
"ENRIPROXY_API_KEY": "YOUR_ENRIPROXY_API_KEY",
"ENRIVISION_DEFAULT_LANGUAGE": "es"
}
}
}Example: no install (always uses whatever npm currently tags as latest)
{
"EnriVision": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@bedolla/enrivision@latest"],
"env": {
"ENRIPROXY_URL": "http://127.0.0.1:8787",
"ENRIPROXY_API_KEY": "YOUR_ENRIPROXY_API_KEY",
"ENRIVISION_DEFAULT_LANGUAGE": "es"
}
}
}{
"EnriVision": {
"type": "stdio",
"command": "node",
"args": ["C:\\Users\\Administrator\\Projects\\EnriVision\\dist\\index.js"],
"env": {
"ENRIPROXY_URL": "http://127.0.0.1:8787",
"ENRIPROXY_API_KEY": "YOUR_ENRIPROXY_API_KEY",
"ENRIVISION_DEFAULT_LANGUAGE": "es"
}
}
}Configuration
EnriVision is configured via environment variables:
ENRIPROXY_URL(string, optional, default:http://127.0.0.1:8787)ENRIPROXY_API_KEY(string, required)ENRIVISION_TIMEOUT_MS(string, optional, default:1800000)Parsed as an integer (milliseconds). This is the operator cap: the per-call analyze timeout is
min(operator, mode budget)withsingle= 10 min (one pass, fast/cheap),multipass/auto= 20 min (per-segment/batch map + reduce;automay escalate to multipass server-side). Uploads are performed in chunks; per-chunk timeouts honormin(operator, derived 30s..300s)floored at 30 s (an operator budget below 30 s never forces tighter single-chunk budgets).
ENRIVISION_DEFAULT_LANGUAGE(string, optional)Default language to send when the tool call does not provide
language.
ENRIVISION_DENY_SYMLINKS(string, optional)Set to
1to reject symlinkedpath/pathsinputs. Strict mode opens withO_NOFOLLOW(POSIX) and compares thedev:inohandle identity fromfstat. On Windows (win32)O_NOFOLLOWis0(advisory only), so strict mode there rests solely on thelstat-vs-fstatcomparison with a small swap window: prefer POSIX hosts when symlink races are in scope.
ENRIVISION_MODEL(string, optional)Model id for server-side dispatch affinity; omit for auto-dispatch.
ENRIVISION_QUIET(string, optional)Set to
1to silence upload/retry progress lines on stderr.
Analysis budgets
The client analyze timeout is min(ENRIVISION_TIMEOUT_MS, mode budget):
single→ 10 min (mirrors EnriCode and the EnriProxy single-pass stage budget).multipass→ 20 min (mirrors EnriCode and the server multipass wall-clock budget).auto(default) → 20 min: the server picks the mode and may escalate to multipass, so the client cannot assume the short budget. If unsure, omit tuning (auto).
Error shape
Tool failures return MCP isError with Spanish-first bilingual text (ES first, EN second) plus machine-readable structuredContent: { code, retryable, httpStatus? } reusing the EnriCode vocabulary:
ENRICODE_ERR_TOOL_INPUT_INVALID— argument/tuning errors (including proxy 400/422). Never retry unchanged (retryable: false).ENRICODE_ERR_TOOL_EXECUTION_FAILED— server/transport failures.retryableis true for 408/429/5xx, false otherwise.ENRICODE_ERR_TOOL_EXECUTION_TIMEOUT— expired upload/analyze budgets (retryable: true; retry with a smaller scope).ENRICODE_ERR_TOOL_EXECUTION_ABORTED— caller-cancelled (retryable: false).
httpStatus is present only when the failure carries a proxy HTTP status.
MCP tools
EnriVision exposes this MCP tool:
analyze_media
General notes:
The tool accepts a single JSON object as its input (the MCP
arguments).At least one of
path,paths, orcursoris required. Whenpathscarries at least one valid entry,pathis ignored (explicit ignore-path contract: sending both is allowed,pathis silently ignored — prefer oneOf semantics and send only one). Acursor(from a truncated response) reads the next window of the list without uploading or analyzing anything.Paths must be absolute on the machine running the MCP server, or http(s) URLs. URLs are downloaded to a temporary directory on the MCP host (up to 64 MiB each; localhost and private-network destinations are blocked) and deleted after analysis. A solitary URL above 64 MiB escalates to EnriProxy's server-side
source_urlingestion (resumable download with extra hops); local files use resumable upload up to 4 GiB.EnriVision does not accept per-call
server_url/api_keyoverrides (these are configured via env vars).
analyze_media
Inputs:
path(string, optional): absolute local file path, or one http(s) URL to download and analyze (up to 64 MiB; a solitary larger URL escalates to server-sidesource_urlingestion).paths(string[], optional): absolute local image paths or http(s) image URLs (useful for UI screenshot sets).context(string, optional): high-level hint (examples:ui,diagram,chart,error,code,meeting,tutorial,photo).question(string, optional): what you want to extract/answer.language(string, optional): preferred response language (ISO 639-1; e.g.,es,en). If omitted, usesENRIVISION_DEFAULT_LANGUAGEwhen set.analysis_mode(string, optional):auto|single|multipass.max_frames(number, optional): single-pass video frames (1..20).model(string, optional): model id for server-side dispatch affinity (max 128 chars; envENRIVISION_MODEL; omit for auto-dispatch).region(object, optional): relative[0,1]zoom box{x, y, width, height}for one image (native-resolution reading of small text); single images only.transcribe(boolean, optional): enable/disable transcription (videos). Has no effect on images/documents (declared inwarnings, ignored).transcription_language(string, optional): whisper hint (auto,es,en, ...).
Continuation:
cursor(string, optional): opaque cursor from a truncated response (segment_summaries_cursorortranscription_segments_cursor); reads the next window without re-analyzing.offset(integer, optional): continuation start index (defaults to the responsenext_offset).limit(integer, optional): continuation window length (1..100; defaults to the server window size).
Video targeting:
video.clip_start_seconds(number, optional)video.clip_duration_seconds(number, optional)video.clip_end_seconds(number, optional; end = start + duration, wins overclip_duration_seconds)
Multipass tuning (advanced; used only for analysis_mode: multipass):
video.segment_seconds(number, optional)video.max_segments(number, optional)video.max_frames_per_segment(number, optional)document.max_pages_total(number, optional)document.pages_per_batch(number, optional)document.max_images_per_batch(number, optional)document.scanned_text_threshold_chars(number, optional)audio.timestamps(boolean, optional)audio.segment_seconds(number, optional)audio.max_segments(number, optional)images.max_images_total(number, optional)images.images_per_batch(number, optional)images.max_dimension(number, optional)
Output:
analysis(string): model-produced analysis.media_type(string): detected media type (video,audio,image,document,image_set).extraction(object): safe metadata summary (internal routing details are stripped).
Example arguments object:
{
"path": "C:\\path\\to\\video.mp4",
"question": "What are the key steps demonstrated?",
"analysis_mode": "auto",
"transcribe": true,
"language": "es"
}Many MCP clients include a built-in Read(...) tool that can ingest local files and attach them to the model request.
This is convenient, but the set of supported formats is limited and can change across client versions.
If the file you need to analyze is not reliably supported by your client (for example .avif, .heic, .svg, videos,
audio, or Office documents), prefer EnriVision MCP so the client can upload bytes and EnriProxy can do extraction reliably.
EnriProxy determines media type using content-type and extension allow-lists.
Videos:
.mp4,.mov,.avi,.mkv,.webm,.m4v,.wmv,.flv,.3gp,.3g2,.ts,.mts,.m2ts,.mpeg,.mpg,.gif
Audio:
.mp3,.mp1,.mp2,.mpa,.mpga,.wav,.aiff,.aif,.aifc,.caf,.flac,.m4a,.m4b,.m4r,.aac,.ogg,.oga,.wma,.opus,.weba,.mka
Images:
.png,.apng,.jpg,.jpeg,.gif,.webp,.avif,.heic,.heif,.tiff,.tif,.bmp,.svg,.ico
Documents:
.pdf,.docx,.pptx,.xlsx
Available Tools
1 toolanalyze_mediaAnálisis de medios EnriVisionARead-onlyIdempotent
Sube y analiza un archivo mediante EnriProxy (extracción del lado servidor + análisis con modelo).
/ Upload and analyze a media file via EnriProxy (server-side extraction + model analysis).
IMPORTANTE para modelos sin visión: la respuesta es SIEMPRE TEXTO (descripción visual generada del lado del servidor y/o transcripción del audio); nunca se devuelven bloques de imagen, así que CUALQUIER modelo puede consumirla — si no puedes ver imágenes, esta herramienta es tu vía para 'ver' archivos multimedia pidiendo la descripción en question.
/ IMPORTANT for models without vision: the response is ALWAYS TEXT (server-side visual description and/or audio transcription); image blocks are never returned, so ANY model can consume it — if you cannot see images, this tool is how you 'see' media by asking for the description in question.
Formatos aceptados — Imágenes estáticas: PNG, JPEG, WebP, TIFF, BMP, AVIF, HEIC/HEIF, SVG. Imágenes animadas: GIF animado, WebP animado, APNG, SVG animado (se extraen fotogramas clave y se describen sus cambios). Video: MP4, WebM, MKV, MOV, AVI y cualquier contenedor/códec decodificable, CON o SIN audio (en video con audio se procesan juntos: fotogramas + transcripción sobre la misma línea de tiempo). Audio: WAV, MP3, OGG (Vorbis/Opus), M4A/AAC, FLAC, ALAC, Opus y cualquier formato común (todo se normaliza a WAV 16 kHz mono antes de transcribir). Documentos: PDF (de texto, escaneado o mixto), Word (.docx), Excel (.xlsx), PowerPoint (.pptx) — con texto e imágenes embebidas — y JSONL. Conjuntos de varias imágenes: use paths.
/ Accepted formats — Static images: PNG, JPEG, WebP, TIFF, BMP, AVIF, HEIC/HEIF, SVG. Animated images: animated GIF, animated WebP, APNG, animated SVG (key frames are extracted and their changes described). Video: MP4, WebM, MKV, MOV, AVI and any decodable container/codec, WITH or WITHOUT audio (video with audio processes both together: frames + transcription on the same timeline). Audio: WAV, MP3, OGG (Vorbis/Opus), M4A/AAC, FLAC, ALAC, Opus and any common format (everything is normalized to 16 kHz mono WAV before transcription). Documents: PDF (text, scanned, or mixed), Word (.docx), Excel (.xlsx), PowerPoint (.pptx) — with embedded text and images — and JSONL. Multiple-image sets: use paths.
Cuándo usarla: PDFs grandes o escaneados donde el Read puede truncar; video/audio u otros binarios que el cliente no puede leer; HEIC/AVIF/TIFF/BMP/APNG/SVG/Office cuando el Read no es confiable; archivos muy grandes con subidas reanudables (hasta 4 GiB).
/ When to use: large or scanned PDFs where client Read may truncate; video/audio or binary media the client cannot Read; HEIC/AVIF/TIFF/BMP/APNG/SVG/Office docs when client Read is unreliable; very large files needing resumable uploads (up to 4 GiB).
Reglas: use path para un archivo, paths para varias imágenes. Cuando paths trae al menos una entrada válida, path se ignora (contrato explícito: mandar ambos se permite, path se ignora en silencio — prefiera semántica oneOf y mande solo uno). Las entradas en blanco se descartan; claves desconocidas en video/audio/document/images se rechazan. question es opcional aquí (obligatoria en EnriCode); attachmentIndex/attachmentId no existen aquí (sólo EnriCode).
/ Rules: use path for one file, paths for several images (UI screenshots/photo sets). When paths carries at least one valid entry, path is ignored (explicit ignore-path contract: sending both is allowed, path is silently ignored — prefer oneOf semantics and send only one). Blank paths entries are discarded; unknown keys inside video/audio/document/images are rejected (check typos like max_pages_totall). question is optional here (required in EnriCode vision.analyze_media); attachmentIndex/attachmentId do not exist here (EnriCode-only).
Presupuestos (timeout = min(ENRIVISION_TIMEOUT_MS del operador, presupuesto del modo)): single = 10 min (una pasada, rápida y barata); multipass = 20 min (por segmentos/lotes + reducción); auto = 20 min (el servidor elige y puede escalar a multipass). Si no sabe cuál usar, omita el afinado (auto).
/ Analysis budgets (client analyze timeout = min(operator ENRIVISION_TIMEOUT_MS, mode budget)): single = 10 min (one pass, fast and cheap, 1 image or simple questions); multipass = 20 min (per-segment/batch map + reduce; PDFs over ~20 pages, long videos, image sets); auto = 20 min (the server picks and may escalate to multipass). If unsure, omit tuning (auto).
Clip de video: para preguntas en un tiempo específico ("¿qué pasa en 12:34?") use video.clip_start_seconds + video.clip_duration_seconds: convierta a segundos (12:34 = 1260+34 = 754), por ejemplo clip_start_seconds=754 y clip_duration_seconds=30. O dé video.clip_end_seconds (fin = inicio + duración, 0-86400 s).
/ Video clip targeting: for time-specific questions ("what happens at 12:34?") use video.clip_start_seconds + video.clip_duration_seconds: convert to seconds (12:34 = 1260+34 = 754) and request a window, e.g. clip_start_seconds=754 and clip_duration_seconds=30. Or give video.clip_end_seconds instead (end = start + duration, 0-86400 s).
Enteros estrictos: los knobs enteros aceptan números o strings enteras completas ("60" vale; "8.0", "8abc" y 7.9 fallan). Los flotantes aceptan decimales ("12.5" vale). transcribe vale true por defecto y no tiene efecto en imágenes/documentos (se declara en warnings, se ignora). Requiere API key válida de EnriProxy (env ENRIPROXY_API_KEY).
/ Strict integers: integer knobs accept numbers or complete integer strings ("60" works; "8.0", "8abc", 7.9 fail). Floats accept decimals ("12.5" works). transcribe defaults to true and has no effect on images/documents (declared in warnings, ignored). Requires a valid EnriProxy API key (env ENRIPROXY_API_KEY, sent as Authorization: Bearer ...).
Errores: las fallas devuelven isError con texto bilingüe más structuredContent {code, retryable, httpStatus?} con el vocabulario EnriCode; retryable marca 429/5xx/timeouts. Si el mensaje trae Detalle del servidor: en el idioma del proxy, repórtelo tal cual. Fotogramas y transcripción comparten la MISMA línea de tiempo. model es el id del modelo para afinidad de dispatch (máximo 128 caracteres o env ENRIVISION_MODEL; omita para auto-dispatch). language controla el idioma de la RESPUESTA; transcription_language aparte el idioma que Whisper espera al TRANSCRIBIR ("auto" = detectar solo).
/ Errors: failures return isError with bilingual text plus structuredContent {code, retryable, httpStatus?} reusing the EnriCode vocabulary (ENRICODE_ERR_TOOL_INPUT_INVALID / EXECUTION_FAILED / EXECUTION_TIMEOUT / EXECUTION_ABORTED); retryable marks 429/5xx/timeouts. If the message carries a Detalle del servidor: fragment in the proxy language, report it verbatim. Video frames and transcription share the SAME video timeline. Animated GIF/WebP/APNG/SVG become representative key frames. model is the active model id for server-side dispatch affinity (max 128 chars, or env ENRIVISION_MODEL; omit for auto-dispatch). Set language (e.g. "es") to match the user language and avoid drift: language controls the analysis RESPONSE language; transcription_language separately controls the language Whisper expects when TRANSCRIBING audio ("auto" = detect only).
Ejemplos mínimos: (1) una imagen: {"path": "/tmp/foto.png", "question": "..."}. (2) clip de video 12:34->754s: {"path": "/tmp/charla.mp4", "question": "...", "video": {"clip_start_seconds": 754, "clip_duration_seconds": 30}}. (3) PDF largo multipass: {"path": "/tmp/manual.pdf", "question": "...", "analysis_mode": "multipass"}. Rutas absolutas del host MCP (en Windows valen C:/...; en POSIX lanzarían error).
/ Minimal examples: (1) single image: {"path": "/tmp/shot.png", "question": "What does each capture show?"}. (2) video clip 12:34->754s: {"path": "/tmp/talk.mp4", "question": "What happens at 12:34?", "video": {"clip_start_seconds": 754, "clip_duration_seconds": 30}}. (3) long PDF multipass: {"path": "/tmp/manual.pdf", "question": "Summarize each chapter.", "analysis_mode": "multipass"}. Absolute MCP-host paths (C:/... drive paths only work on a Windows host; POSIX rejects them).
Continuación: si la respuesta trae has_more con cursor (segment_summaries_cursor o transcription_segments_cursor), pida el resto con solo cursor (+ offset opcional, por defecto next_offset; también limit opcional 1-100 para acotar la ventana). Con cursor no mande path/paths. / Continuation: when the response carries has_more with a cursor (segment_summaries_cursor or transcription_segments_cursor), ask for the rest with only cursor (+ optional offset, defaults to next_offset; optional limit 1-100 bounds the window); never send path/paths with cursor.
Depuración de capturas de UI: abra con un veredicto de una línea; describa zona por zona; aproxime colores como hex; cuantifique defectos de layout; transcriba etiquetas, botones y errores visibles; compare observado vs esperado cuando aplique. / UI-screenshot debugging (when the media are app screenshots): open with a one-line plain verdict; describe zone by zone (header, sidebar, main content, modals, notifications), not as a general scene; approximate colors as hex values (e.g. #1F6FEB) and name them; quantify layout defects (overflows, clipping, overlaps, misalignments, missing spacing, cut text) estimating pixel magnitudes when possible; transcribe labels, buttons, and any visible error/status text; when the request states what was expected, compare observed vs expected explicitly.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Ruta absoluta a un archivo local en la máquina donde corre el servidor MCP (por ejemplo, C:\\Users\\User\\Downloads\\video.mp4), o una URL http(s) de imagen/video/audio/PDF para descargar y analizar (hasta 64 MiB; hosts locales y redes privadas bloqueados). Una URL solitaria que excede 64 MiB escala a la ingesta `source_url` del servidor (descarga reanudable del lado de EnriProxy con más hops y techo mayor); los archivos locales usan subida reanudable hasta 4 GiB. Cuando `paths` trae al menos una entrada válida, `path` se ignora. / Absolute local file path on the machine running this MCP server (e.g. C:\Users\User\Downloads\video.mp4), or one http(s) URL of image/video/audio/PDF to download and analyze (up to 64 MiB; localhost and private networks blocked). A solitary URL above 64 MiB escalates to the server's `source_url` ingestion (resumable server-side download with extra hops and a higher ceiling); local files use resumable upload up to 4 GiB. When `paths` carries at least one valid entry, `path` is ignored. | |
| audio | No | Ajuste opcional de multipass para audio (se usa sólo al analizar archivos de audio). Dentro de audio valen timestamps, audioTimestamps o audio_timestamps, segment_seconds o segmentSeconds, max_segments o maxSegments, y los planos audioTimestamps/audio_timestamps/segmentSeconds/segment_seconds/maxSegments/max_segments valen igual (el plano gana sobre ambos anidados). Sin plano, valores distintos entre video y audio para el mismo knob se rechazan. timestamps acepta true/false y "true"/"false". / Optional multipass tuning for audio (only used when analyzing audio files). Inside audio timestamps, audioTimestamps, or audio_timestamps work, as do segment_seconds/segmentSeconds and max_segments/maxSegments; flat aliases work the same (flat wins over both nested). Without a flat, differing video vs audio values for the same knob are rejected. timestamps accepts true/false and "true"/"false". | |
| limit | No | Máximo de entradas a leer en esta continuación (1-100; por defecto el tamaño de ventana del servidor). / Maximum entries to read in this continuation (1-100; defaults to the server window size). | |
| model | No | Id opcional del modelo activo para afinidad de dispatch del lado servidor (incluido el reroute Muse Spark); texto no vacío de máximo 128 caracteres. También acepta el env ENRIVISION_MODEL. Omita para auto-dispatch. / Optional active model id for server-side dispatch affinity (including the Muse Spark reroute); non-empty text, max 128 chars. Also accepts env ENRIVISION_MODEL. Omit for auto-dispatch. | |
| paths | No | Rutas absolutas a varios archivos de imagen locales o URLs http(s) (capturas de UI/sets de fotos; cada URL hasta 64 MiB). Cuando se proporcionan, EnriVision sube un único archivo de conjunto para procesamiento por lotes y reducción del lado servidor. Las entradas en blanco se descartan. / Absolute local paths to several image files, or http(s) image URLs (UI screenshots/photo sets; each URL up to 64 MiB). When provided, EnriVision uploads a single set archive for server-side batching + reduce. | |
| video | No | Ajuste opcional de multipass para video. Se usa sólo al analizar videos. Dentro de video valen snake_case y camelCase (clip_start_seconds o clipStartSeconds, segment_seconds o segmentSeconds, max_segments o maxSegments, max_frames_per_segment o maxFramesPerSegment), y los planos clipStartSeconds/clipEndSeconds/clipDurationSeconds/segmentSeconds/maxSegments/maxFramesPerSegment valen igual (el plano gana sobre ambos anidados). Sin plano, video.segment_seconds y audio.segment_seconds (o max_segments) con valores distintos se rechazan: use el plano o solo uno de los dos objetos. / Optional multipass tuning for video. Only used when analyzing videos. Inside video both snake_case and camelCase work, and the flat clipStartSeconds/clipEndSeconds/clipDurationSeconds/segmentSeconds/maxSegments/maxFramesPerSegment aliases work the same (flat wins over both nested). Without a flat, differing video.segment_seconds vs audio.segment_seconds (or max_segments) values are rejected: use the flat or only one of the two objects. | |
| cursor | No | Cursor opaco de continuación de una respuesta truncada (segment_summaries_cursor o transcription_segments_cursor). Con cursor NO se sube ni analiza nada: solo lee la siguiente ventana de la lista. No se combina con 'path'/'paths'. / Opaque continuation cursor from a truncated response (segment_summaries_cursor or transcription_segments_cursor). With cursor nothing is uploaded or analyzed: it only reads the next window of the list. Cannot be combined with 'path'/'paths'. | |
| images | No | Ajuste opcional de multipass para conjuntos de imágenes (se usa sólo con `paths`). Dentro de images valen snake_case y camelCase (max_images_total o maxImagesTotal, images_per_batch o imagesPerBatch, max_dimension o maxDimension). / Optional multipass tuning for image sets (only used with `paths`). Inside images snake_case and camelCase work. | |
| offset | No | Índice inicial de la continuación (entero >= 0; por defecto, el next_offset de la respuesta). / Continuation start index (integer >= 0; defaults to the response next_offset). | |
| region | No | Región relativa de la IMAGEN original para analizar a resolución nativa (zoom; acepta números y strings numéricas como "0.1"). Coordenadas entre 0 y 1; (0,0) es la esquina superior izquierda. Use las cajas devueltas en 'elements' de un análisis previo de la misma imagen: NUNCA invente coordenadas. Ideal para leer texto pequeño (etiquetas, código) que en la imagen completa comprimida resulta ilegible. Regla única: sólo imágenes individuales (`path` o `paths` con un solo elemento); con conjuntos de varias imágenes, video, PDF u otra media no-imagen la llamada se rechaza con error. / Relative REGION of the ORIGINAL image for native-resolution zoom (accepts numbers and numeric strings like "0.1"). Coords between 0 and 1; (0,0) is the top-left corner. Use the boxes returned in 'elements' of a previous analysis of the same image: NEVER invent coordinates. Ideal for small text (labels, code) illegible in the compressed full image. Single images only (`path` or single-entry `paths`); multi-image sets, video, PDF, or other non-image media are rejected. | |
| context | No | Pista opcional de análisis: ui, diagram, chart, error, code, meeting, tutorial, photo. Déjelo vacío para detección automática. Máximo 2000 caracteres; si los excede falla antes de subir. / Optional analysis hint: ui, diagram, chart, error, code, meeting, tutorial, photo. Leave empty for auto-detect. Max 2000 chars; longer fails before upload. | |
| document | No | Ajuste opcional de multipass para documentos (PDF). Dentro de document valen snake_case, camelCase y los legados max_pages/maxPages/documentMaxPages/document_max_pages. Los planos documentMaxPages/document_max_pages valen igual que document.max_pages_total (el plano gana). / Optional multipass tuning for documents (PDF). Inside document snake_case, camelCase, and legacy max_pages/maxPages/documentMaxPages/document_max_pages work. The flat documentMaxPages/document_max_pages aliases equal document.max_pages_total (flat wins). | |
| language | No | Código de idioma preferido de la RESPUESTA del análisis (ISO 639-1), por ejemplo 'es', 'en'. No afecta la transcripción: para eso use 'transcription_language'. Precedencia: parámetro explícito > ENRIVISION_DEFAULT_LANGUAGE > servidor. / Preferred RESPONSE language code of the analysis (ISO 639-1), e.g. 'es', 'en'. Does not affect transcription: use 'transcription_language' for that. Precedence: explicit param > ENRIVISION_DEFAULT_LANGUAGE > server. | |
| question | No | Pregunta explícita opcional que responder sobre el archivo (opcional aquí; en EnriCode vision.analyze_media es obligatoria). Máximo 2000 caracteres; si los excede falla antes de subir. / Optional explicit question to answer about the file (optional here; required in EnriCode vision.analyze_media). Max 2000 chars; longer fails before upload. | |
| maxFrames | No | Alias de max_frames (entero 1-20, por defecto 20, modo single). Las strings enteras completas valen. / Alias of max_frames (integer 1-20, default 20, single mode). Complete integer strings work. | |
| max_frames | No | Máximo opcional de fotogramas para videos, entero 1-20 (por defecto 20), en modo 'single' (pasada única). También acepta maxFrames. Para tiempos específicos, prefiera video.clip_start_seconds + video.clip_duration_seconds. Para multipass, use video.max_frames_per_segment. / Optional max frames for videos, integer 1-20 (default 20), in 'single' mode. Also accepts maxFrames; complete integer strings work. | |
| transcribe | No | Sobreescritura opcional para activar/desactivar la transcripción de audio en videos. Acepta true/false y "true"/"false" (los demás valores se rechazan). / Optional override to enable/disable audio transcription on videos. Accepts true/false and "true"/"false" (other values are rejected). Has no effect on images/documents (declared in warnings, ignored). | |
| maxSegments | No | Atajo plano de max_segments (entero 1-60; también vale max_segments). Misma precedencia que segmentSeconds: sin objetos alimenta a ambos, con uno alimenta a ese, con ambos distintos sin plano se rechaza. El plano gana. / Flat shortcut for max_segments (integer 1-60; max_segments also works). Same precedence as segmentSeconds. Flat wins. | |
| analysisMode | No | Alias de analysis_mode (mismo selector, mismos presupuestos). / Alias of analysis_mode (same selector, same budgets). | |
| max_segments | No | Alias plano de maxSegments (entero 1-60). El plano gana sobre video.max_segments y audio.max_segments. / Flat alias for maxSegments (integer 1-60). Flat wins over video.max_segments and audio.max_segments. | |
| analysis_mode | No | También acepta analysisMode. Selector opcional de modo de análisis. 'single' = una sola pasada, rápida y barata (1 imagen, preguntas simples). 'multipass' = por segmentos/lotes + reducción (PDFs de más de 20 páginas, videos largos, conjuntos). 'auto' = el servidor elige (prefiere multipass para PDFs de más de 20 páginas). Omita si no sabe cuál usar. / Also accepts analysisMode. Optional analysis-mode selector. 'single' = one pass, fast and cheap (1 image, simple questions). 'multipass' = per-segment/batch + reduce (PDFs over ~20 pages, long videos, sets). 'auto' = the server picks (prefers multipass for PDFs over ~20 pages). Omit if unsure. | |
| clipEndSeconds | No | Atajo plano de video.clip_end_seconds (0-86400 s, mayor que el inicio; también vale clip_end_seconds). La duración se calcula como fin menos inicio. El plano gana sobre el anidado. / Flat shortcut for video.clip_end_seconds (0-86400 s, greater than start; clip_end_seconds also works). Duration derives as end minus start. Flat wins over nested. | |
| segmentSeconds | No | Atajo plano de segment_seconds (5-600 s; también vale segment_seconds). Sin objetos video/audio alimenta a ambos y el servidor aplica el que corresponda; con un solo objeto alimenta a ese; con ambos y sin plano, valores distintos se rechazan. El plano gana sobre ambos anidados. / Flat shortcut for segment_seconds (5-600 s; segment_seconds also works). Without video/audio objects it feeds both and the server applies the matching one; with one object it feeds that one; with both and differing values (no flat) it is rejected. Flat wins over both nested. | |
| audioTimestamps | No | Atajo plano de audio.timestamps (también vale audio_timestamps; acepta true/false y "true"/"false"). Solo aplica a audio. El plano gana sobre el anidado. / Flat shortcut for audio.timestamps (audio_timestamps also works; accepts true/false and "true"/"false"). Audio only. Flat wins over nested. | |
| segment_seconds | No | Alias plano de segmentSeconds (5-600 s). El plano gana sobre video.segment_seconds y audio.segment_seconds. / Flat alias for segmentSeconds (5-600 s). Flat wins over video.segment_seconds and audio.segment_seconds. | |
| audio_timestamps | No | Alias plano de audioTimestamps (solo audio). El plano gana sobre el anidado. / Flat alias for audioTimestamps (audio only). Flat wins over nested. | |
| clipStartSeconds | No | Atajo plano de video.clip_start_seconds (0-86400 s; también vale clip_start_seconds). Para 12:34 use 754. El plano gana sobre el anidado. / Flat shortcut for video.clip_start_seconds (0-86400 s; clip_start_seconds also works). For 12:34 use 754. Flat wins over nested. | |
| clip_end_seconds | No | Alias plano de clipEndSeconds (0-86400 s). El plano gana sobre el anidado. / Flat alias for clipEndSeconds (0-86400 s). Flat wins over nested. | |
| documentMaxPages | No | Atajo plano de document.max_pages_total (entero 1-200; también vale document_max_pages). Solo aplica a documentos. El plano gana sobre el anidado. / Flat shortcut for document.max_pages_total (integer 1-200; document_max_pages also works). Documents only. Flat wins over nested. | |
| clip_start_seconds | No | Alias plano de clipStartSeconds (0-86400 s). El plano gana sobre el anidado. / Flat alias for clipStartSeconds (0-86400 s). Flat wins over nested. | |
| document_max_pages | No | Alias plano de documentMaxPages (entero 1-200, solo documentos). El plano gana sobre el anidado. / Flat alias for documentMaxPages (integer 1-200, documents only). Flat wins over nested. | |
| clipDurationSeconds | No | Atajo plano de video.clip_duration_seconds (mayor que 0, hasta 86400 s; también vale clip_duration_seconds). Úselo junto a clipStartSeconds. El plano gana sobre el anidado. / Flat shortcut for video.clip_duration_seconds (greater than 0, up to 86400 s; clip_duration_seconds also works). Use with clipStartSeconds. Flat wins over nested. | |
| maxFramesPerSegment | No | Atajo plano de video.max_frames_per_segment (entero 1-20; también vale max_frames_per_segment). Solo aplica a video; con audio se rechaza. El plano gana sobre el anidado. / Flat shortcut for video.max_frames_per_segment (integer 1-20; max_frames_per_segment also works). Video only; rejected with audio. Flat wins over nested. | |
| clip_duration_seconds | No | Alias plano de clipDurationSeconds (hasta 86400 s). El plano gana sobre el anidado. / Flat alias for clipDurationSeconds (up to 86400 s). Flat wins over nested. | |
| transcriptionLanguage | No | Alias de transcription_language (misma pista, misma precedencia). / Alias of transcription_language (same hint, same precedence). | |
| max_frames_per_segment | No | Alias plano de maxFramesPerSegment (entero 1-20, solo video). El plano gana sobre el anidado. / Flat alias for maxFramesPerSegment (integer 1-20, video only). Flat wins over nested. | |
| transcription_language | No | También acepta transcriptionLanguage. Pista opcional de idioma que Whisper espera al TRANSCRIBIR el audio/video (por ejemplo, 'auto', 'es', 'en'; 'auto' = detectar solo). No cambia el idioma de la respuesta: para eso use 'language'. / Also accepts transcriptionLanguage. Optional hint for the language Whisper expects when TRANSCRIBING audio/video (e.g. 'auto', 'es', 'en'; 'auto' = detect only). Does not change the response language: use 'language' for that. |
Output Schema
| Name | Required | Description |
|---|---|---|
| analysis | Yes | Análisis en texto producido por EnriProxy. / Text analysis produced by EnriProxy. |
| elements | No | Cajas de elementos detectados en análisis de imagen (coordenadas relativas 0-1, reutilizables como `region`). / Detected element boxes in image analyses (relative 0-1 coords, reusable as `region`). |
| warnings | No | Avisos de honestidad bilingües (por ejemplo, ventana de clip recortada al límite de 24 h). Solo presente cuando el parseo ajustó un valor pedido. / Honesty warnings (bilingual, e.g. clip window trimmed to the 24 h limit). Only present when parsing adjusted a requested value. |
| extraction | Yes | Metadatos de extracción devueltos por el servidor (sin identificadores internos). Las strings muy largas se recortan principio+fin con el marcador […truncado…]; la forma del objeto se preserva. / Extraction metadata returned by the server (no internal ids). Very long strings are head+tail trimmed with a […truncated…] marker; object shape is preserved. |
| media_type | Yes | Tipo de media detectado. / Detected media type. |
| analysis_truncated | No | `true` cuando `analysis` se truncó al tope de `structuredContent` (262144 caracteres en puntos de código; se conservan principio y fin). / Flag that is true when `analysis` was truncated to the `structuredContent` cap (262144 chars in code points; head and tail kept). |
| analysis_total_chars | No | Total de caracteres (puntos de código) del análisis completo antes de truncar. / Total chars (code points) of the full analysis before truncation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description discloses: response is ALWAYS text and never image blocks, error responses include isError plus structuredContent with retryable codes, integer knobs reject non-integral strings, `transcribe` is ignored for images/documents, video and audio share the same timeline, and budgets define timeouts. It also documents precedence rules ('flat wins over nested') and the `Detalle del servidor` error reporting requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very long, amplified by full bilingual duplicationahan, but extremely well organized with clear headings (Cuando usarla, Reglas, Presupuestos, Errores, Ejemplos minimos, Continuacion, Depuracion). Every sentence carries practical information, and the length is justified by 37 parameters and complex edge cases. It could be slightly trimmed but the density and structure earn a high score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all necessary operational aspects: accepted formats per media type, analysis budgets, strict integer/float rules, error vocabulary, continuation cursors, UI-debugging guidance, and three minimal examples. With an output schema present, nothing needed to invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While the schema already describes each parameter (100% coverage), the description adds critical meaning beyond it: the `path` vs `paths` oneOf contract (path ignored when paths is valid), flat-vs-nested precedence rules, alias resolution, concrete examples (12:34 -> 754 seconds, clip_start_seconds + clip_duration_seconds), and the multipass budget concept. It also clarifies that `question` is optional here but required in EnriCode, and warns about ignored parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Upload and analyze a media file via EnriProxy (server-side extraction + model analysis)' and immediately clarifies that the response is always text, positioning it as the tool for models without vision. It also distinguishes its role by enumerating when client Read is unreliable (large/scanned PDFs, video/audio, HEIC/AVIF/TIFF, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'Cuándo usarla' section lists precise conditions (large/scanned PDFs, video/audio, binary media, very large files with resumable uploads). It also explains when to use `path` vs `paths`, when continuation with `cursor` is mandatory, and contrasts with EnriCode (where `question` is required). Clear exclusions and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.7- Changed
analyze_media4 fields changed- changed
Input schema / anyOfPrevious value: -[ - { - "required": [ - "path" - ] - }, - { - "required": [ - "paths" - ] - } -]New value: +[ + { + "required": [ + "path" + ] + }, + { + "required": [ + "paths" + ] + }, + { + "required": [ + "cursor" + ] + } +] - added
Input schema / properties / limitAdded value: +{ + "description": "Máximo de entradas a leer en esta continuación (1-100; por defecto el tamaño de ventana del servidor). / Maximum entries to read in this continuation (1-100; defaults to the server window size).", + "maximum": 100, + "minimum": 1, + "type": "integer" +} - changed
Input schema / properties / offset / typePrevious value: -[ - "number", - "string" -]New value: +"integer" - changed
Input schema / properties / path / descriptionPrevious value: -"Ruta absoluta a un archivo local en la máquina donde corre el servidor MCP (por ejemplo, C:\\\\Users\\\\User\\\\Downloads\\\\video.mp4), o una URL http(s) de imagen/video/audio/PDF para descargar y analizar (hasta 64 MiB; hosts locales y redes privadas bloqueados). Cuando `paths` trae al menos una entrada válida, `path` se ignora. / Absolute local file path on the machine running this MCP server (e.g. C:\\Users\\User\\Downloads\\video.mp4), or one http(s) URL of image/video/audio/PDF to download and analyze (up to 64 MiB; localhost and private networks blocked)."New value: +"Ruta absoluta a un archivo local en la máquina donde corre el servidor MCP (por ejemplo, C:\\\\Users\\\\User\\\\Downloads\\\\video.mp4), o una URL http(s) de imagen/video/audio/PDF para descargar y analizar (hasta 64 MiB; hosts locales y redes privadas bloqueados). Una URL solitaria que excede 64 MiB escala a la ingesta `source_url` del servidor (descarga reanudable del lado de EnriProxy con más hops y techo mayor); los archivos locales usan subida reanudable hasta 4 GiB. Cuando `paths` trae al menos una entrada válida, `path` se ignora. / Absolute local file path on the machine running this MCP server (e.g. C:\\Users\\User\\Downloads\\video.mp4), or one http(s) URL of image/video/audio/PDF to download and analyze (up to 64 MiB; localhost and private networks blocked). A solitary URL above 64 MiB escalates to the server's `source_url` ingestion (resumable server-side download with extra hops and a higher ceiling); local files use resumable upload up to 4 GiB. When `paths` carries at least one valid entry, `path` is ignored."
1 tool update
v0.1.6- Changed
analyze_media100 fields changed- added
Input schema / properties / analysisModeAdded value: +{ + "description": "Alias de analysis_mode (mismo selector, mismos presupuestos). / Alias of analysis_mode (same selector, same budgets).", + "enum": [ + "auto", + "single", + "multipass" + ], + "type": "string" +} - changed
Input schema / properties / analysis_mode / descriptionPrevious value: -"Selector opcional de modo de análisis: auto, single o multipass."New value: +"También acepta analysisMode. Selector opcional de modo de análisis. 'single' = una sola pasada, rápida y barata (1 imagen, preguntas simples). 'multipass' = por segmentos/lotes + reducción (PDFs de más de 20 páginas, videos largos, conjuntos). 'auto' = el servidor elige (prefiere multipass para PDFs de más de 20 páginas). Omita si no sabe cuál usar. / Also accepts analysisMode. Optional analysis-mode selector. 'single' = one pass, fast and cheap (1 image, simple questions). 'multipass' = per-segment/batch + reduce (PDFs over ~20 pages, long videos, sets). 'auto' = the server picks (prefers multipass for PDFs over ~20 pages). Omit if unsure." - changed
Input schema / properties / audio / descriptionPrevious value: -"Ajuste opcional de multipass para audio (se usa sólo al analizar archivos de audio)."New value: +"Ajuste opcional de multipass para audio (se usa sólo al analizar archivos de audio). Dentro de audio valen timestamps, audioTimestamps o audio_timestamps, segment_seconds o segmentSeconds, max_segments o maxSegments, y los planos audioTimestamps/audio_timestamps/segmentSeconds/segment_seconds/maxSegments/max_segments valen igual (el plano gana sobre ambos anidados). Sin plano, valores distintos entre video y audio para el mismo knob se rechazan. timestamps acepta true/false y \"true\"/\"false\". / Optional multipass tuning for audio (only used when analyzing audio files). Inside audio timestamps, audioTimestamps, or audio_timestamps work, as do segment_seconds/segmentSeconds and max_segments/maxSegments; flat aliases work the same (flat wins over both nested). Without a flat, differing video vs audio values for the same knob are rejected. timestamps accepts true/false and \"true\"/\"false\"." - added
Input schema / properties / audio / properties / audioTimestampsAdded value: +{ + "description": "Alias de timestamps (acepta true/false y \"true\"/\"false\"). El plano gana sobre el anidado. / Alias of timestamps (accepts true/false and \"true\"/\"false\"). Flat wins over nested.", + "type": [ + "boolean", + "string" + ] +} - added
Input schema / properties / audio / properties / audio_timestampsAdded value: +{ + "description": "Alias de timestamps (solo audio). El plano gana sobre el anidado. / Alias of timestamps (audio only). Flat wins over nested.", + "type": [ + "boolean", + "string" + ] +} - added
Input schema / properties / audio / properties / maxSegmentsAdded value: +{ + "description": "Alias de max_segments para audio (entero 1-60). El plano gana sobre el anidado. / Alias of max_segments for audio (integer 1-60). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / audio / properties / max_segments / descriptionPrevious value: -"Número máximo de segmentos de audio a analizar."New value: +"Número máximo de segmentos de audio a analizar (entero 1-60; el servidor rechaza valores mayores). / Max audio segments to analyze (integer 1-60; the server rejects larger values)." - changed
Input schema / properties / audio / properties / max_segments / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / audio / properties / segmentSecondsAdded value: +{ + "description": "Alias de segment_seconds para audio (5-600; por defecto 60). El plano gana sobre el anidado. / Alias of segment_seconds for audio (5-600; default 60). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / audio / properties / segment_seconds / descriptionPrevious value: -"Duración del segmento en segundos para multipass de audio."New value: +"Duración del segmento en segundos para multipass de audio (5-600; por defecto 60). / Segment duration in seconds for audio multipass (5-600; default 60)." - changed
Input schema / properties / audio / properties / segment_seconds / typePrevious value: -"number"New value: +[ + "number", + "string" +] - changed
Input schema / properties / audio / properties / timestamps / descriptionPrevious value: -"Si incluir segmentos con marca de tiempo en la extracción de audio."New value: +"Si incluir segmentos con marca de tiempo en la extracción de audio. / Whether to include timestamped segments in the audio extraction." - changed
Input schema / properties / audio / properties / timestamps / typePrevious value: -"boolean"New value: +[ + "boolean", + "string" +] - added
Input schema / properties / audioTimestampsAdded value: +{ + "description": "Atajo plano de audio.timestamps (también vale audio_timestamps; acepta true/false y \"true\"/\"false\"). Solo aplica a audio. El plano gana sobre el anidado. / Flat shortcut for audio.timestamps (audio_timestamps also works; accepts true/false and \"true\"/\"false\"). Audio only. Flat wins over nested.", + "type": [ + "boolean", + "string" + ] +} - added
Input schema / properties / audio_timestampsAdded value: +{ + "description": "Alias plano de audioTimestamps (solo audio). El plano gana sobre el anidado. / Flat alias for audioTimestamps (audio only). Flat wins over nested.", + "type": [ + "boolean", + "string" + ] +} - added
Input schema / properties / clipDurationSecondsAdded value: +{ + "description": "Atajo plano de video.clip_duration_seconds (mayor que 0, hasta 86400 s; también vale clip_duration_seconds). Úselo junto a clipStartSeconds. El plano gana sobre el anidado. / Flat shortcut for video.clip_duration_seconds (greater than 0, up to 86400 s; clip_duration_seconds also works). Use with clipStartSeconds. Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / clipEndSecondsAdded value: +{ + "description": "Atajo plano de video.clip_end_seconds (0-86400 s, mayor que el inicio; también vale clip_end_seconds). La duración se calcula como fin menos inicio. El plano gana sobre el anidado. / Flat shortcut for video.clip_end_seconds (0-86400 s, greater than start; clip_end_seconds also works). Duration derives as end minus start. Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / clipStartSecondsAdded value: +{ + "description": "Atajo plano de video.clip_start_seconds (0-86400 s; también vale clip_start_seconds). Para 12:34 use 754. El plano gana sobre el anidado. / Flat shortcut for video.clip_start_seconds (0-86400 s; clip_start_seconds also works). For 12:34 use 754. Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / clip_duration_secondsAdded value: +{ + "description": "Alias plano de clipDurationSeconds (hasta 86400 s). El plano gana sobre el anidado. / Flat alias for clipDurationSeconds (up to 86400 s). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / clip_end_secondsAdded value: +{ + "description": "Alias plano de clipEndSeconds (0-86400 s). El plano gana sobre el anidado. / Flat alias for clipEndSeconds (0-86400 s). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / clip_start_secondsAdded value: +{ + "description": "Alias plano de clipStartSeconds (0-86400 s). El plano gana sobre el anidado. / Flat alias for clipStartSeconds (0-86400 s). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / context / descriptionPrevious value: -"Pista opcional de análisis: ui, diagram, chart, error, code, meeting, tutorial, photo. Déjelo vacío para detección automática."New value: +"Pista opcional de análisis: ui, diagram, chart, error, code, meeting, tutorial, photo. Déjelo vacío para detección automática. Máximo 2000 caracteres; si los excede falla antes de subir. / Optional analysis hint: ui, diagram, chart, error, code, meeting, tutorial, photo. Leave empty for auto-detect. Max 2000 chars; longer fails before upload." - added
Input schema / properties / cursorAdded value: +{ + "description": "Cursor opaco de continuación de una respuesta truncada (segment_summaries_cursor o transcription_segments_cursor). Con cursor NO se sube ni analiza nada: solo lee la siguiente ventana de la lista. No se combina con 'path'/'paths'. / Opaque continuation cursor from a truncated response (segment_summaries_cursor or transcription_segments_cursor). With cursor nothing is uploaded or analyzed: it only reads the next window of the list. Cannot be combined with 'path'/'paths'.", + "type": "string" +} - changed
Input schema / properties / document / descriptionPrevious value: -"Ajuste opcional de multipass para documentos (PDF)."New value: +"Ajuste opcional de multipass para documentos (PDF). Dentro de document valen snake_case, camelCase y los legados max_pages/maxPages/documentMaxPages/document_max_pages. Los planos documentMaxPages/document_max_pages valen igual que document.max_pages_total (el plano gana). / Optional multipass tuning for documents (PDF). Inside document snake_case, camelCase, and legacy max_pages/maxPages/documentMaxPages/document_max_pages work. The flat documentMaxPages/document_max_pages aliases equal document.max_pages_total (flat wins)." - added
Input schema / properties / document / properties / documentMaxPagesAdded value: +{ + "description": "Alias legado de max_pages_total (entero 1-200). El plano gana sobre el anidado. / Legacy alias of max_pages_total (integer 1-200). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / document / properties / document_max_pagesAdded value: +{ + "description": "Alias legado de max_pages_total (entero 1-200). El plano gana sobre el anidado. / Legacy alias of max_pages_total (integer 1-200). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / document / properties / maxImagesPerBatchAdded value: +{ + "description": "Alias de max_images_per_batch (entero 0-20; 0 = sin render). / Alias of max_images_per_batch (integer 0-20; 0 = no render).", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / document / properties / maxPagesAdded value: +{ + "description": "Alias legado de max_pages_total (entero 1-200). El plano gana sobre el anidado. / Legacy alias of max_pages_total (integer 1-200). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / document / properties / maxPagesTotalAdded value: +{ + "description": "Alias de max_pages_total (entero 1-200; por defecto 20). El plano gana sobre el anidado. / Alias of max_pages_total (integer 1-200; default 20). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / document / properties / max_images_per_batch / descriptionPrevious value: -"Máximo de páginas renderizadas (imágenes) por lote."New value: +"Máximo de páginas renderizadas (imágenes) por lote (entero 0-20; 0 = sin render). Más imágenes = más costo de visión; omita para el valor del servidor. / Max rendered (image) pages per batch (integer 0-20; 0 = no render). More images = more vision cost; omit for the server value." - changed
Input schema / properties / document / properties / max_images_per_batch / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / document / properties / max_pagesAdded value: +{ + "description": "Alias legado de max_pages_total (entero 1-200). El plano gana sobre el anidado. / Legacy alias of max_pages_total (integer 1-200). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / document / properties / max_pages_total / descriptionPrevious value: -"Número máximo de páginas a analizar en total."New value: +"Número máximo de páginas a analizar en total (entero 1-200; por defecto 20). Más páginas = más costo y tiempo; omita para pocas páginas. / Max pages to analyze in total (integer 1-200; default 20). More pages = more cost and time; omit for few pages." - changed
Input schema / properties / document / properties / max_pages_total / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / document / properties / pagesPerBatchAdded value: +{ + "description": "Alias de pages_per_batch (entero 1-200). / Alias of pages_per_batch (integer 1-200).", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / document / properties / pages_per_batch / descriptionPrevious value: -"Páginas por lote para las llamadas map de multipass."New value: +"Páginas por lote para las llamadas map de multipass (entero 1-200). Lotes chicos = más llamadas pero menos memoria; omita para el valor del servidor. / Pages per batch for multipass map calls (integer 1-200). Smaller batches = more calls but less memory; omit for the server value." - changed
Input schema / properties / document / properties / pages_per_batch / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / document / properties / scannedTextThresholdCharsAdded value: +{ + "description": "Alias de scanned_text_threshold_chars (entero 0-5000). / Alias of scanned_text_threshold_chars (integer 0-5000).", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / document / properties / scanned_text_threshold_chars / descriptionPrevious value: -"Longitud mínima de texto extraído para tratar una página como textual."New value: +"Longitud mínima de texto extraído para tratar una página como textual en vez de escaneada (entero 0-5000). Sólo afecta el enrutamiento texto-vs-visión; omita para el valor del servidor. / Min extracted-text length to treat a page as textual instead of scanned (integer 0-5000). Only affects text-vs-vision routing; omit for the server value." - changed
Input schema / properties / document / properties / scanned_text_threshold_chars / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / documentMaxPagesAdded value: +{ + "description": "Atajo plano de document.max_pages_total (entero 1-200; también vale document_max_pages). Solo aplica a documentos. El plano gana sobre el anidado. / Flat shortcut for document.max_pages_total (integer 1-200; document_max_pages also works). Documents only. Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / document_max_pagesAdded value: +{ + "description": "Alias plano de documentMaxPages (entero 1-200, solo documentos). El plano gana sobre el anidado. / Flat alias for documentMaxPages (integer 1-200, documents only). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / images / descriptionPrevious value: -"Ajuste opcional de multipass para conjuntos de imágenes (se usa sólo con `paths`)."New value: +"Ajuste opcional de multipass para conjuntos de imágenes (se usa sólo con `paths`). Dentro de images valen snake_case y camelCase (max_images_total o maxImagesTotal, images_per_batch o imagesPerBatch, max_dimension o maxDimension). / Optional multipass tuning for image sets (only used with `paths`). Inside images snake_case and camelCase work." - added
Input schema / properties / images / properties / imagesPerBatchAdded value: +{ + "description": "Alias de images_per_batch (entero 1-20). / Alias of images_per_batch (integer 1-20).", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / images / properties / images_per_batch / descriptionPrevious value: -"Imágenes por lote para las llamadas map de multipass."New value: +"Imágenes por lote para las llamadas map de multipass (entero 1-20). Lotes chicos = más llamadas pero menos memoria; omita para el valor del servidor. / Images per batch for multipass map calls (integer 1-20). Smaller batches = more calls but less memory; omit for the server value." - changed
Input schema / properties / images / properties / images_per_batch / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / images / properties / maxDimensionAdded value: +{ + "description": "Alias de max_dimension (entero 256-4096). / Alias of max_dimension (integer 256-4096).", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / images / properties / maxImagesTotalAdded value: +{ + "description": "Alias de max_images_total (entero 1-500). / Alias of max_images_total (integer 1-500).", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / images / properties / max_dimension / descriptionPrevious value: -"Dimensión máxima para las imágenes (ancho/alto)."New value: +"Dimensión máxima de cada imagen en píxeles, ancho/alto (entero 256-4096). Valores grandes = más detalle y más costo; omita para el valor del servidor. / Max image dimension in pixels, width/height (integer 256-4096). Larger values = more detail and more cost; omit for the server value." - changed
Input schema / properties / images / properties / max_dimension / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - changed
Input schema / properties / images / properties / max_images_total / descriptionPrevious value: -"Número máximo de imágenes a analizar en total."New value: +"Número máximo de imágenes del conjunto a analizar (entero 1-500). Más imágenes = más costo y tiempo; omita para analizarlas todas. / Max set images to analyze (integer 1-500). More images = more cost and time; omit to analyze all." - changed
Input schema / properties / images / properties / max_images_total / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - changed
Input schema / properties / language / descriptionPrevious value: -"Código de idioma preferido de respuesta (ISO 639-1), por ejemplo 'es', 'en'."New value: +"Código de idioma preferido de la RESPUESTA del análisis (ISO 639-1), por ejemplo 'es', 'en'. No afecta la transcripción: para eso use 'transcription_language'. Precedencia: parámetro explícito > ENRIVISION_DEFAULT_LANGUAGE > servidor. / Preferred RESPONSE language code of the analysis (ISO 639-1), e.g. 'es', 'en'. Does not affect transcription: use 'transcription_language' for that. Precedence: explicit param > ENRIVISION_DEFAULT_LANGUAGE > server." - added
Input schema / properties / maxFramesAdded value: +{ + "description": "Alias de max_frames (entero 1-20, por defecto 20, modo single). Las strings enteras completas valen. / Alias of max_frames (integer 1-20, default 20, single mode). Complete integer strings work.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / maxFramesPerSegmentAdded value: +{ + "description": "Atajo plano de video.max_frames_per_segment (entero 1-20; también vale max_frames_per_segment). Solo aplica a video; con audio se rechaza. El plano gana sobre el anidado. / Flat shortcut for video.max_frames_per_segment (integer 1-20; max_frames_per_segment also works). Video only; rejected with audio. Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / maxSegmentsAdded value: +{ + "description": "Atajo plano de max_segments (entero 1-60; también vale max_segments). Misma precedencia que segmentSeconds: sin objetos alimenta a ambos, con uno alimenta a ese, con ambos distintos sin plano se rechaza. El plano gana. / Flat shortcut for max_segments (integer 1-60; max_segments also works). Same precedence as segmentSeconds. Flat wins.", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / max_frames / descriptionPrevious value: -"Máximo opcional de fotogramas para videos (1-20) en modo single-pass. Para tiempos específicos, prefiera video.clip_start_seconds + video.clip_duration_seconds. Para multipass, use video.max_frames_per_segment."New value: +"Máximo opcional de fotogramas para videos, entero 1-20 (por defecto 20), en modo 'single' (pasada única). También acepta maxFrames. Para tiempos específicos, prefiera video.clip_start_seconds + video.clip_duration_seconds. Para multipass, use video.max_frames_per_segment. / Optional max frames for videos, integer 1-20 (default 20), in 'single' mode. Also accepts maxFrames; complete integer strings work." - changed
Input schema / properties / max_frames / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / max_frames_per_segmentAdded value: +{ + "description": "Alias plano de maxFramesPerSegment (entero 1-20, solo video). El plano gana sobre el anidado. / Flat alias for maxFramesPerSegment (integer 1-20, video only). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / max_segmentsAdded value: +{ + "description": "Alias plano de maxSegments (entero 1-60). El plano gana sobre video.max_segments y audio.max_segments. / Flat alias for maxSegments (integer 1-60). Flat wins over video.max_segments and audio.max_segments.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / modelAdded value: +{ + "description": "Id opcional del modelo activo para afinidad de dispatch del lado servidor (incluido el reroute Muse Spark); texto no vacío de máximo 128 caracteres. También acepta el env ENRIVISION_MODEL. Omita para auto-dispatch. / Optional active model id for server-side dispatch affinity (including the Muse Spark reroute); non-empty text, max 128 chars. Also accepts env ENRIVISION_MODEL. Omit for auto-dispatch.", + "type": "string" +} - added
Input schema / properties / offsetAdded value: +{ + "description": "Índice inicial de la continuación (entero >= 0; por defecto, el next_offset de la respuesta). / Continuation start index (integer >= 0; defaults to the response next_offset).", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / path / descriptionPrevious value: -"Ruta absoluta a un archivo local en la máquina donde corre el servidor MCP (por ejemplo, C:\\\\Users\\\\User\\\\Downloads\\\\video.mp4), o una URL http(s) de imagen/video/audio/PDF para descargar y analizar (hasta 64 MiB; hosts locales y redes privadas bloqueados)."New value: +"Ruta absoluta a un archivo local en la máquina donde corre el servidor MCP (por ejemplo, C:\\\\Users\\\\User\\\\Downloads\\\\video.mp4), o una URL http(s) de imagen/video/audio/PDF para descargar y analizar (hasta 64 MiB; hosts locales y redes privadas bloqueados). Cuando `paths` trae al menos una entrada válida, `path` se ignora. / Absolute local file path on the machine running this MCP server (e.g. C:\\Users\\User\\Downloads\\video.mp4), or one http(s) URL of image/video/audio/PDF to download and analyze (up to 64 MiB; localhost and private networks blocked)." - changed
Input schema / properties / paths / descriptionPrevious value: -"Rutas absolutas a varios archivos de imagen locales o URLs http(s) (capturas de UI/sets de fotos; cada URL hasta 64 MiB). Cuando se proporcionan, EnriVision sube un único archivo de conjunto para procesamiento por lotes y reducción del lado servidor."New value: +"Rutas absolutas a varios archivos de imagen locales o URLs http(s) (capturas de UI/sets de fotos; cada URL hasta 64 MiB). Cuando se proporcionan, EnriVision sube un único archivo de conjunto para procesamiento por lotes y reducción del lado servidor. Las entradas en blanco se descartan. / Absolute local paths to several image files, or http(s) image URLs (UI screenshots/photo sets; each URL up to 64 MiB). When provided, EnriVision uploads a single set archive for server-side batching + reduce." - added
Input schema / properties / paths / items / descriptionAdded value: +"Una imagen: ruta absoluta local o URL http(s) (hasta 64 MiB; hosts locales y redes privadas bloqueados). / One image: absolute local path or http(s) URL (up to 64 MiB; localhost and private networks blocked)." - changed
Input schema / properties / question / descriptionPrevious value: -"Pregunta explícita opcional que responder sobre el archivo."New value: +"Pregunta explícita opcional que responder sobre el archivo (opcional aquí; en EnriCode vision.analyze_media es obligatoria). Máximo 2000 caracteres; si los excede falla antes de subir. / Optional explicit question to answer about the file (optional here; required in EnriCode vision.analyze_media). Max 2000 chars; longer fails before upload." - changed
Input schema / properties / region / descriptionPrevious value: -"Región relativa de la IMAGEN original para analizar a resolución nativa (zoom). Coordenadas entre 0 y 1; (0,0) es la esquina superior izquierda. Use las cajas devueltas en 'elements' de un análisis previo de la misma imagen: NUNCA invente coordenadas. Ideal para leer texto pequeño (labels, código) que en la imagen completa comprimida resulta ilegible. Sólo imágenes (path, no paths)."New value: +"Región relativa de la IMAGEN original para analizar a resolución nativa (zoom; acepta números y strings numéricas como \"0.1\"). Coordenadas entre 0 y 1; (0,0) es la esquina superior izquierda. Use las cajas devueltas en 'elements' de un análisis previo de la misma imagen: NUNCA invente coordenadas. Ideal para leer texto pequeño (etiquetas, código) que en la imagen completa comprimida resulta ilegible. Regla única: sólo imágenes individuales (`path` o `paths` con un solo elemento); con conjuntos de varias imágenes, video, PDF u otra media no-imagen la llamada se rechaza con error. / Relative REGION of the ORIGINAL image for native-resolution zoom (accepts numbers and numeric strings like \"0.1\"). Coords between 0 and 1; (0,0) is the top-left corner. Use the boxes returned in 'elements' of a previous analysis of the same image: NEVER invent coordinates. Ideal for small text (labels, code) illegible in the compressed full image. Single images only (`path` or single-entry `paths`); multi-image sets, video, PDF, or other non-image media are rejected." - changed
Input schema / properties / region / properties / height / descriptionPrevious value: -"Alto relativo (1 = alto completo)."New value: +"Alto relativo (1 = alto completo). / Relative height (1 = full height)." - changed
Input schema / properties / region / properties / height / typePrevious value: -"number"New value: +[ + "number", + "string" +] - changed
Input schema / properties / region / properties / width / descriptionPrevious value: -"Ancho relativo (1 = ancho completo)."New value: +"Ancho relativo (1 = ancho completo). / Relative width (1 = full width)." - changed
Input schema / properties / region / properties / width / typePrevious value: -"number"New value: +[ + "number", + "string" +] - changed
Input schema / properties / region / properties / x / descriptionPrevious value: -"Coordenada horizontal relativa de la esquina superior izquierda (0 = borde izquierdo)."New value: +"Coordenada horizontal relativa de la esquina superior izquierda (0 = borde izquierdo). / Relative horizontal coord of the top-left corner (0 = left edge)." - changed
Input schema / properties / region / properties / x / typePrevious value: -"number"New value: +[ + "number", + "string" +] - changed
Input schema / properties / region / properties / y / descriptionPrevious value: -"Coordenada vertical relativa de la esquina superior izquierda (0 = borde superior)."New value: +"Coordenada vertical relativa de la esquina superior izquierda (0 = borde superior). / Relative vertical coord of the top-left corner (0 = top edge)." - changed
Input schema / properties / region / properties / y / typePrevious value: -"number"New value: +[ + "number", + "string" +] - added
Input schema / properties / segmentSecondsAdded value: +{ + "description": "Atajo plano de segment_seconds (5-600 s; también vale segment_seconds). Sin objetos video/audio alimenta a ambos y el servidor aplica el que corresponda; con un solo objeto alimenta a ese; con ambos y sin plano, valores distintos se rechazan. El plano gana sobre ambos anidados. / Flat shortcut for segment_seconds (5-600 s; segment_seconds also works). Without video/audio objects it feeds both and the server applies the matching one; with one object it feeds that one; with both and differing values (no flat) it is rejected. Flat wins over both nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / segment_secondsAdded value: +{ + "description": "Alias plano de segmentSeconds (5-600 s). El plano gana sobre video.segment_seconds y audio.segment_seconds. / Flat alias for segmentSeconds (5-600 s). Flat wins over video.segment_seconds and audio.segment_seconds.", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / transcribe / descriptionPrevious value: -"Sobreescritura opcional para activar/desactivar la transcripción de audio en videos."New value: +"Sobreescritura opcional para activar/desactivar la transcripción de audio en videos. Acepta true/false y \"true\"/\"false\" (los demás valores se rechazan). / Optional override to enable/disable audio transcription on videos. Accepts true/false and \"true\"/\"false\" (other values are rejected). Has no effect on images/documents (declared in warnings, ignored)." - changed
Input schema / properties / transcribe / typePrevious value: -"boolean"New value: +[ + "boolean", + "string" +] - added
Input schema / properties / transcriptionLanguageAdded value: +{ + "description": "Alias de transcription_language (misma pista, misma precedencia). / Alias of transcription_language (same hint, same precedence).", + "type": "string" +} - changed
Input schema / properties / transcription_language / descriptionPrevious value: -"Pista opcional de idioma para la transcripción de audio/video (por ejemplo, 'auto', 'es', 'en')."New value: +"También acepta transcriptionLanguage. Pista opcional de idioma que Whisper espera al TRANSCRIBIR el audio/video (por ejemplo, 'auto', 'es', 'en'; 'auto' = detectar solo). No cambia el idioma de la respuesta: para eso use 'language'. / Also accepts transcriptionLanguage. Optional hint for the language Whisper expects when TRANSCRIBING audio/video (e.g. 'auto', 'es', 'en'; 'auto' = detect only). Does not change the response language: use 'language' for that." - changed
Input schema / properties / video / descriptionPrevious value: -"Ajuste opcional de multipass para video. Se usa sólo al analizar videos."New value: +"Ajuste opcional de multipass para video. Se usa sólo al analizar videos. Dentro de video valen snake_case y camelCase (clip_start_seconds o clipStartSeconds, segment_seconds o segmentSeconds, max_segments o maxSegments, max_frames_per_segment o maxFramesPerSegment), y los planos clipStartSeconds/clipEndSeconds/clipDurationSeconds/segmentSeconds/maxSegments/maxFramesPerSegment valen igual (el plano gana sobre ambos anidados). Sin plano, video.segment_seconds y audio.segment_seconds (o max_segments) con valores distintos se rechazan: use el plano o solo uno de los dos objetos. / Optional multipass tuning for video. Only used when analyzing videos. Inside video both snake_case and camelCase work, and the flat clipStartSeconds/clipEndSeconds/clipDurationSeconds/segmentSeconds/maxSegments/maxFramesPerSegment aliases work the same (flat wins over both nested). Without a flat, differing video.segment_seconds vs audio.segment_seconds (or max_segments) values are rejected: use the flat or only one of the two objects." - added
Input schema / properties / video / properties / clipDurationSecondsAdded value: +{ + "description": "Alias de clip_duration_seconds (mayor que 0, hasta 86400). El plano gana sobre el anidado. / Alias of clip_duration_seconds (greater than 0, up to 86400). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / video / properties / clipEndSecondsAdded value: +{ + "description": "Alias de clip_end_seconds (0-86400, debe ser mayor que el inicio). El plano gana sobre el anidado. / Alias of clip_end_seconds (0-86400, must exceed start). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - added
Input schema / properties / video / properties / clipStartSecondsAdded value: +{ + "description": "Alias de clip_start_seconds (0-86400). El plano clipStartSeconds gana sobre el anidado. / Alias of clip_start_seconds (0-86400). Flat clipStartSeconds wins over nested.", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / video / properties / clip_duration_seconds / descriptionPrevious value: -"Duración opcional del clip en segundos para análisis de video dirigido a un tiempo."New value: +"Duración opcional del clip en segundos (mayor que 0, hasta 86400). Úsela junto a clip_start_seconds; si da clip_end_seconds, no la necesita. Si el fin implícito (inicio + duración) excede 86400 segundos, la duración se recorta al límite con un aviso en 'warnings'. / Optional clip duration in seconds (greater than 0, up to 86400). Use with clip_start_seconds; not needed with clip_end_seconds. When start + duration exceeds 86400 s the duration is trimmed to the limit with a 'warnings' note." - changed
Input schema / properties / video / properties / clip_duration_seconds / typePrevious value: -"number"New value: +[ + "number", + "string" +] - added
Input schema / properties / video / properties / clip_end_secondsAdded value: +{ + "description": "Fin opcional del clip en segundos (0-86400, debe ser mayor que el inicio). Si se da, la duración se calcula como fin menos inicio e ignora clip_duration_seconds. / Optional clip end in seconds (0-86400, must exceed start). When given, duration derives as end minus start and clip_duration_seconds is ignored.", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / video / properties / clip_start_seconds / descriptionPrevious value: -"Offset opcional de inicio del clip en segundos para análisis de video dirigido a un tiempo."New value: +"Inicio opcional del clip en segundos (0-86400). Para 12:34 use 754 (= 12*60+34). Con clip_end_seconds, fin = inicio + duración. / Optional clip start in seconds (0-86400). For 12:34 use 754 (= 12*60+34). With clip_end_seconds, end = start + duration." - changed
Input schema / properties / video / properties / clip_start_seconds / typePrevious value: -"number"New value: +[ + "number", + "string" +] - added
Input schema / properties / video / properties / maxFramesPerSegmentAdded value: +{ + "description": "Alias de max_frames_per_segment (entero 1-20; por defecto 8). El plano gana sobre el anidado. / Alias of max_frames_per_segment (integer 1-20; default 8). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - added
Input schema / properties / video / properties / maxSegmentsAdded value: +{ + "description": "Alias de max_segments para video (entero 1-60). El plano gana sobre el anidado. / Alias of max_segments for video (integer 1-60). Flat wins over nested.", + "type": [ + "integer", + "string" + ] +} - changed
Input schema / properties / video / properties / max_frames_per_segment / descriptionPrevious value: -"Máximo de fotogramas a extraer por segmento."New value: +"Máximo de fotogramas a extraer por segmento de video (entero 1-20; por defecto 8). / Max frames to extract per video segment (integer 1-20; default 8)." - changed
Input schema / properties / video / properties / max_frames_per_segment / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - changed
Input schema / properties / video / properties / max_segments / descriptionPrevious value: -"Número máximo de segmentos a analizar."New value: +"Número máximo de segmentos de video a analizar (entero 1-60; el servidor rechaza valores mayores). / Max video segments to analyze (integer 1-60; the server rejects larger values)." - changed
Input schema / properties / video / properties / max_segments / typePrevious value: -"integer"New value: +[ + "integer", + "string" +] - added
Input schema / properties / video / properties / segmentSecondsAdded value: +{ + "description": "Alias de segment_seconds para video (5-600; por defecto 60). El plano gana sobre el anidado. / Alias of segment_seconds for video (5-600; default 60). Flat wins over nested.", + "type": [ + "number", + "string" + ] +} - changed
Input schema / properties / video / properties / segment_seconds / descriptionPrevious value: -"Duración del segmento en segundos."New value: +"Duración del segmento en segundos para video (5-600; por defecto 60). / Segment duration in seconds for video (5-600; default 60)." - changed
Input schema / properties / video / properties / segment_seconds / typePrevious value: -"number"New value: +[ + "number", + "string" +] - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "analysis": { + "description": "Análisis en texto producido por EnriProxy. / Text analysis produced by EnriProxy.", + "type": "string" + }, + "analysis_total_chars": { + "description": "Total de caracteres (puntos de código) del análisis completo antes de truncar. / Total chars (code points) of the full analysis before truncation.", + "type": "integer" + }, + "analysis_truncated": { + "description": "`true` cuando `analysis` se truncó al tope de `structuredContent` (262144 caracteres en puntos de código; se conservan principio y fin). / Flag that is true when `analysis` was truncated to the `structuredContent` cap (262144 chars in code points; head and tail kept).", + "type": "boolean" + }, + "elements": { + "description": "Cajas de elementos detectados en análisis de imagen (coordenadas relativas 0-1, reutilizables como `region`). / Detected element boxes in image analyses (relative 0-1 coords, reusable as `region`).", + "items": { + "properties": { + "box": { + "properties": { + "height": { + "type": "number" + }, + "width": { + "type": "number" + }, + "x": { + "type": "number" + }, + "y": { + "type": "number" + } + }, + "required": [ + "x", + "y", + "width", + "height" + ], + "type": "object" + }, + "label": { + "type": "string" + } + }, + "required": [ + "label", + "box" + ], + "type": "object" + }, + "type": "array" + }, + "extraction": { + "description": "Metadatos de extracción devueltos por el servidor (sin identificadores internos). Las strings muy largas se recortan principio+fin con el marcador […truncado…]; la forma del objeto se preserva. / Extraction metadata returned by the server (no internal ids). Very long strings are head+tail trimmed with a […truncated…] marker; object shape is preserved.", + "type": "object" + }, + "media_type": { + "description": "Tipo de media detectado. / Detected media type.", + "type": "string" + }, + "warnings": { + "description": "Avisos de honestidad bilingües (por ejemplo, ventana de clip recortada al límite de 24 h). Solo presente cuando el parseo ajustó un valor pedido. / Honesty warnings (bilingual, e.g. clip window trimmed to the 24 h limit). Only present when parsing adjusted a requested value.", + "items": { + "type": "string" + }, + "type": "array" + } + }, + "required": [ + "analysis", + "media_type", + "extraction" + ], + "type": "object" +}
1 tool update
v0.1.5- Changed
analyze_media29 fields changed- changed
Input schema / properties / analysis_mode / descriptionPrevious value: -"Optional analysis mode selector: auto, single, or multipass."New value: +"Selector opcional de modo de análisis: auto, single o multipass." - changed
Input schema / properties / audio / descriptionPrevious value: -"Optional audio multipass tuning (used only when analyzing audio files)."New value: +"Ajuste opcional de multipass para audio (se usa sólo al analizar archivos de audio)." - changed
Input schema / properties / audio / properties / max_segments / descriptionPrevious value: -"Maximum number of audio segments to analyze."New value: +"Número máximo de segmentos de audio a analizar." - changed
Input schema / properties / audio / properties / segment_seconds / descriptionPrevious value: -"Segment duration in seconds for audio multipass."New value: +"Duración del segmento en segundos para multipass de audio." - changed
Input schema / properties / audio / properties / timestamps / descriptionPrevious value: -"Whether to include timestamped segments in audio extraction."New value: +"Si incluir segmentos con marca de tiempo en la extracción de audio." - changed
Input schema / properties / context / descriptionPrevious value: -"Optional analysis hint: ui, diagram, chart, error, code, meeting, tutorial, photo. Leave empty for auto-detection."New value: +"Pista opcional de análisis: ui, diagram, chart, error, code, meeting, tutorial, photo. Déjelo vacío para detección automática." - changed
Input schema / properties / document / descriptionPrevious value: -"Optional document multipass tuning (PDF)."New value: +"Ajuste opcional de multipass para documentos (PDF)." - changed
Input schema / properties / document / properties / max_images_per_batch / descriptionPrevious value: -"Maximum rendered pages (images) per batch."New value: +"Máximo de páginas renderizadas (imágenes) por lote." - changed
Input schema / properties / document / properties / max_pages_total / descriptionPrevious value: -"Maximum number of pages to analyze in total."New value: +"Número máximo de páginas a analizar en total." - changed
Input schema / properties / document / properties / pages_per_batch / descriptionPrevious value: -"Pages per batch for multipass map calls."New value: +"Páginas por lote para las llamadas map de multipass." - changed
Input schema / properties / document / properties / scanned_text_threshold_chars / descriptionPrevious value: -"Minimum extracted text length to treat a page as textual."New value: +"Longitud mínima de texto extraído para tratar una página como textual." - changed
Input schema / properties / images / descriptionPrevious value: -"Optional image-set multipass tuning (used only with `paths`)."New value: +"Ajuste opcional de multipass para conjuntos de imágenes (se usa sólo con `paths`)." - changed
Input schema / properties / images / properties / images_per_batch / descriptionPrevious value: -"Images per batch for multipass map calls."New value: +"Imágenes por lote para las llamadas map de multipass." - changed
Input schema / properties / images / properties / max_dimension / descriptionPrevious value: -"Maximum dimension for images (width/height)."New value: +"Dimensión máxima para las imágenes (ancho/alto)." - changed
Input schema / properties / images / properties / max_images_total / descriptionPrevious value: -"Maximum number of images to analyze in total."New value: +"Número máximo de imágenes a analizar en total." - changed
Input schema / properties / language / descriptionPrevious value: -"Preferred response language code (ISO 639-1), e.g. 'es', 'en'."New value: +"Código de idioma preferido de respuesta (ISO 639-1), por ejemplo 'es', 'en'." - changed
Input schema / properties / max_frames / descriptionPrevious value: -"Optional max frames for videos (1-20) in single-pass mode. For targeted timestamps, prefer video.clip_start_seconds + video.clip_duration_seconds. For multipass, use video.max_frames_per_segment."New value: +"Máximo opcional de fotogramas para videos (1-20) en modo single-pass. Para tiempos específicos, prefiera video.clip_start_seconds + video.clip_duration_seconds. Para multipass, use video.max_frames_per_segment." - changed
Input schema / properties / path / descriptionPrevious value: -"Absolute path to a local file on the machine running the MCP server (e.g., C:\\\\Users\\\\User\\\\Downloads\\\\video.mp4)."New value: +"Ruta absoluta a un archivo local en la máquina donde corre el servidor MCP (por ejemplo, C:\\\\Users\\\\User\\\\Downloads\\\\video.mp4), o una URL http(s) de imagen/video/audio/PDF para descargar y analizar (hasta 64 MiB; hosts locales y redes privadas bloqueados)." - changed
Input schema / properties / paths / descriptionPrevious value: -"Absolute paths to multiple local image files (UI screenshots/photo sets). When provided, EnriVision uploads a single media-set archive for server-side batching + reduce."New value: +"Rutas absolutas a varios archivos de imagen locales o URLs http(s) (capturas de UI/sets de fotos; cada URL hasta 64 MiB). Cuando se proporcionan, EnriVision sube un único archivo de conjunto para procesamiento por lotes y reducción del lado servidor." - changed
Input schema / properties / question / descriptionPrevious value: -"Optional explicit question to answer about the file."New value: +"Pregunta explícita opcional que responder sobre el archivo." - added
Input schema / properties / regionAdded value: +{ + "description": "Región relativa de la IMAGEN original para analizar a resolución nativa (zoom). Coordenadas entre 0 y 1; (0,0) es la esquina superior izquierda. Use las cajas devueltas en 'elements' de un análisis previo de la misma imagen: NUNCA invente coordenadas. Ideal para leer texto pequeño (labels, código) que en la imagen completa comprimida resulta ilegible. Sólo imágenes (path, no paths).", + "properties": { + "height": { + "description": "Alto relativo (1 = alto completo).", + "type": "number" + }, + "width": { + "description": "Ancho relativo (1 = ancho completo).", + "type": "number" + }, + "x": { + "description": "Coordenada horizontal relativa de la esquina superior izquierda (0 = borde izquierdo).", + "type": "number" + }, + "y": { + "description": "Coordenada vertical relativa de la esquina superior izquierda (0 = borde superior).", + "type": "number" + } + }, + "required": [ + "x", + "y", + "width", + "height" + ], + "type": "object" +} - changed
Input schema / properties / transcribe / descriptionPrevious value: -"Optional override to enable/disable audio transcription for videos."New value: +"Sobreescritura opcional para activar/desactivar la transcripción de audio en videos." - changed
Input schema / properties / transcription_language / descriptionPrevious value: -"Optional Whisper language hint for audio/video transcription (e.g., 'auto', 'es', 'en')."New value: +"Pista opcional de idioma para la transcripción de audio/video (por ejemplo, 'auto', 'es', 'en')." - changed
Input schema / properties / video / descriptionPrevious value: -"Optional video multipass tuning. Used only when analyzing videos."New value: +"Ajuste opcional de multipass para video. Se usa sólo al analizar videos." - changed
Input schema / properties / video / properties / clip_duration_seconds / descriptionPrevious value: -"Optional clip duration in seconds for time-targeted video analysis."New value: +"Duración opcional del clip en segundos para análisis de video dirigido a un tiempo." - changed
Input schema / properties / video / properties / clip_start_seconds / descriptionPrevious value: -"Optional clip start offset in seconds for time-targeted video analysis."New value: +"Offset opcional de inicio del clip en segundos para análisis de video dirigido a un tiempo." - changed
Input schema / properties / video / properties / max_frames_per_segment / descriptionPrevious value: -"Maximum frames to extract per segment."New value: +"Máximo de fotogramas a extraer por segmento." - changed
Input schema / properties / video / properties / max_segments / descriptionPrevious value: -"Maximum number of segments to analyze."New value: +"Número máximo de segmentos a analizar." - changed
Input schema / properties / video / properties / segment_seconds / descriptionPrevious value: -"Segment duration in seconds."New value: +"Duración del segmento en segundos."
1 tool update
v0.1.1- Changed
analyze_media1 field changed- added
Input schema / properties / analysis_mode / enumAdded value: +[ + "auto", + "single", + "multipass" +]
1 tool update
v0.1.0- First observed
analyze_media
TDQS
Scored across 1 tool
Only one tool exists, so there is zero risk of an agent selecting the wrong tool for a task. The single analyze_media tool unambiguously covers any media-analysis request, with no overlapping purposes to confuse.
analyze_media follows the standard verb_noun convention used across well-designed MCP servers. While a single tool offers little evidence of a broader pattern, the name is self-consistent, descriptive, and aligned with common conventions.
A single tool sits at the thin end of the calibration range, where 1-2 tools feels sparse. This is partially redeemed by the tool being a deliberately consolidated mega-tool that absorbs image, video, audio, and document analysis into one entry, but that pushes substantial complexity into a single schema.
For the server's stated purpose — media analysis — coverage is remarkably complete: static and animated images, video with shared audio timeline, transcription, PDFs, Office documents, multi-image sets, video clipping, multipass budgets, and continuation cursors for long outputs. An agent can fully accomplish the server's job with this one tool and no obvious dead ends.
Maintenance
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
The Listenetic MCP server is a remote, cloud-hosted server that enables AI assistants like ChatGPT and Claude to convert articles, documents, websites, and videos into high-quality AI-generated audio. It provides multi-format support for text and binary files, natural-sounding text-to-audio conversion using AI, and specialized processing for SSML, markup, markdown, and various media formats through three core tools: listentic_supported_mimetypes, listentic_add_content_text, and listentic_add_content_binary.
The CustomGPT.ai MCP server is a fully managed, RAG-powered endpoint that connects large language models with private knowledge bases and external data sources. It provides tools for retrieval-augmented generation queries (send_message), data ingestion (upload_file), and source listing, enabling AI agents to query private documents like PDFs with high accuracy and real-time citations.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server that lets AI assistants read and visually analyze local documents — PDFs, Excel spreadsheets, CSV files, Word documents, PowerPoint presentations, and images.425 npm43 PyPIMIT
- FlicenseBqualityDmaintenanceMCP server for analyzing local audio and video files with Google Gen AI, returning structured summaries, timelines, transcripts, and observations.11-

Augentofficial
AlicenseBqualityCmaintenanceMCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.224MIT- AlicenseNot gradedqualityDmaintenanceAn MCP server for comprehensive video analysis — AI-powered transcription, visual frame analysis, and metadata extraction from 1000+ platforms.1MIT