Skip to main content
Glama
hermoso-ai

Hermoso

Official

Clip a long video

clip_video

Cut one long video into ranked, ready-to-post short clips with burned captions and subject-aware vertical reframing, using transcription to pick self-contained moments.

Instructions

Cut ONE long video into several RANKED, ready-to-post short clips (podcast, webinar, interview, talk, long ad cut -> Reels/Shorts/TikTok). Transcribes with timestamps, picks the strongest SELF-CONTAINED moments, then cuts + reframes each with ffmpeg — no video model renders anything, so it is fast and cheap. THE VERTICAL REFRAME IS SUBJECT-AWARE: one cheap vision call per clip (billed as its own event) picks a SINGLE crop offset held for the whole clip, so a speaker sitting camera-left is not cropped out and the framing never drifts inside a clip; with nothing to discard or no single subject it stays dead centre — read reframedToSubject and each clip's reframeWhy back rather than assuming either way. ACCEPTS: a YouTube link (or Vimeo / Loom / Dailymotion / Streamable / Rumble / Wistia / Twitch / TED), a direct https .mp4/.mov/.webm, or a Hermoso /generated/ URL (upload_file turns a local file into one). NOT supported: TikTok / Instagram / Facebook links, and anything age-restricted, private, members-only, geo-blocked or still LIVE — those fail fast with the real reason and are fully refunded, so ask for a direct file or an upload rather than retrying. Source ~15s to ~600MB; only the first ~40 minutes is analysed (truncated:true says so). Cost: a ~7-credit hold settled to the exact transcription + encode cost, plus the clip-selection model's tokens as their own small event. RETURNS clips[] — each its OWN served mp4 URL, title, hook, ready-to-post caption, 0-100 score and source timecode. SUBTITLES ARE BURNED IN BY DEFAULT (slim white CAPS, thin black outline, bottom safe band, no box) because short-form is watched on mute; captions:false for clean footage. TIMING IS APPROXIMATE, NOT WORD-LEVEL — cues follow the transcript's per-sentence timestamps, split by character count; never promise frame-accurate sync. captionsBurned counts the clips that really carry a burned track and captionNote says why any are bare.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
countNohow many clips to cut, 1-8 (default 4)
videoYesthe long video to clip — a YouTube/Vimeo/Loom/Dailymotion/Streamable/Rumble/Wistia/Twitch/TED watch URL, a direct https .mp4/.mov/.webm, or a Hermoso /generated/ URL
captionsNoburn subtitles into every clip. DEFAULT TRUE — a clip cut from a podcast or a talk is watched on mute, and the words are the product. Set false for clean footage. A clip whose window carries no readable speech is delivered bare rather than captioned with a guess, and the result says which.
aspectRatioNoclip shape: '9:16' (default), '1:1', '16:9', any 'W:H', or 'keep' for the source framing

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.1.285
    • changedInput schema / properties / aspectRatio / description
      Previous value: -"clip shape — '9:16' (default) vertical for Reels/Shorts/TikTok; 'keep' leaves the source framing untouched"New value: +"clip shape: '9:16' (default), '1:1', '16:9', any 'W:H', or 'keep' for the source framing"
    • removedInput schema / properties / aspectRatio / enum
      Removed value: -[
      -  "9:16",
      -  "1:1",
      -  "16:9",
      -  "keep"
      -]
  2. Addedv0.1.161

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnly=false/openWorld/idempotent=false/destructive=false. The description adds real behavioral context the annotations cannot: the ~7-credit hold settled to actual cost plus a separate vision-call event, fast/cheap because no video model renders, only first ~40 min analysed with a truncated flag, failed/unsupported inputs 'fail fast... and are fully refunded', and captions burned in by default. That is exceptional depth beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely information-dense and mostly front-loaded, but the wall of CAPS-emphasised clauses, parenthetical engine lists, and refund/truncation/caption caveats make it hard to scan; several sentences are doing three jobs at once. Some of this could be trimmed without losing the substance, so it lands at adequate-but-overstuffed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A no-output-schema mutation-ish tool that nonetheless explains inputs accepted/rejected, cost model and refund behavior, truncation, reframe logic (subject-aware single-offset crop, centre fallback, reframedToSubject/reframeWhy fields), captions default, and return shape (clips[] with URL/title/hook/caption/score/timecode). Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description still adds real value: it explains the default and rationale for captions:false, notes the SUBTITLES default in prose, and mentions captionsBurned/captionNote outcomes tied to the captions parameter. It adds meaning beyond the schema rather than repeating it, warranting a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('cut ONE long video into several RANKED, ready-to-post short clips') and resource, and the parenthetical domain list (podcast/webinar/interview/talk/ad -> Reels/Shorts/TikTok) pins the exact use case. A reader can distinguish it from reframe_video, edit_video, and stitch_video without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit accept/reject conditions ('YouTube link... or direct https .mp4', 'NOT supported: TikTok / Instagram / Facebook links, and anything age-restricted, private...'), including the correct fallback ('ask for a direct file or an upload rather than retrying'). It does not name sibling alternatives (e.g. reframe_video for an existing short, edit_video) as when-to-use forks, so it stops just short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools