Skip to main content
Glama

fablecut_auto_duck

Lower music under dialogue by detecting voice activity and writing duck keyframes on music clips, with adjustable dip, attack, release, and threshold.

Instructions

Duck music under dialogue: finds where the voice tracks have sound (RMS above threshold, short gaps bridged so the music doesn't pump between words) and writes duck keyframes (dB) on the given music clips — a dip of amount dB ramping down attack s before speech and back up release s after. The duck multiplies the clip's volume, so its own level, volume keyframes and fades are untouched, and re-running replaces the previous dips (amount 0 clears them). Voice = every audio clip on under tracks (default: all audio tracks the music isn't on). Linked stems of a music clip get the same keys. The same pipeline as the editor's Auto-duck. Needs ffmpeg + ffprobe on PATH. Refuses locked clips unless force:true.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
forceNoAlso change clips the user locked (only when they asked)
underNoAudio tracks holding the voice, e.g. ["A1"] (default: every other audio track)
amountNoDip depth in dB, -40…0 (default -12; 0 removes the ducking)
attackNoSeconds to ramp down before speech (default 0.3)
clipIdsYesThe music / bed clips to duck
releaseNoSeconds to ramp back up after speech (default 0.6)
thresholdNoVoice detection level in dBFS (default -40; raise it if room noise triggers ducks)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.10.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it is a mutation (writes keyframes), it multiplies clip volume so level/keyframes/fades are untouched, re-running replaces previous dips, linked stems get the same keys, requires ffmpeg+ffprobe on PATH, and refuses locked clips unless force. These are exactly the side-effect and prerequisite facts an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and free of filler, but very dense with long multi-clause sentences (em dashes, parentheticals) that pack several behaviors into each line. Efficient rather than wasteful, though slightly heavy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation with no output schema, it covers what the agent needs: side effects, idempotency, dependency requirements, lock/force behavior, stem propagation, and what it writes. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema — the interpolation/ramp semantics of attack and release, the gap-bridging rationale behind threshold, the defaulting logic of `under` ('every audio track the music isn't on'), and that amount 0 removes ducking.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Duck music under dialogue') plus the exact mechanism (finds voice via RMS threshold, writes duck keyframes in dB on the music clips). It clearly separates this from sibling audio tools like normalize_audio or denoise by naming the ducking/auto-duck pipeline explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when this applies (ducking music under voice) and gives operational triggers: raise threshold if room noise triggers ducks, amount 0 clears, force to touch locked clips. It does not explicitly name sibling alternatives to prefer in other situations, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.