Skip to main content
Glama
kvoltmer

Audionaut MCP Server

Separate stems

separate_stems

Splits audio clips into Drums, Bass, Other, and Vocals stems using the Demucs source separator, keeping each stem aligned with the original clip.

Instructions

Splits a clip into Drums/Bass/Other/Vocals tracks with the Demucs (htdemucs) source separator, aligned with the source clip. Slow: minutes for a full song, CPU only. Needs the model weights, which the Audionaut app downloads once (Settings > Separation); fails with model_missing until then, and with demucs_unavailable in builds without Demucs.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
clipNoPlaylist item id (default: the track's first clip)
trackNoTrack id (default 0; imported audio lands on a new track)
projectYesPath to the .audium project package (absolute paths recommended)
threadsNoParallel segments (default: physical cores)
mute_sourceNoMute the source track's channels (default true)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses performance characteristics (slow, CPU only), a prerequisite (downloaded model weights), and two precise failure modes (model_missing, demucs_unavailable) with their causes. It also notes output alignment with the source clip.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and output, then adds the cost and failure conditions in two dense sentences. No filler or repetition; every clause adds operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations or output schema, the description covers the operation, its results, runtime cost, prerequisites, and error codes. An agent has everything needed to call it and to interpret likely failures.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds context about what the separation produces but no syntax or default details beyond what the schema already documents for clip, track, project, threads, and mute_source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (splits) and resource (clip), names the exact output tracks (Drums/Bass/Other/Vocals), and identifies the model (Demucs htdemucs). It is clearly distinct from siblings like split and analyze.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong context: it is slow (minutes for a full song, CPU only) and requires model weights downloaded via Settings > Separation. It does not explicitly compare against alternatives like split, but the performance and prerequisite cues are enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.