Skip to main content
Glama

Edit parts of a track

edit_track

Re-roll or remove parts of an existing track — "swap the bass", "drop the vocals", "give me a different lead".

This is the tool for refining a track step by step. Three things govern how to use it in a chain:

  1. Each edit makes a NEW track. The original is never modified. Pass the id returned by the previous edit into the next one; reusing the original id silently discards everything done so far and still answers success.

  2. Each edit costs a full track off the quota. A five-step session is five tracks. Check get_capabilities before starting a long one.

  3. Replacement is best-effort. If the underlying collection has no alternative for a part at this tempo and key, replacing it removes it instead — a success response with the instrument simply gone. Offer the result for a listen rather than asserting the change landed.

One part at a time is the right granularity: it keeps each step reviewable and avoids spending a track on a combination the user did not ask for. mode cannot be changed by editing — regenerate for that.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
formatNo
bitrateNo
durationNo
track_idYesThe track to edit. In a chain of edits this is the id returned by the PREVIOUS edit, not the original — each edit makes a new track.
intensityNo
delete_stemsNo
wait_secondsNo
replace_stemsNoWhole sections to re-roll: DRUMS (drums, percs, hats, claps), BASS, LEADS (mids, leads, pads), VOCALS, FX (fx, riser, impact). Coarser than replace_instruments — use it when the user names a section rather than one part.
delete_instrumentsNo
replace_instrumentsNoSingle parts to re-roll: DRUMS, PERCS, HATS, CLAPS, BASS, MIDS, LEADS, FX, VOCALS, PADS, RISER, IMPACT. Map what the user says onto these — a melody, lead line, riff or 'that horn/synth' is LEADS; a pad, string or chord bed is PADS; a countermelody or stab sitting under the lead is MIDS; a sweep into a drop is RISER; a hit on the drop is IMPACT.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
bpmNo
keyNo
urlNoDownload link once done.
modeNo
formatNo
promptNo
statusYesdone, failed or pending.
bitrateNo
durationNo
track_idNo
importantNo
expires_atNoAfter this the audio is deleted and the URL stops working.
session_idNo
edited_fromNoThe track this one was derived from. Edits and variations never modify the original — they mint a new track, and this is the link back.
playlist_indexNo
how_it_was_servedYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: each call produces a NEW track and never mutates the original, each call burns a full track of quota, and replacement is best-effort so a 'success' can silently drop a part. This is exactly the non-obvious, chain-affecting context annotations cannot express; no contradiction with destructiveHint=false since the source track is preserved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by three numbered constraints — dense and well-structured, with no filler. It does run long for a tool description, but nearly every sentence carries actionable guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return shape needn't be covered, and the description correctly calls out the misleading success semantics. Remaining gap is the undocumented output/encoding params (format, bitrate, duration) and wait_seconds behavior, which an agent must infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Coverage is only 30% across 10 params. The description reinforces track_id chaining and part-level granularity (mapping 'bass'/'vocals' to stems), but says nothing about format, bitrate, duration, intensity, delete_* variants, or wait_seconds, leaving most params to the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Re-roll or remove parts of an existing track') and grounds it in user-level examples ('swap the bass', 'drop the vocals'). It also distinguishes itself from the sibling regenerate_similar/regenerate path with 'mode cannot be changed by editing — regenerate for that.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit chaining rules (pass the previous edit's id), a recommended granularity ('one part at a time'), a quota caution pointing at get_capabilities, and a clear exclusion for mode changes. An agent knows when to use this and when not to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources