Skip to main content
Glama

create_song

Turn lyrics and a chord progression into a UTAU .ust file plus a matching .mid and .wav backing track that starts at 0:00.

Instructions

Turn lyrics and a chord progression into a song: a UTAU .ust file plus a backing track (.mid and .wav) of the same length, so both line up at 0:00.

The vocal itself is not rendered here: open the .ust in UTAU or OpenUtau with a voicebank, and import the backing .wav alongside it.

lyrics: one sung line per text line; a blank line is a bar of rest. Write 'syl-la-ble' to force syllable breaks and 'la~' to hold a syllable. chords: bars separated by '|', e.g. 'C | Am | F G | C'. Chords in one bar share it equally, '%' repeats the previous bar. The progression loops. key: e.g. 'C', 'Bb', 'F#m'. Off-beat notes use this scale. lyric_mode: 'auto' (kana by mora; other words keep the whole word on their first note and '+' on the rest, which is what OpenUtau's English phonemizers expect), 'syllables' (word fragments like 'win' 'dow'), 'romaji' (convert romaji to hiragana for Japanese voicebanks), 'raw'. voice_range: lowest-highest note for the melody, e.g. 'A3-C5'. guide_volume: 0-1 level of a flute doubling the melody in the backing, as a pitch reference for the voice; 0 leaves it out. seed: change it to get a different melody over the same chords. soundfont: path to an .sf2 file; used if FluidSynth is installed, otherwise the built-in synth renders the backing.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
keyNoC
bassNo
nameNosong
seedNo
drumsNo
styleNoarpeggio
tempoNo
chordsYes
lyricsYes
soundfontNo
voice_dirNo
instrumentNopiano
intro_barsNo
lyric_modeNoauto
output_dirNo
voice_rangeNoA3-C5
guide_volumeNo
bars_per_lineNo
beats_per_barNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the vocal is NOT rendered, that both outputs are length-matched at 0:00, and that a soundfont is only used if FluidSynth is installed (otherwise a built-in synth renders). It omits write-location/overwrite behavior and any failure modes, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-sentence summary is front-loaded and states the core contract before the per-parameter notes, and each parameter line is compact and earns its place. It runs long overall, but the length is driven by genuinely non-obvious parameter syntax rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a file-generating tool with no output schema and no annotations, the description covers the returned artifacts, their alignment, and the external dependency (FluidSynth) plus the required follow-up step in UTAU. The main remaining gap is the undocumented half of the parameter set (tempo, style, drums, output_dir, etc.).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 19 parameters, so the description must compensate. It richly documents lyrics (syllable/held-note syntax), chords ('|', '%', looping), key, lyric_mode (with per-mode semantics), voice_range, guide_volume, seed, and soundfont — but leaves bass, drums, style, tempo, instrument, name, intro_bars, bars_per_line, beats_per_bar, voice_dir, and output_dir entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource ('Turn lyrics and a chord progression into a song') and names the concrete artifacts produced (.ust plus .mid/.wav backing of matching length). This implicitly distinguishes it from siblings lyrics_to_ust and render_backing, which each produce only one of those outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the post-generation workflow ('open the .ust in UTAU or OpenUtau with a voicebank, and import the backing .wav alongside it'), which clarifies the intended pipeline. However, it never states when to choose this tool over lyrics_to_ust or render_backing, so the alternative-selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.