utau-lyrics
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| UTAU_LYRICS_SOUNDFONT | No | Path to a .sf2 SoundFont file for FluidSynth; if missing, the built-in synth is used. Can also be passed as soundfont. | |
| UTAU_LYRICS_OUTPUT_DIR | No | Directory where generated files go; defaults to utau-lyrics-mcp-output in your user folder unless output_dir is passed. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| create_songA | Turn lyrics and a chord progression into a song: a UTAU .ust file plus a backing track (.mid and .wav) of the same length, so both line up at 0:00. The vocal itself is not rendered here: open the .ust in UTAU or OpenUtau with a voicebank, and import the backing .wav alongside it. lyrics: one sung line per text line; a blank line is a bar of rest. Write 'syl-la-ble' to force syllable breaks and 'la~' to hold a syllable. chords: bars separated by '|', e.g. 'C | Am | F G | C'. Chords in one bar share it equally, '%' repeats the previous bar. The progression loops. key: e.g. 'C', 'Bb', 'F#m'. Off-beat notes use this scale. lyric_mode: 'auto' (kana by mora; other words keep the whole word on their first note and '+' on the rest, which is what OpenUtau's English phonemizers expect), 'syllables' (word fragments like 'win' 'dow'), 'romaji' (convert romaji to hiragana for Japanese voicebanks), 'raw'. voice_range: lowest-highest note for the melody, e.g. 'A3-C5'. guide_volume: 0-1 level of a flute doubling the melody in the backing, as a pitch reference for the voice; 0 leaves it out. seed: change it to get a different melody over the same chords. soundfont: path to an .sf2 file; used if FluidSynth is installed, otherwise the built-in synth renders the backing. |
| lyrics_to_ustB | Write only the UTAU .ust file (melody fitted to the chords), no backing. Arguments are the same as create_song. |
| render_backingC | Render only a backing track (.mid and .wav) from a chord progression. bars defaults to one pass through the progression. |
| preview_syllablesB | Show how each lyric line will be split into notes, before making a song. Notes are separated by spaces; '+' continues the word before it and '(xN)' marks a held note. |
| list_optionsC | Instruments, accompaniment styles, lyric modes and chord qualities accepted. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
create_song, lyrics_to_ust, and render_backing have intentional overlap (full song vs UST-only vs backing-only), but descriptions clearly state output boundaries. list_options and preview_syllables are distinct and unlikely to be confused.
Four tools follow a verb_noun snake_case pattern (list_options, create_song, render_backing, preview_syllables). lyrics_to_ust is noun_to_noun but stays in the same snake_case style and is unambiguous.
Five tools are well-scoped: option discovery, full song generation, two focused component generators, and a preview utility. No tool feels redundant or missing for the stated purpose.
The set covers option discovery, full song generation, UST-only generation, backing-only generation, and syllable preview. Vocal rendering is explicitly out of scope, so the missing capability is not a gap.