Skip to main content
Glama

utau-lyrics-mcp

Test version. Give it lyrics and a chord progression, and it writes:

  • a .ust file for UTAU or OpenUtau, one note per syllable, with pitches chosen to fit the chords

  • a backing track of the same length as .mid and .wav, so the two line up at 0:00

It runs as an MCP server (for Claude or any other MCP client), as an OpenUtau or UTAU plugin, or from the command line.

It does not render the voice. You still open the .ust in UTAU or OpenUtau with a voicebank and import the backing .wav next to it.

Install

Python 3.10 or newer.

pip install -e .

Related MCP server: midi-composer-mcp

Try it without an MCP client

python -m utau_lyrics_mcp song examples/lyrics_en.txt --chords "C | Am | F G | C" --output-dir output

That writes output/lyrics_en.ust, output/lyrics_en_backing.mid and output/lyrics_en_backing.wav, and prints the melody it picked.

A Japanese voicebank needs kana lyrics. Write romaji and pass --lyric-mode romaji:

python -m utau_lyrics_mcp song examples/lyrics_romaji.txt --chords "Am | F | C | G" --key Am --lyric-mode romaji --instrument guitar --style strum --drums --output-dir output

Use it as an MCP server

Claude Code:

claude mcp add utau-lyrics -- python -m utau_lyrics_mcp

Claude Desktop, in claude_desktop_config.json:

{
  "mcpServers": {
    "utau-lyrics": { "command": "python", "args": ["-m", "utau_lyrics_mcp"] }
  }
}

Files go to utau-lyrics-mcp-output in your user folder unless you pass output_dir or set UTAU_LYRICS_OUTPUT_DIR.

Tool

What it does

create_song

Lyrics and chords in, .ust plus backing .mid and .wav out

lyrics_to_ust

The .ust only

render_backing

A backing track from chords alone

preview_syllables

Shows how each line will be split into notes

list_options

Instruments, styles, lyric modes, chord qualities

Writing lyrics and chords

Each line of text is one sung line and gets two bars by default. A line with too many syllables takes more bars. A blank line is one bar of rest. The song starts with a one-bar intro and ends with one bar on the first chord of the progression.

English words are counted in syllables, one note each. The whole word goes on the first note and each later note gets +, so "window" becomes window +. OpenUtau's English phonemizers read that as "spread this word over these notes". For word fragments instead (win dow), use lyric_mode="syllables".

The syllable count is a rough guess based on vowel groups. It gets "window" and "little" right and counts "quiet" as one. Fix a word by hyphenating it yourself: qui-et, beau-ti-ful. Add ~ to hold a syllable longer: slow~. Run preview_syllables first to see the split.

Kana is split by mora. ー and っ lengthen the note before them.

Chords are bars separated by |. Chords inside one bar share it equally, and % repeats the previous bar:

C | Am | F G | %

The progression loops until the lyrics run out. Supported qualities: major, m, 5, dim, aug, sus2, sus4, 6, m6, 7, maj7, m7, m7b5, dim7, 9, add9, plus slash bass like D/F#.

How the melody is chosen

Notes that start on a beat, and the last note of every line, use a tone from the chord playing at that moment. Notes between beats can use any note of the key. The picker prefers small steps from the previous note and stays inside voice_range (default A3-C5). The last note of the song lands on the tonic when the final chord contains it.

The same input always gives the same melody. Change seed for a different one.

The backing track and the voice

The backing has a chord part (block, strum or arpeggio), a bass, optional drums, and a flute that doubles the vocal melody quietly. The flute is there as a pitch reference while you tune the voice. Set guide_volume to 0 for a backing without it.

Instruments: piano, epiano, organ, guitar, strings, pad.

There are two renderers.

Built-in synth. Used by default. It builds each instrument from sine partials in numpy, so it needs no downloads and sounds like a synth.

FluidSynth with a SoundFont. For sampled instruments, install FluidSynth and get a free SoundFont such as FluidR3_GM (MIT), MuseScore_General (MIT) or GeneralUser GS (its own permissive licence). Then pass soundfont or set UTAU_LYRICS_SOUNDFONT to the .sf2 path. The instrument names map to General MIDI programs. If FluidSynth or the file is missing, the built-in synth takes over.

The .mid is always written, so you can also load it into any DAW and pick your own instruments.

Getting it to sing in OpenUtau

Open the .ust, pick a singer for the track, then pick a phonemizer that matches both the voicebank and the lyric language. A mismatch gives silence or a hum, and the log fills with "phonemizer error" lines.

  • English classic voicebank: use the phonemizer its readme names, usually EN X-SAMPA, EN ARPA+ or EN VCCV.

  • Japanese voicebank: write the lyrics in kana or use lyric_mode="romaji", with a JA phonemizer. A Japanese bank cannot sing English words.

  • The DiffSinger phonemizers only work with DiffSinger voicebanks.

Then import the backing .wav on a second track.

OpenUtau and UTAU plugin

Windows only for now.

python -m utau_lyrics_mcp install-plugin

This copies the plugin into OpenUtau's Plugins folder (Documents/OpenUtau/Plugins/utau-lyrics-mcp) and writes a run.bat that calls the Python you ran the command with. For classic UTAU, or an OpenUtau data folder somewhere else, pass the folder: --dir "C:\path o\plugins". Restart the editor afterwards.

In OpenUtau, select notes in the piano roll and pick "Fit notes to chords + backing track" from the legacy plugin menu. The plugin keeps your lyrics and note lengths, changes the pitches to fit the chords in settings.ini, and writes plugin_backing.wav and .mid for the selection to utau-lyrics-mcp-output in your user folder. Chords start at the first selected note.

Edit settings.ini in the installed folder to change chords, key, range, seed and instrument. Reinstalling keeps your settings.ini.

Not tested yet

  • The plugin has been run through its run.bat on a temp file in OpenUtau's format, but not from inside OpenUtau or UTAU.

  • The FluidSynth path has not been run on a real install.

  • Rhythm is an even eighth-note grid. There is no syncopation and no melisma.

  • 4/4 is the only time signature that has been tried, though beats_per_bar exists.

Tests

pip install -e ".[dev]"
pytest

Licence

MIT

Available Tools

5 tools
create_songA

Turn lyrics and a chord progression into a song: a UTAU .ust file plus a backing track (.mid and .wav) of the same length, so both line up at 0:00.

The vocal itself is not rendered here: open the .ust in UTAU or OpenUtau with a voicebank, and import the backing .wav alongside it.

lyrics: one sung line per text line; a blank line is a bar of rest. Write 'syl-la-ble' to force syllable breaks and 'la~' to hold a syllable. chords: bars separated by '|', e.g. 'C | Am | F G | C'. Chords in one bar share it equally, '%' repeats the previous bar. The progression loops. key: e.g. 'C', 'Bb', 'F#m'. Off-beat notes use this scale. lyric_mode: 'auto' (kana by mora; other words keep the whole word on their first note and '+' on the rest, which is what OpenUtau's English phonemizers expect), 'syllables' (word fragments like 'win' 'dow'), 'romaji' (convert romaji to hiragana for Japanese voicebanks), 'raw'. voice_range: lowest-highest note for the melody, e.g. 'A3-C5'. guide_volume: 0-1 level of a flute doubling the melody in the backing, as a pitch reference for the voice; 0 leaves it out. seed: change it to get a different melody over the same chords. soundfont: path to an .sf2 file; used if FluidSynth is installed, otherwise the built-in synth renders the backing.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoC
bassNo
nameNosong
seedNo
drumsNo
styleNoarpeggio
tempoNo
chordsYes
lyricsYes
soundfontNo
voice_dirNo
instrumentNopiano
intro_barsNo
lyric_modeNoauto
output_dirNo
voice_rangeNoA3-C5
guide_volumeNo
bars_per_lineNo
beats_per_barNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the vocal is NOT rendered, that both outputs are length-matched at 0:00, and that a soundfont is only used if FluidSynth is installed (otherwise a built-in synth renders). It omits write-location/overwrite behavior and any failure modes, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-sentence summary is front-loaded and states the core contract before the per-parameter notes, and each parameter line is compact and earns its place. It runs long overall, but the length is driven by genuinely non-obvious parameter syntax rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a file-generating tool with no output schema and no annotations, the description covers the returned artifacts, their alignment, and the external dependency (FluidSynth) plus the required follow-up step in UTAU. The main remaining gap is the undocumented half of the parameter set (tempo, style, drums, output_dir, etc.).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 19 parameters, so the description must compensate. It richly documents lyrics (syllable/held-note syntax), chords ('|', '%', looping), key, lyric_mode (with per-mode semantics), voice_range, guide_volume, seed, and soundfont — but leaves bass, drums, style, tempo, instrument, name, intro_bars, bars_per_line, beats_per_bar, voice_dir, and output_dir entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource ('Turn lyrics and a chord progression into a song') and names the concrete artifacts produced (.ust plus .mid/.wav backing of matching length). This implicitly distinguishes it from siblings lyrics_to_ust and render_backing, which each produce only one of those outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the post-generation workflow ('open the .ust in UTAU or OpenUtau with a voicebank, and import the backing .wav alongside it'), which clarifies the intended pipeline. However, it never states when to choose this tool over lyrics_to_ust or render_backing, so the alternative-selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_optionsC

Instruments, accompaniment styles, lyric modes and chord qualities accepted.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it discloses no behavior: no indication that it is a read-only enumeration, what the return format is, or whether results are static. It only lists categories of options.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler, front-loading the four option categories. It is efficient, though as a noun phrase fragment it lacks a clear subject/verb structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should at least sketch the return shape (e.g., a mapping of category to accepted values). It enumerates the categories it will return, which is adequate for the caller to know what to expect, but leaves the structure unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to disambiguate and the baseline is 4. The description's mention of accepted categories is the only semantic content needed for a parameterless lookup.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the option categories (instruments, accompaniment styles, lyric modes, chord qualities) but never states a verb or that it returns the accepted values. An agent can infer enumeration from the name 'list_options', but the fragment alone doesn't distinguish it as a lookup versus a mutation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. It doesn't say to call this before create_song to discover valid option values, nor mention any alternative or prerequisite, leaving the agent to guess its role relative to create_song and render_backing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lyrics_to_ustB

Write only the UTAU .ust file (melody fitted to the chords), no backing. Arguments are the same as create_song.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoC
nameNosong
seedNo
tempoNo
chordsYes
lyricsYes
voice_dirNo
intro_barsNo
lyric_modeNoauto
output_dirNo
voice_rangeNoA3-C5
bars_per_lineNo
beats_per_barNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the key trait — a file is written and no backing audio is produced — but says nothing about where the file lands, whether existing files are overwritten, or what permissions/paths are required for a write operation with 13 parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the output artifact and followed by the argument cross-reference. Nothing is padded or repeated, and the scope constraint leads.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter file-writing tool with no annotations and no output schema, the definition is far too thin. It never explains the parameters, the output location, or how the generated melody relates to the supplied chords and lyrics, leaving major gaps an agent would need filled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 13 parameters, so the schema documents names and defaults only, with no meaning for key, lyric_mode, voice_range, bars_per_line, or the rest. The description offers no per-parameter semantics and only delegates via 'same as create_song', which leaves the agent dependent on a sibling's definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and artifact: write only the UTAU .ust file, melody fitted to chords. The phrase 'no backing' distinguishes it from the sibling render_backing, so the agent can separate it from at least one alternative. It stops short of explaining what a .ust file is or how this differs from create_song beyond the missing backing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'no backing' and 'Arguments are the same as create_song' imply this is the UST-only variant of create_song, which is usable routing information. However, there is no explicit when-to-use/when-not statement, no prerequisite on voice_dir or output_dir, and no guidance on choosing this over create_song.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_syllablesB

Show how each lyric line will be split into notes, before making a song. Notes are separated by spaces; '+' continues the word before it and '(xN)' marks a held note.

ParametersJSON Schema
NameRequiredDescriptionDefault
lyricsYes
lyric_modeNoauto

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral burden. It usefully discloses the notation conventions (space-separated notes, '+' to continue a word, '(xN)' for held notes), but 'preview'/'show' only implicitly conveys that this is a non-mutating inspection step — no explicit statement that nothing is written.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the purpose before the notation details. Every clause carries information; the only minor nit is that the notation rules describe output rather than the tool's own behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value structure need not be re-explained, and the notation gloss helps interpret it. However, with no annotations and fully undocumented parameters, an agent still lacks guidance on lyric_mode and on confirming this call has no side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema itself explains neither the required 'lyrics' string nor the 'lyric_mode' option (which has a default of 'auto' but unknown accepted values). The description explains output notation but adds nothing about what these parameters mean or how lyric_mode changes behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and resource: it shows how each lyric line will be split into notes, and frames the operation as happening 'before making a song'. This clearly separates it from create_song, but it never names or contrasts with the closer sibling lyrics_to_ust, so sibling differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'before making a song' implies the intended workflow slot (a preview step preceding creation), which is useful context, but there is no explicit when-to-use/when-not guidance and no routing to alternatives like lyrics_to_ust or render_backing. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

render_backingC

Render only a backing track (.mid and .wav) from a chord progression. bars defaults to one pass through the progression.

ParametersJSON Schema
NameRequiredDescriptionDefault
barsNo
bassNo
nameNobacking
drumsNo
styleNoarpeggio
tempoNo
chordsYes
soundfontNo
instrumentNopiano
output_dirNo
beats_per_barNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the output formats and one default, but says nothing about where files are written, whether existing files are overwritten, naming behavior, or anything about a process that writes two files to disk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and output, then the one default worth calling out. No filler and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter, side-effect-producing tool with no annotations and no output schema, the description covers a small slice. It never explains the resulting files' location or how the many musical options shape the render, leaving the agent under-informed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 11 parameters. The description clarifies only one of them — that 'bars' defaults to one pass through the progression (a meaning the null default does not convey). The remaining ten parameters (style, tempo, instrument, soundfont, drums, bass, output_dir, etc.) receive no elaboration anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (render) and resource (backing track from a chord progression) plus the output artifacts (.mid and .wav). The word 'only' weakly signals contrast with a fuller-song render, but no sibling is named, so differentiation from create_song is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no alternatives are named. The adverb 'only' implies a contrast with other render paths, but the agent must guess which sibling to prefer and under what conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedcreate_song
    • First observedlist_options
    • First observedlyrics_to_ust
    • First observedpreview_syllables
    • First observedrender_backing

TDQS

B3.4/5.0

Scored across 5 tools

Disambiguation4/5

create_song, lyrics_to_ust, and render_backing have intentional overlap (full song vs UST-only vs backing-only), but descriptions clearly state output boundaries. list_options and preview_syllables are distinct and unlikely to be confused.

Naming Consistency4/5

Four tools follow a verb_noun snake_case pattern (list_options, create_song, render_backing, preview_syllables). lyrics_to_ust is noun_to_noun but stays in the same snake_case style and is unambiguous.

Tool Count5/5

Five tools are well-scoped: option discovery, full song generation, two focused component generators, and a preview utility. No tool feels redundant or missing for the stated purpose.

Completeness5/5

The set covers option discovery, full song generation, UST-only generation, backing-only generation, and syllable preview. Vocal rendering is explicitly out of scope, so the missing capability is not a gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers