utau-lyrics
# utau-lyrics-mcp
Test version. Give it lyrics and a chord progression, and it writes:
- a `.ust` file for UTAU or OpenUtau, one note per syllable, with pitches chosen to fit the chords
- a backing track of the same length as `.mid` and `.wav`, so the two line up at 0:00
It runs as an MCP server (for Claude or any other MCP client), as an OpenUtau or UTAU plugin, or from the command line.
It does not render the voice. You still open the `.ust` in UTAU or OpenUtau with a voicebank and import the backing `.wav` next to it.
## Install
Python 3.10 or newer.
```bash
pip install -e .
```
## Try it without an MCP client
```bash
python -m utau_lyrics_mcp song examples/lyrics_en.txt --chords "C | Am | F G | C" --output-dir output
```
That writes `output/lyrics_en.ust`, `output/lyrics_en_backing.mid` and `output/lyrics_en_backing.wav`, and prints the melody it picked.
A Japanese voicebank needs kana lyrics. Write romaji and pass `--lyric-mode romaji`:
```bash
python -m utau_lyrics_mcp song examples/lyrics_romaji.txt --chords "Am | F | C | G" --key Am --lyric-mode romaji --instrument guitar --style strum --drums --output-dir output
```
## Use it as an MCP server
Claude Code:
```bash
claude mcp add utau-lyrics -- python -m utau_lyrics_mcp
```
Claude Desktop, in `claude_desktop_config.json`:
```json
{
"mcpServers": {
"utau-lyrics": { "command": "python", "args": ["-m", "utau_lyrics_mcp"] }
}
}
```
Files go to `utau-lyrics-mcp-output` in your user folder unless you pass `output_dir` or set `UTAU_LYRICS_OUTPUT_DIR`.
| Tool | What it does |
| --- | --- |
| `create_song` | Lyrics and chords in, `.ust` plus backing `.mid` and `.wav` out |
| `lyrics_to_ust` | The `.ust` only |
| `render_backing` | A backing track from chords alone |
| `preview_syllables` | Shows how each line will be split into notes |
| `list_options` | Instruments, styles, lyric modes, chord qualities |
## Writing lyrics and chords
Each line of text is one sung line and gets two bars by default. A line with too many syllables takes more bars. A blank line is one bar of rest. The song starts with a one-bar intro and ends with one bar on the first chord of the progression.
English words are counted in syllables, one note each. The whole word goes on the first note and each later note gets `+`, so "window" becomes `window` `+`. OpenUtau's English phonemizers read that as "spread this word over these notes". For word fragments instead (`win` `dow`), use `lyric_mode="syllables"`.
The syllable count is a rough guess based on vowel groups. It gets "window" and "little" right and counts "quiet" as one. Fix a word by hyphenating it yourself: `qui-et`, `beau-ti-ful`. Add `~` to hold a syllable longer: `slow~`. Run `preview_syllables` first to see the split.
Kana is split by mora. `ー` and `っ` lengthen the note before them.
Chords are bars separated by `|`. Chords inside one bar share it equally, and `%` repeats the previous bar:
```
C | Am | F G | %
```
The progression loops until the lyrics run out. Supported qualities: major, `m`, `5`, `dim`, `aug`, `sus2`, `sus4`, `6`, `m6`, `7`, `maj7`, `m7`, `m7b5`, `dim7`, `9`, `add9`, plus slash bass like `D/F#`.
## How the melody is chosen
Notes that start on a beat, and the last note of every line, use a tone from the chord playing at that moment. Notes between beats can use any note of the key. The picker prefers small steps from the previous note and stays inside `voice_range` (default `A3-C5`). The last note of the song lands on the tonic when the final chord contains it.
The same input always gives the same melody. Change `seed` for a different one.
## The backing track and the voice
The backing has a chord part (`block`, `strum` or `arpeggio`), a bass, optional drums, and a flute that doubles the vocal melody quietly. The flute is there as a pitch reference while you tune the voice. Set `guide_volume` to `0` for a backing without it.
Instruments: `piano`, `epiano`, `organ`, `guitar`, `strings`, `pad`.
There are two renderers.
**Built-in synth.** Used by default. It builds each instrument from sine partials in numpy, so it needs no downloads and sounds like a synth.
**FluidSynth with a SoundFont.** For sampled instruments, install [FluidSynth](https://www.fluidsynth.org/) and get a free SoundFont such as FluidR3_GM (MIT), MuseScore_General (MIT) or GeneralUser GS (its own permissive licence). Then pass `soundfont` or set `UTAU_LYRICS_SOUNDFONT` to the `.sf2` path. The instrument names map to General MIDI programs. If FluidSynth or the file is missing, the built-in synth takes over.
The `.mid` is always written, so you can also load it into any DAW and pick your own instruments.
## Getting it to sing in OpenUtau
Open the `.ust`, pick a singer for the track, then pick a phonemizer that matches both the voicebank and the lyric language. A mismatch gives silence or a hum, and the log fills with "phonemizer error" lines.
- English classic voicebank: use the phonemizer its readme names, usually `EN X-SAMPA`, `EN ARPA+` or `EN VCCV`.
- Japanese voicebank: write the lyrics in kana or use `lyric_mode="romaji"`, with a `JA` phonemizer. A Japanese bank cannot sing English words.
- The `DiffSinger` phonemizers only work with DiffSinger voicebanks.
Then import the backing `.wav` on a second track.
## OpenUtau and UTAU plugin
Windows only for now.
```bash
python -m utau_lyrics_mcp install-plugin
```
This copies the plugin into OpenUtau's `Plugins` folder (`Documents/OpenUtau/Plugins/utau-lyrics-mcp`) and writes a `run.bat` that calls the Python you ran the command with. For classic UTAU, or an OpenUtau data folder somewhere else, pass the folder: `--dir "C:\path o\plugins"`. Restart the editor afterwards.
In OpenUtau, select notes in the piano roll and pick "Fit notes to chords + backing track" from the legacy plugin menu. The plugin keeps your lyrics and note lengths, changes the pitches to fit the chords in `settings.ini`, and writes `plugin_backing.wav` and `.mid` for the selection to `utau-lyrics-mcp-output` in your user folder. Chords start at the first selected note.
Edit `settings.ini` in the installed folder to change chords, key, range, seed and instrument. Reinstalling keeps your `settings.ini`.
## Not tested yet
- The plugin has been run through its `run.bat` on a temp file in OpenUtau's format, but not from inside OpenUtau or UTAU.
- The FluidSynth path has not been run on a real install.
- Rhythm is an even eighth-note grid. There is no syncopation and no melisma.
- 4/4 is the only time signature that has been tried, though `beats_per_bar` exists.
## Tests
```bash
pip install -e ".[dev]"
pytest
```
## Licence
MIT
TDQS
Scored across 5 tools
create_song, lyrics_to_ust, and render_backing have intentional overlap (full song vs UST-only vs backing-only), but descriptions clearly state output boundaries. list_options and preview_syllables are distinct and unlikely to be confused.
Four tools follow a verb_noun snake_case pattern (list_options, create_song, render_backing, preview_syllables). lyrics_to_ust is noun_to_noun but stays in the same snake_case style and is unambiguous.
Five tools are well-scoped: option discovery, full song generation, two focused component generators, and a preview utility. No tool feels redundant or missing for the stated purpose.
The set covers option discovery, full song generation, UST-only generation, backing-only generation, and syllable preview. Vocal rendering is explicitly out of scope, so the missing capability is not a gap.