Skip to main content
Glama
Dosugamea

Voicevox MCP Server

by Dosugamea

Voicevox MCP Server

This is a server for using VOICEVOX compatible speech synthesis servers (AivisSpeech / VOICEVOX / COEIROINK) via MCP (Model Context Protocol). It can be used for speech synthesis in agent mode using Claude 3.7 in Cursor, etc.

Prerequisites

Windows environment

Docker environment (WSL2)

  • Docker and Docker Compose

  • WSL2

  • VOICEVOX ENGINE etc. (run locally or in Docker)

  • sudo apt install libsdl2-dev pulseaudio-utils pulseaudio -enabled Linux environment

  • Permission to access /mnt/wslg

Related MCP server: TTS-MCP

Installation and Configuration

  1. Clone the repository

git clone https://github.com/Dosugamea/voicevox-mcp-server.git
cd voicevox-mcp-server
  1. Installing dependencies

npm install
  1. Setting environment variables Create a .env file by copying .env_example and modifying the settings as needed:

VOICEVOX_API_URL=http://localhost:50021
VOICEVOX_SPEAKER_ID=1

How to do it

Execution in Windows environment

Please launch a server separately from the editor by following the steps below.

npm run build
npm start

Execution in Docker environment

No need to use an editor or any other operations. It starts in stdio mode so it cannot be executed directly.

How to set it up

When running in a Windows environment

Please add the following to mcp.json. The connection is unstable, so please reconnect if it is disconnected.

        "voicevox": {
            "url": "http://localhost:10100/sse"
        }

When running in a Docker environment

Please add the following to mcp.json. (The author's environment has not been tested.)

{
    "tools": {
        "voicevox": {
            "command": "cmd",
            "args": [
                "/c",
                "docker",
                "run",
                "-i",
                "--rm",
                "-v",
                "/mnt/wslg:/mnt/wslg",
                "-e",
                "PULSE_SERVER",
                "-e",
                "SDL_AUDIODRIVER",
                "-e",
                "VOICEVOX_API_URL",
                "-e",
                "VOICEVOX_SPEAKER_ID",
                "your-local-docker-image-name"
            ],
            "env": {
                "PULSE_SERVER": "unix:/mnt/wslg/PulseServer",
                "SDL_AUDIODRIVER": "pulseaudio",
                "VOICEVOX_API_URL": "http://host.docker.internal:50031",
                "VOICEVOX_SPEAKER_ID": "919692871"
            }
        }
    }
}

About Speaker ID

The speaker ID varies depending on the model of VOICEVOX you are using. By default, "1" (Shikoku Metan) is used. If you want to use another speaker ID, change the environment variable VOICEVOX_SPEAKER_ID .

You can check the list of speaker IDs at the /speakers endpoint of the VOICEVOX ENGINE API. Example: curl http://localhost:50021/speakers

troubleshooting

  • Connection error with VOICEVOX : Please make sure that VOICEVOX ENGINE is running and that the API URL is set correctly.

  • No sound playing : Make sure VLC is properly installed and in your path.

  • Audio output problem in Docker environment : Please check that pulseaudio is configured correctly.

Developer Information

  • To contribute to the source code, please create an issue or submit a pull request.

  • To report bugs or request features, please use the Issues feature on GitHub.

license

MIT License

A
license - permissive license
B
quality
F
maintenance

Maintenance

UpdatingMaintainers
UpdatingResponse time
Release cycle
0Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Tools

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    A Model Context Protocol server that integrates high-quality text-to-speech capabilities with Claude Desktop and other MCP-compatible clients, supporting multiple voice options and audio formats.
    17
    1
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    A Node.js server that enables AI assistants to interact with Bouyomi-chan's text-to-speech functionality through Model Context Protocol (MCP), allowing for voice reading of text with adjustable parameters.
    1
    2
    MIT

View all related MCP servers

Related MCP Connectors

  • A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…

  • MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)

  • MCP server for AI dialogue using various LLM models via AceDataCloud

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Dosugamea/voicevox-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server