Skip to main content
Glama
README.md
# voicevox-mcp

このプロジェクトは、VOICEVOXエンジンと連携して音声合成やスピーカー情報の取得ができるMCP(Model Context Protocol)サーバーです。TypeScriptで実装されており、MCP SDKを利用しています。

<a href="https://glama.ai/mcp/servers/@Yuki10Kobayashi/voicevox-mcp">
  <img width="380" height="200" src="https://glama.ai/mcp/servers/@Yuki10Kobayashi/voicevox-mcp/badge" alt="VOICEVOX Server MCP server" />
</a>

# 機能
- VOICEVOXエンジンのスピーカー情報取得(/speakers)
- 指定したスピーカーでテキストを音声合成し、ローカルで再生(/speak)
  - Macのみ対応

# セットアップ

## VOICEVOXエンジンの起動(Docker推奨)

```sh
docker compose up -d
```

これで localhost:50021 でVOICEVOXエンジンが起動します。


## 依存パッケージのインストール & ビルド

```sh
npm install
npm run build 
```



# 使い方

## Cursorの設定例

```.cursor/mcp.json
{
  "mcpServers": {
    "voicevox-mcp": {
      "command": "node",
      "args": ["${Path to Repository}/dist/index.js"],
      "env": {
        "SPEAKER_ID": 8,
        "SPEED_SCALE": 1.2,
        "VOICEVOX_API_URL": "http://localhost:50021" 
      }
    }
  }
}
```

VOICEVOX_API_URLは必要に応じて設定


- MCPクライアントから speakers ツールでスピーカー一覧を取得できます。
- speak ツールでテキストを音声合成し、ローカルで再生できます(afplayコマンドを使用しているため、Mac環境推奨)。

主な依存パッケージ

- `@modelcontextprotocol/sdk`
- `zod`
- `typescript`


# 注意事項

- 今後改善
  - VOICEVOXエンジンが localhost:50021 で動作していないと音声合成は利用できません。
  - Mac以外の環境では afplay の部分を適宜変更してください。


# ライセンス

MIT License

TDQS

D1.8/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: 'speak' implies text-to-speech synthesis, while 'speakers' likely lists available voice options. There is no overlap or ambiguity between them.

Naming Consistency5/5

Both tools use simple, consistent naming: 'speak' and 'speakers' are both lowercase nouns, with 'speak' as a verb-like noun and 'speakers' as a plural noun. The pattern is uniform and predictable.

Tool Count2/5

With only two tools, the server feels thin for a VOICEVOX text-to-speech domain. Expected operations like adjusting voice parameters, controlling playback, or managing audio output are missing, making the set under-scoped.

Completeness2/5

The tool surface is severely incomplete for a VOICEVOX server. Core functionalities such as voice customization, audio format settings, playback control, or synthesis status checks are absent, leaving significant gaps that will hinder agent workflows.

Maintenance

ActivityInactive
ResponsivenessNo issues