Bouyomi-chan MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Bouyomi-chan MCP Serverread this text aloud with a female voice at normal speed"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Boyomi-chan MCP Server (Node.js version)
This is a server that uses the Model Context Protocol (MCP) to provide text-to-speech functionality using Bokuyomi-chan (a slow voice) to AI assistants. It is implemented in Node.js/TypeScript.
overview
This server is an MCP server that allows AI assistants such as Claude to use Boyomi-chan.
Related MCP server: MCP Simple AivisSpeech
function
Text to speech
Select voice type (female, male, etc.)
Volume adjustment
Adjustable speech speed
Pitch Adjustment
Prerequisites
Node.js 16 or higher
npm 7 or higher
Boyomi-chan must be installed.
The HTTP link for Boyomi-chan is running on port 50080.
How to install
Clone this repository:
git clone https://github.com/uraoz/bouyomichan-mcp-nodejs.git
cd bouyomichan-mcp-nodejsInstall the dependencies:
npm installWhich compiles:
npm run buildHow to use
Starting the Server
npm startIntegration with Claude for Desktop
To work with Claude for Desktop you need to edit the configuration file:
Open the Claude for Desktop configuration file:
MacOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
Add the following content (replace the paths with the actual file paths):
{
"mcpServers": {
"bouyomi": {
"command": "node",
"args": [
"/絶対パス/bouyomichan-mcp-nodejs/build/index.js"
]
}
}
}Restart Claude for Desktop.
Usage Example
Claude for Desktop will read text aloud to you by:
Read out "Hello, World"
A male voice reads out "This is a test."
Speed up and read "I'm in a hurry"
Parameter Description
Parameters | explanation | Default value | Scope |
text | Read text | Required | Any text |
voice | Audio Type | 0 (1 female) | 0: Female 1, 1: Male 1, 2: Female 2, ... |
volume | volume | -1 (default) | -1: default, 0-100: volume level |
speed | speed | -1 (default) | -1: default, 50-200: speed level |
tone | Pitch | -1 (default) | -1: default, 50-200: pitch level |
license
MIT
Available Tools
1 toolread_textC
テキストを棒読みちゃんで読み上げます
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the behavioral trait of using a monotone voice (棒読みちゃん), which adds some context beyond basic functionality. However, it doesn't disclose other important behaviors like whether it requires internet access, has rate limits, handles different languages, or what happens on errors. The description is minimal and lacks comprehensive behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence in Japanese that directly states the tool's function. It's front-loaded with the core action and includes a stylistic detail (monotone voice). There's no wasted text, but it could be slightly more informative without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is somewhat complete but has gaps. It explains what the tool does but lacks details on output format (e.g., audio file, stream), error handling, or performance characteristics. For a text-to-speech tool, more context on behavior and results would be helpful, though the low complexity mitigates this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the inputs. The description doesn't need to add parameter information, and it doesn't contradict the schema. Baseline is 4 for zero parameters, as the description appropriately focuses on functionality rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool's purpose as 'reading text aloud with a monotone voice' (棒読みちゃん), which is clear but somewhat vague. It specifies the action (read aloud) and the style (monotone), but doesn't distinguish from siblings since none exist. The description isn't tautological but lacks specific details about the resource or output format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description implies usage for text-to-speech with a monotone voice, but there's no mention of prerequisites, limitations, or context for when this specific style is appropriate. With no sibling tools, differentiation isn't needed, but general usage context is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
read_text
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools. The tool 'read_text' has a singular, clear purpose that cannot be confused with any other tool in this set.
A single tool inherently exhibits perfect naming consistency, as there are no other tools to compare it against. The name 'read_text' follows a clear verb_noun pattern, which is consistent with itself.
A single tool for a text-to-speech server is too minimal for practical use. While it covers the core functionality, typical MCP servers benefit from additional tools (e.g., for configuration, status checks, or voice control), making this count feel thin and limiting for agent interactions.
The tool 'read_text' provides the essential action for a text-to-speech server, but there are notable gaps. Missing operations might include stopping speech, adjusting speed or volume, checking status, or managing voice settings, which could hinder agent workflows in more complex scenarios.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for AI dialogue using various LLM models via AceDataCloud
AI voice generation: text-to-speech and voice cloning from any MCP client.
MCP server for Text-to-Speech
Related MCP Servers
- AlicenseBqualityFmaintenanceA server that enables Claude 3.7 and other AI agents to access VOICEVOX-compatible speech synthesis engines (AivisSpeech, VOICEVOX, COEIROINK) through the Model Context Protocol.112MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that integrates with AivisSpeech to enable AI assistants to convert text to natural-sounding Japanese speech with customizable voice parameters.1358Apache 2.0
- AlicenseAqualityBmaintenanceA text-to-speech MCP server that enables AI assistants to speak using the VOICEVOX engine with support for multi-character conversations. It features queue management, low-latency streaming via FFplay, and cross-platform playback across Windows, macOS, and Linux.714916ISC
- FlicenseAqualityDmaintenanceAn MCP server that enables text-to-speech generation and phonetic kana conversion using VOICEROID2 via voiceroid_daemon. It supports customizable voice parameters and provides cross-platform audio playback for synthesized speech.3-