ms-tts
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ms-ttsConvert 'Hello, how are you?' to French speech."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Text-to-Speech Server
Model Context Protocol server for text-to-speech synthesis using Azure Speech Services.
Features
🎵 High-Quality Speech: Azure Neural voices with natural sound
🌍 6 Languages: English, Finnish, Spanish, German, French, Swedish
🗣️ Smart Voice Selection: Auto-select optimal voices or specify manually
📊 Performance Metrics: Synthesis timing and audio stats
đź”§ MCP Compatible: Works with Claude Desktop, VS Code, other MCP clients
Related MCP server: voiceroid_daemon-mcp
Supported Voices
Language | Default Voice | Alternatives |
English (en-US) |
|
|
Finnish (fi-FI) |
|
|
Spanish (es-ES) |
|
|
German (de-DE) |
|
|
French (fr-FR) |
|
|
Swedish (sv-SE) |
|
|
Quick Start
# 1. Install
npm install
# 2. Configure (copy from parent or create .env)
cp ../.env .env
# 3. Test
npm run test:basic
# 4. Start
npm startRequired .env:
AZURE_SPEECH_KEY=your-key
AZURE_SPEECH_REGION=westeuropeMCP Integration
Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"audio-tts": {
"command": "node",
"args": ["/path/to/mcp-server/mcp-server.mjs"],
"env": {"AZURE_SPEECH_KEY": "your-key", "AZURE_SPEECH_REGION": "westeurope"}
}
}
}VS Code
Use included .vscode/mcp.json or install MCP extension.
Usage
Natural language: "Convert to Finnish speech: Hei kaikki, olen Jenny."
Direct tool call:
{
"tool": "synthesize_speech",
"parameters": {
"sentence": "Hei kaikki, olen Jenny ja puhun suomea.",
"language": "fi-FI",
"voice": "en-US-JennyMultilingualNeural"
}
}Tool Parameters
Parameter | Required | Description |
| âś… | Text to convert (1-1000 chars) |
| âś… | Language code ( |
| ❌ | Specific voice (uses language default if not specified) |
Output
Audio saved to ./audio/mcp-generated/ as:
mcp-tts-fi_FI-en-US-JennyMultilingualNeural-2025-08-17T16-30-45-123Z.wavReturns: file path, voice used, performance metrics (synthesis time, duration, etc.)
Troubleshooting
Server won't start: Check Azure credentials in .env, ensure Node.js 16+, run npm install
No audio output: Verify output directory exists, check Azure quota/billing, confirm supported language
Voice issues: Use exact voice names from table above, try language default, check Azure region support
Debug mode: DEBUG=* npm start
Requirements
Node.js 16+
Azure Speech Services API key
MCP-compatible client (Claude Desktop, VS Code with MCP extension)
Built with Model Context Protocol for universal AI integration
Available Tools
1 toolsynthesize_speechA
Convert text to speech using Microsoft Azure Speech Services. Supports multiple languages and voices.
| Name | Required | Description | Default |
|---|---|---|---|
| voice | No | Optional specific voice name. If not provided, uses the best voice for the language. | |
| language | Yes | Language code (e.g., en-US, fi-FI, es-ES, de-DE, fr-FR, sv-SE) | en-US |
| sentence | Yes | The text to convert to speech |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavior. However, it only states the basic conversion function and language/voice support, without mentioning output format, latency, quotas, or any side effects. This is a significant transparency gap, so a score of 2 is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundancy, front-loading the core action ('Convert text to speech') and adding a concise capability note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description is moderately complete: it explains the core function and the schema covers all parameters, but it omits any explanation of the return format or audio output behavior. This limits completeness, so a score of 3 is justified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters are fully described in the schema (100% coverage), so the description doesn't need to add parameter details. The mention of language/voice support in the description merely echoes the schema and adds no new semantic information. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'convert' and clearly identifies the resource ('text to speech') and implementation ('Microsoft Azure Speech Services'). It also mentions multi-language and voice support, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There are no sibling tools, so explicit alternatives aren't required. The description provides clear context for when to use the tool (when text-to-speech synthesis is needed) but doesn't include exclusions or prerequisites, so a score of 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only one tool in the server, there is no possibility of confusion or overlap. The single tool has a clear and unique purpose of converting text to speech.
The tool name 'synthesize_speech' follows a clear verb_noun pattern, which is consistent and predictable. Although there is only one tool, the naming convention is sound and would align with a well-structured set.
The server has a single tool, which feels thin for a text-to-speech service. While a minimal server could focus solely on synthesis, a typical TTS backend would benefit from additional tools such as listing available voices or managing audio outputs, making the current count borderline.
The core action of synthesizing speech is covered, but there is a notable gap in metadata discovery—agents cannot query available voices or languages programmatically. This limits the surface's ability to fully support dynamic voice selection, though the synthesis itself is functional.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Text-to-Speech
MCP server for Speech-to-Text
MCP server exposing the AceDataCloud Fish Audio API (text-to-speech with voice conditioning)
AI voice generation: text-to-speech and voice cloning from any MCP client.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server integrated with Microsoft Edge's high-quality speech synthesis capabilities, supporting multilingual speech generation, audio merging, and cloud storage.12Apache 2.0
- FlicenseAqualityDmaintenanceAn MCP server that enables text-to-speech generation and phonetic kana conversion using VOICEROID2 via voiceroid_daemon. It supports customizable voice parameters and provides cross-platform audio playback for synthesized speech.3
- FlicenseNot gradedqualityDmaintenanceAn MCP server that leverages the Microsoft Edge TTS service to provide high-quality text-to-speech capabilities across over 80 languages. It enables users to generate audio files, query available voices, and create subtitle files using natural language commands.
- AlicenseNot gradedqualityDmaintenanceAn MCP server that converts text into lifelike speech using Microsoft Edge's Text-to-Speech service, supporting customizable voice, rate, volume, and pitch.4MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/sheikkinen/ms-tts'
If you have feedback or need assistance with the MCP directory API, please join our Discord server