Skip to main content
Glama

MCP Text-to-Speech Server

Model Context Protocol server for text-to-speech synthesis using Azure Speech Services.

Features

  • 🎵 High-Quality Speech: Azure Neural voices with natural sound

  • 🌍 6 Languages: English, Finnish, Spanish, German, French, Swedish

  • 🗣️ Smart Voice Selection: Auto-select optimal voices or specify manually

  • 📊 Performance Metrics: Synthesis timing and audio stats

  • đź”§ MCP Compatible: Works with Claude Desktop, VS Code, other MCP clients

Related MCP server: voiceroid_daemon-mcp

Supported Voices

Language

Default Voice

Alternatives

English (en-US)

en-US-RyanMultilingualNeural

en-US-JennyMultilingualNeural, en-US-AndrewMultilingualNeural

Finnish (fi-FI)

en-US-RyanMultilingualNeural

en-US-JennyMultilingualNeural, fi-FI-SelmaNeural, fi-FI-NooraNeural

Spanish (es-ES)

es-ES-AlvaroNeural

es-ES-ElviraNeural

German (de-DE)

de-DE-ConradNeural

de-DE-KatjaNeural

French (fr-FR)

fr-FR-DeniseNeural

fr-FR-HenriNeural

Swedish (sv-SE)

sv-SE-MattiasNeural

sv-SE-SofieNeural

Quick Start

# 1. Install
npm install

# 2. Configure (copy from parent or create .env)
cp ../.env .env

# 3. Test
npm run test:basic

# 4. Start
npm start

Required .env:

AZURE_SPEECH_KEY=your-key
AZURE_SPEECH_REGION=westeurope

MCP Integration

Claude Desktop (claude_desktop_config.json)

{
  "mcpServers": {
    "audio-tts": {
      "command": "node",
      "args": ["/path/to/mcp-server/mcp-server.mjs"],
      "env": {"AZURE_SPEECH_KEY": "your-key", "AZURE_SPEECH_REGION": "westeurope"}
    }
  }
}

VS Code

Use included .vscode/mcp.json or install MCP extension.

Usage

Natural language: "Convert to Finnish speech: Hei kaikki, olen Jenny."

Direct tool call:

{
  "tool": "synthesize_speech",
  "parameters": {
    "sentence": "Hei kaikki, olen Jenny ja puhun suomea.",
    "language": "fi-FI", 
    "voice": "en-US-JennyMultilingualNeural"
  }
}

Tool Parameters

Parameter

Required

Description

sentence

âś…

Text to convert (1-1000 chars)

language

âś…

Language code (en-US, fi-FI, es-ES, de-DE, fr-FR, sv-SE)

voice

❌

Specific voice (uses language default if not specified)

Output

Audio saved to ./audio/mcp-generated/ as:

mcp-tts-fi_FI-en-US-JennyMultilingualNeural-2025-08-17T16-30-45-123Z.wav

Returns: file path, voice used, performance metrics (synthesis time, duration, etc.)

Troubleshooting

Server won't start: Check Azure credentials in .env, ensure Node.js 16+, run npm install

No audio output: Verify output directory exists, check Azure quota/billing, confirm supported language

Voice issues: Use exact voice names from table above, try language default, check Azure region support

Debug mode: DEBUG=* npm start

Requirements

  • Node.js 16+

  • Azure Speech Services API key

  • MCP-compatible client (Claude Desktop, VS Code with MCP extension)


Built with Model Context Protocol for universal AI integration

Available Tools

1 tool
synthesize_speechA

Convert text to speech using Microsoft Azure Speech Services. Supports multiple languages and voices.

ParametersJSON Schema
NameRequiredDescriptionDefault
voiceNoOptional specific voice name. If not provided, uses the best voice for the language.
languageYesLanguage code (e.g., en-US, fi-FI, es-ES, de-DE, fr-FR, sv-SE)en-US
sentenceYesThe text to convert to speech

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. However, it only states the basic conversion function and language/voice support, without mentioning output format, latency, quotas, or any side effects. This is a significant transparency gap, so a score of 2 is given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundancy, front-loading the core action ('Convert text to speech') and adding a concise capability note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description is moderately complete: it explains the core function and the schema covers all parameters, but it omits any explanation of the return format or audio output behavior. This limits completeness, so a score of 3 is justified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters are fully described in the schema (100% coverage), so the description doesn't need to add parameter details. The mention of language/voice support in the description merely echoes the schema and adds no new semantic information. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'convert' and clearly identifies the resource ('text to speech') and implementation ('Microsoft Azure Speech Services'). It also mentions multi-language and voice support, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are no sibling tools, so explicit alternatives aren't required. The description provides clear context for when to use the tool (when text-to-speech synthesis is needed) but doesn't include exclusions or prerequisites, so a score of 4 is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.8/5.0
Disambiguation5/5

With only one tool in the server, there is no possibility of confusion or overlap. The single tool has a clear and unique purpose of converting text to speech.

Naming Consistency5/5

The tool name 'synthesize_speech' follows a clear verb_noun pattern, which is consistent and predictable. Although there is only one tool, the naming convention is sound and would align with a well-structured set.

Tool Count3/5

The server has a single tool, which feels thin for a text-to-speech service. While a minimal server could focus solely on synthesis, a typical TTS backend would benefit from additional tools such as listing available voices or managing audio outputs, making the current count borderline.

Completeness3/5

The core action of synthesizing speech is covered, but there is a notable gap in metadata discovery—agents cannot query available voices or languages programmatically. This limits the surface's ability to fully support dynamic voice selection, though the synthesis itself is functional.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    An MCP server integrated with Microsoft Edge's high-quality speech synthesis capabilities, supporting multilingual speech generation, audio merging, and cloud storage.
    1
    2
    Apache 2.0
  • F
    license
    A
    quality
    D
    maintenance
    An MCP server that enables text-to-speech generation and phonetic kana conversion using VOICEROID2 via voiceroid_daemon. It supports customizable voice parameters and provides cross-platform audio playback for synthesized speech.
    3
  • F
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that leverages the Microsoft Edge TTS service to provide high-quality text-to-speech capabilities across over 80 languages. It enables users to generate audio files, query available voices, and create subtitle files using natural language commands.
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that converts text into lifelike speech using Microsoft Edge's Text-to-Speech service, supporting customizable voice, rate, volume, and pitch.
    4
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/sheikkinen/ms-tts'

If you have feedback or need assistance with the MCP directory API, please join our Discord server