ttsstt-mcp-server
Provides Speech-to-Text (STT) and Text-to-Speech (TTS) tools using OpenAI-compatible APIs, allowing conversion of audio files to text and text to speech with configurable models, voices, and parameters.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ttsstt-mcp-serverTranscribe the audio recording at /path/to/audio.mp3"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
TTSSTT MCP Server
A Model Context Protocol (MCP) server that provides Speech-to-Text (STT) and Text-to-Speech (TTS) tools using OpenAI-compatible APIs.
Features
STT (Speech-to-Text): Convert audio files to text using OpenAI-compatible STT APIs
TTS (Text-to-Speech): Convert text to audio files using OpenAI-compatible TTS APIs
Related MCP server: Open AI Text To Speech1 MCP Server
Quick Start
Using npx (Recommended)
npx github:ptbsare/ttsstt-mcp-serverLocal Development
# Clone the repository
git clone https://github.com/ptbsare/ttsstt-mcp-server.git
cd ttsstt-mcp-server
# Install dependencies
npm install
# Build
npm run build
# Run
npm startEnvironment Variables
STT Configuration
Variable | Description | Default |
| OpenAI-compatible STT API endpoint URL | (required) |
| API key for STT service | (required) |
| STT model name |
|
| Response format ( |
|
TTS Configuration
Variable | Description | Default |
| OpenAI-compatible TTS API endpoint URL | (required) |
| API key for TTS service | (required) |
| TTS model name ( |
|
| Voice to use ( |
|
| Speech speed (0.25 to 4.0) |
|
| Pitch adjustment |
|
| Directory to save generated audio files |
|
Alternative Environment Variables
You can also use the following prefixed versions:
OPENAI_STT_URL,OPENAI_STT_API_KEY,OPENAI_STT_MODEL,OPENAI_STT_RESPONSE_FORMATOPENAI_TTS_URL,OPENAI_TTS_API_KEY,OPENAI_TTS_MODEL,OPENAI_TTS_VOICE,OPENAI_TTS_SPEED,OPENAI_TTS_PITCH,OPENAI_TTS_OUTPUT_DIR
Usage
Configure MCP Client
Add to your MCP client configuration (e.g., Claude Desktop, Cursor, VS Code):
{
"mcpServers": {
"ttsstt": {
"command": "npx",
"args": ["github:ptbsare/ttsstt-mcp-server"],
"env": {
"STT_URL": "http://192.168.195.210:10500/v1/audio/transcriptions",
"STT_API_KEY": "your-stt-api-key",
"STT_MODEL": "whisper-1",
"TTS_URL": "http://192.168.195.210:10500/v1/audio/speech",
"TTS_API_KEY": "your-tts-api-key",
"TTS_MODEL": "tts-1",
"TTS_VOICE": "alloy",
"TTS_OUTPUT_DIR": "./audio_output"
}
}
}
}STT Tool
Convert audio to text:
{
"tool": "stt",
"arguments": {
"audio_path": "/path/to/audio.mp3",
"language": "zh"
}
}Parameters:
audio_path(required): Path to the audio file to transcribelanguage(optional): Language code for transcription (e.g., 'en', 'zh', 'ja')
TTS Tool
Convert text to audio:
{
"tool": "tts",
"arguments": {
"text": "Hello, world! This is a test.",
"voice": "nova",
"speed": 1.0
}
}Parameters:
text(required): Text content to convert to speechvoice(optional): Voice to use (alloy, echo, fable, onyx, nova, shimmer)speed(optional): Speech speed (0.25 to 4.0)pitch(optional): Pitch adjustment
Returns: Path to the generated audio file
API Compatibility
This server is compatible with OpenAI's Audio API format:
STT: Uses
/v1/audio/transcriptionsendpoint with multipart/form-dataTTS: Uses
/v1/audio/speechendpoint with JSON body
Testing
Test server with the provided test endpoint:
STT_URL=http://192.168.195.210:10500/v1/audio/transcriptions \
STT_API_KEY=test \
TTS_URL=http://192.168.195.210:10500/v1/audio/speech \
TTS_API_KEY=test \
TTS_OUTPUT_DIR=/tmp/tts_output \
npm run build && npm startProject Structure
ttsstt-mcp-server/
├── src/
│ └── index.ts # Main server code
├── dist/ # Compiled JavaScript
├── package.json
├── tsconfig.json
├── LICENSE
└── README.mdDependencies
@modelcontextprotocol/sdk- MCP SDKaxios- HTTP clientzod- Schema validationform-data- Multipart form data for STT
License
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
You should have received a copy of the GNU General Public License along with this program. If not, see https://www.gnu.org/licenses/.
This server cannot be deployed
Maintenance
Related MCP Connectors
AI voice generation: text-to-speech and voice cloning from any MCP client.
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
Speech, transcription, voice agents, Trace, Recap, dubbing and narration with browser OAuth.
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables advanced audio transcription, text-to-speech generation, and audio processing using OpenAI's Whisper and GPT-4o models with support for multiple audio formats, file management, and parallel processing.821 PyPI60MIT
- AlicenseCqualityDmaintenanceEnables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.1MIT
- AlicenseAqualityFmaintenanceEnables text-to-speech conversion using ElevenLabs API with voice management, streaming support, and multiple models.51MIT
- AlicenseNot gradedqualityDmaintenanceProvides text-to-speech functionality using OpenAI's TTS API, enabling text-to-speech conversion, voice listing, and model listing.MIT