@theyahia/yandex-speechkit-mcp
Related Servers
Alternatives to @theyahia/yandex-speechkit-mcp
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityDmaintenanceProvides speech recognition and synthesis tools via SaluteSpeech API, enabling AI assistants to handle voice input and output.4MIT
- AlicenseAqualityBmaintenanceProvides speech recognition (STT) and synthesis (TTS) tools via the Sber SaluteSpeech API, enabling audio transcription and voice generation through natural language.554 npm1MIT

VoiceLab MCP Serverofficial
AlicenseNot gradedqualityCmaintenanceEnables AI agents to use speech AI capabilities including LLM chat completions, text-to-speech in Uzbek, Russian, and English, speech-to-text transcription with speaker labels, voice isolation for noise removal, and realtime WebSocket streaming.MIT- AlicenseAqualityBmaintenanceEnables AI agents to generate speech using Gemini TTS models, with tools for text-to-speech, task polling, and pricing checks.4246 npmApache 2.0
- FlicenseNot gradedqualityDmaintenanceProvides tools for generating speech from text using the ElevenLabs API, including voice listing, text-to-speech conversion, and quota checking.-
- AlicenseAqualityDmaintenanceEnables AI assistants to manage Yandex Direct advertising campaigns, ads, keywords, and reports via natural language using the Yandex Direct API v5.21MIT
TDQS
Scored across 5 tools
Tools are mostly distinct: list_voices (voice listing), recognize (low-level STT), synthesize (low-level TTS), skill_synthesize (high-level TTS), skill_transcribe (high-level STT). The high-level vs low-level distinction is clear in descriptions, but an agent might hesitate between skill_synthesize and synthesize.
Naming is inconsistent: list_voices follows verb_noun pattern, recognize and synthesize are single verbs, skill_synthesize and skill_transcribe have a 'skill_' prefix. Mix of patterns could confuse agents.
5 tools is appropriate for a speech kit server. It covers both STT and TTS with low-level and high-level options, without being overwhelming.
Core STT and TTS functionality is covered. Minor gaps like explicit language detection or streaming aren't present but are not critical given the high-level tools handle auto-detection.