@theyahia/yandex-speechkit-mcp
> ## 🗄 Репозиторий заархивирован
>
> Разработка переехала в **[theYahia/YaAll](https://github.com/theYahia/YaAll)** — сборку, где весь яндексовский слой лежит в одном месте: свои MCP-серверы, скиллы Claude Code и материалы официальных наборов Яндекса.
>
> Актуальная версия того, что лежало здесь: [`mcp/yandex-speechkit-mcp/`](https://github.com/theYahia/YaAll/tree/main/mcp/yandex-speechkit-mcp)
>
> Пакет в npm прежний — [`@theyahia/yandex-speechkit-mcp`](https://www.npmjs.com/package/@theyahia/yandex-speechkit-mcp), ставится и работает как раньше.
> Здесь больше ничего не обновляется. Задачи и pull request'ы — в YaAll.
>
> **Archived — development moved to [theYahia/YaAll](https://github.com/theYahia/YaAll),** a single repository bundling the whole Yandex stack.
> The current version of this package now lives at [`mcp/yandex-speechkit-mcp/`](https://github.com/theYahia/YaAll/tree/main/mcp/yandex-speechkit-mcp).
> The npm package [`@theyahia/yandex-speechkit-mcp`](https://www.npmjs.com/package/@theyahia/yandex-speechkit-mcp) is unchanged.
> Please open issues and pull requests there.
> Этот сервер входит в сборку **[theYahia/YaAll](https://github.com/theYahia/YaAll)** —
> весь яндексовский слой в одном репозитории: десять MCP-серверов, скиллы Claude Code
> под SEO и валидацию спроса, плюс материалы официальных серверов Яндекса.
> Здесь он живёт отдельно, там — рядом с остальными: [`mcp/yandex-speechkit-mcp/`](https://github.com/theYahia/YaAll/tree/main/mcp/yandex-speechkit-mcp)
>
> *Part of [theYahia/YaAll](https://github.com/theYahia/YaAll) — the whole Yandex stack in one repo.*
# @theyahia/yandex-speechkit-mcp
MCP server for Yandex SpeechKit API — speech recognition, synthesis, and voice listing. **5 tools.**
[](https://www.npmjs.com/package/@theyahia/yandex-speechkit-mcp)
[](https://opensource.org/licenses/MIT)
Part of the [Russian API MCP](https://github.com/theYahia/russian-mcp) series by [@theYahia](https://github.com/theYahia).
## Installation
### Claude Desktop
```json
{
"mcpServers": {
"yandex-speechkit": {
"command": "npx",
"args": ["-y", "@theyahia/yandex-speechkit-mcp"],
"env": {
"YANDEX_SPEECHKIT_API_KEY": "your-api-key",
"FOLDER_ID": "your-folder-id"
}
}
}
}
```
### Claude Code
```bash
claude mcp add yandex-speechkit \
-e YANDEX_SPEECHKIT_API_KEY=your-api-key \
-e FOLDER_ID=your-folder-id \
-- npx -y @theyahia/yandex-speechkit-mcp
```
### Streamable HTTP (remote / Docker)
```bash
YANDEX_SPEECHKIT_API_KEY=... FOLDER_ID=... npx @theyahia/yandex-speechkit-mcp --http
# Listens on :8080/mcp (override with PORT env var)
```
### Smithery
Deploy via [smithery.ai](https://smithery.ai) — config in `smithery.yaml`.
## Authentication
| Variable | Description |
|----------|-------------|
| `YANDEX_SPEECHKIT_API_KEY` | Yandex Cloud API key (preferred) |
| `YANDEX_API_KEY` | Legacy alias (still works) |
| `IAM_TOKEN` | Short-lived IAM token (alternative to API key) |
| `FOLDER_ID` | Yandex Cloud folder ID (required) |
| `YANDEX_FOLDER_ID` | Legacy alias for FOLDER_ID |
> Get credentials at [Yandex Cloud Console](https://console.cloud.yandex.ru/).
## Tools (5)
| Tool | Type | Description |
|------|------|-------------|
| `recognize` | Core | Speech recognition (STT) — Base64 audio to text |
| `synthesize` | Core | Speech synthesis (TTS) — text to Base64 audio |
| `list_voices` | Core | List available TTS voices, filter by language |
| `skill_transcribe` | Skill | High-level transcription — returns clean text |
| `skill_synthesize` | Skill | High-level synthesis — smart defaults, auto-detects language from voice |
## Examples
```
Transcribe this audio file
Synthesize "Hello, how are you?" with voice filipp
What voices are available in Russian?
Speak this text using the alena voice
```
## Development
```bash
npm install
npm run build
npm test
npm run dev # stdio mode
```
## License
MIT
TDQS
Scored across 5 tools
Tools are mostly distinct: list_voices (voice listing), recognize (low-level STT), synthesize (low-level TTS), skill_synthesize (high-level TTS), skill_transcribe (high-level STT). The high-level vs low-level distinction is clear in descriptions, but an agent might hesitate between skill_synthesize and synthesize.
Naming is inconsistent: list_voices follows verb_noun pattern, recognize and synthesize are single verbs, skill_synthesize and skill_transcribe have a 'skill_' prefix. Mix of patterns could confuse agents.
5 tools is appropriate for a speech kit server. It covers both STT and TTS with low-level and high-level options, without being overwhelming.
Core STT and TTS functionality is covered. Minor gaps like explicit language detection or streaming aren't present but are not critical given the high-level tools handle auto-detection.