Skip to main content
Glama
theYahia

@theyahia/yandex-speechkit-mcp

by theYahia

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
FOLDER_IDYesYandex Cloud folder ID (required)
IAM_TOKENNoShort-lived IAM token (alternative to API key)
YANDEX_SPEECHKIT_API_KEYNoYandex Cloud API key (preferred)

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
recognizeC

Speech recognition (STT) via Yandex SpeechKit. Takes Base64 audio, returns text.

synthesizeC

Speech synthesis (TTS) via Yandex SpeechKit. Takes text, returns Base64 audio.

list_voicesB

List available TTS voices. Optionally filter by language.

skill_transcribeC

High-level transcription skill. Returns clean text from audio.

skill_synthesizeB

High-level speech synthesis skill. Smart voice defaults, auto-detects language.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.2/5.0

Scored across 5 tools

Disambiguation4/5

Tools are mostly distinct: list_voices (voice listing), recognize (low-level STT), synthesize (low-level TTS), skill_synthesize (high-level TTS), skill_transcribe (high-level STT). The high-level vs low-level distinction is clear in descriptions, but an agent might hesitate between skill_synthesize and synthesize.

Naming Consistency3/5

Naming is inconsistent: list_voices follows verb_noun pattern, recognize and synthesize are single verbs, skill_synthesize and skill_transcribe have a 'skill_' prefix. Mix of patterns could confuse agents.

Tool Count5/5

5 tools is appropriate for a speech kit server. It covers both STT and TTS with low-level and high-level options, without being overwhelming.

Completeness4/5

Core STT and TTS functionality is covered. Minor gaps like explicit language detection or streaming aren't present but are not critical given the high-level tools handle auto-detection.

Maintenance

ActivitySlowing
ResponsivenessNo issues