Whissle MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| speech_to_textA | Convert speech to text with a given model and save the output text file to a given directory. Directory is optional, if not provided, the output file will be saved to $HOME/Desktop. |
| diarize_speechA | Convert speech to text with speaker diarization and save the output text file to a given directory. Directory is optional, if not provided, the output file will be saved to $HOME/Desktop. |
| translate_textA | Translate text from one language to another. |
| summarize_textA | Summarize text using an LLM model. |
| list_asr_modelsB | List all available ASR models and their capabilities. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Most tools have distinct purposes: diarize_speech adds speaker identification to transcription, speech_to_text is basic transcription, list_asr_models provides metadata, summarize_text and translate_text handle text processing. However, diarize_speech and speech_to_text share significant overlap in core transcription functionality, which could cause confusion about when to use each.
All tools follow a consistent verb_noun naming pattern with snake_case throughout: diarize_speech, list_asr_models, speech_to_text, summarize_text, and translate_text. The naming is predictable and follows the same grammatical structure across all five tools.
Five tools is reasonable for a speech/text processing server, covering transcription (with and without diarization), model listing, summarization, and translation. The count feels slightly thin for a comprehensive MCP server but adequately covers core functionality without being overwhelming.
The server covers basic speech-to-text workflows and text processing operations, but has notable gaps. There's no text-to-speech capability, no audio file manipulation tools, and no batch processing operations. While the existing tools handle individual tasks, the surface feels incomplete for a comprehensive speech/text processing domain.