gemini-tts
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| GEMINI_API_KEY | Yes | Google Gemini API key, stored in the .env file as GEMINI_API_KEY=your_key. Required to authenticate with the Gemini TTS API. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| generate_speechA | Generate natural English narration audio using Gemini TTS. The default voice is Algenib with a rapid-fire, slightly smiling, energetic delivery and a neutral English accent. The resulting audio is saved as a WAV file. |
| get_audio_filesA | List all WAV audio files generated by the Gemini TTS MCP. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools are entirely distinct: one generates speech, the other lists generated files. There is no overlap or ambiguity in their purposes.
Both tools follow the verb_noun pattern: generate_speech and get_audio_files. The naming is consistent, clear, and predictable.
Two tools is on the thin side for a TTS server, but they cover the core generate and retrieve workflow. It feels slightly minimal rather than unreasonably sparse.
The server covers the essential TTS lifecycle: generating speech and listing outputs. Minor gaps exist such as no delete operation or voice selection flexibility, but agents can work around them for basic use.