Fish Audio MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| FISH_API_KEY | Yes | Your Fish Audio API key | |
| FISH_LATENCY | No | Latency mode (low, balanced, normal) | balanced |
| FISH_MODEL_ID | No | TTS model to use (s2-pro, s1) | s2-pro |
| FISH_AUTO_PLAY | No | Auto-play audio | false |
| FISH_STREAMING | No | Enable streaming mode | false |
| FISH_REFERENCES | No | Multiple voice references as JSON array string | |
| AUDIO_OUTPUT_DIR | No | Directory for audio file output | ~/.fish-audio-mcp/audio_output |
| FISH_MP3_BITRATE | No | MP3 bitrate (64, 128, 192) | 128 |
| FISH_REFERENCE_ID | No | Default voice reference ID (single reference mode) | |
| FISH_OUTPUT_FORMAT | No | Default audio format (mp3, wav, pcm, opus) | mp3 |
| FISH_REFERENCE_1_ID | No | First voice reference ID (individual format) | |
| FISH_REFERENCE_2_ID | No | Second voice reference ID (individual format) | |
| FISH_REFERENCE_1_NAME | No | First voice reference name (individual format) | |
| FISH_REFERENCE_1_TAGS | No | First voice reference tags (comma-separated) (individual format) | |
| FISH_REFERENCE_2_NAME | No | Second voice reference name (individual format) | |
| FISH_REFERENCE_2_TAGS | No | Second voice reference tags (comma-separated) (individual format) | |
| FISH_DEFAULT_REFERENCE | No | Default reference ID when using multiple references |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| fish_audio_ttsB | Generate speech from text using Fish Audio TTS API |
| fish_audio_list_referencesA | List all configured voice references |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: one manages references, the other generates speech. No overlap or ambiguity.
Both tools follow the consistent 'fish_audio_verb_noun' pattern with snake_case, making them predictable and easy to understand.
With only 2 tools, the server feels thin for a TTS service, but it may be appropriate for a minimal integration. The count is borderline but not extreme.
The server lacks essential operations like creating/deleting references, listing voices, or setting voice parameters. Users cannot fully manage the TTS workflow, leading to significant gaps.