MCP Make Sound
The MCP Make Sound server provides sound playback capabilities for macOS, integrating with MCP clients to offer multiple audio functions:
Pre-configured Sounds: Play informational (Glass.aiff), warning (Purr.aiff), and error (Sosumi.aiff) notification sounds.
System Sounds: Access any of the 14 built-in macOS system sounds by name.
Text-to-Speech (TTS): Convert text to speech with customizable voices.
Custom Audio Files: Play audio files from a specified absolute path on disk.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Make Soundplay a success sound for the completed task"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
π MCP Make Sound
A Model Context Protocol (MCP) server that provides comprehensive sound playback capabilities for macOS. This server allows AI assistants and other MCP clients to play system sounds, text-to-speech, and custom audio files for rich audio feedback.
β¨ Features
π Simple Sound Methods: Pre-configured info, warning, and error sounds
π΅ Custom System Sounds: Play any of the 14 built-in macOS sounds
π£οΈ Text-to-Speech: Convert text to speech with customizable voices
π File Playback: Play custom audio files from disk
π Built with TypeScript and the MCP SDK
πͺΆ Lightweight and easy to integrate
Related MCP server: MCP Sound Tool
π Requirements
π macOS (uses
afplayand system sounds)π’ Node.js 18+
π TypeScript
π Installation
Clone this repository:
git clone <repository-url>
cd mcp-make-soundInstall dependencies:
npm installBuild the project:
npm run buildπ‘ Usage
π΅ Running the Server
Start the MCP server:
npm startFor development with auto-reload:
npm run devπ― Example: Claude Integration with Warp Terminal
Here's how you can set up the MCP sound server to provide audio feedback when AI tasks complete in Warp terminal:
Configuration Rule: "When AI is done, use mcp-make-sound to play a sound. The MCP supports error, info and success. Play the right sound based on AI task outcome."
This setup allows you to:
π Hear a pleasant chime when tasks complete successfully
β οΈ Get an alert sound for warnings or partial completions
β Receive clear audio feedback for errors or failures
The audio feedback helps you stay focused on other work while knowing immediately when your AI assistant has finished processing your requests.
π οΈ Available Tools
The server provides four tools:
Simple Sound Methods (Legacy)
play_info_sound
Description: Play an informational system sound
Parameters: None
Sound: Glass.aiff
play_warning_sound
Description: Play a warning system sound
Parameters: None
Sound: Purr.aiff
play_error_sound
Description: Play an error system sound
Parameters: None
Sound: Sosumi.aiff
Advanced Sound Method
play_sound
Description: Play various types of sounds with customizable parameters
Parameters:
type(required):"system","tts", or"file"Additional based on type (see examples below)
π Usage Examples
System Sounds
Play any of the 14 built-in macOS sounds:
{
"name": "play_sound",
"arguments": {
"type": "system",
"name": "Basso"
}
}Available system sounds:
Basso, Blow, Bottle, Frog, Funk, Glass, Hero, Morse, Ping, Pop, Purr, Sosumi, Submarine, Tink
Text-to-Speech
Convert text to speech with optional voice selection:
{
"name": "play_sound",
"arguments": {
"type": "tts",
"text": "Hello, this is a test message",
"voice": "Albert"
}
}Without voice (uses system default):
{
"name": "play_sound",
"arguments": {
"type": "tts",
"text": "Task completed successfully"
}
}Supported voices:
English:
Albert,Alice,Bad News,Bahh,Bells,Boing,Bruce,Bubbles,Cellos,Daniel,Deranged,Fred,Good News,Hysterical,Junior,Kathy,Pipe Organ,Princess,Ralph,Trinoids,Whisper,ZarvoxInternational:
Anna,AmΓ©lie,Daria,Eddy,Fiona,Jorge,Juan,Luca,Marie,Moira,Nora,Rishi,Samantha,Serena,Tessa,Thomas,Veena,Victoria,Xander,Yelda,Zosia
Note: If an unsupported voice is specified, the system will gracefully fall back to the default voice and continue playback.
Custom Audio Files
Play audio files from disk:
{
"name": "play_sound",
"arguments": {
"type": "file",
"path": "/Users/username/Music/notification.mp3"
}
}Supports common audio formats: .aiff, .wav, .mp3, .m4a, etc.
π Security & Limitations
System Sounds: Only the 14 official macOS sounds are allowed
Text-to-Speech:
Text limited to 1000 characters maximum
Voice validation with graceful fallback to system default
Curated list of 43+ supported voices for security
File Playback: Requires absolute paths and validates file existence
π Integration with MCP Clients
This server can be integrated with any MCP-compatible client, such as:
π€ Claude Desktop
π οΈ Custom MCP clients
π§ AI assistants that support MCP
MCP Configuration Example
Add this to your MCP client configuration:
{
"mcp-make-sound": {
"command": "node",
"args": [
"/Users/nocoo/Workspace/mcp-make-sound/dist/index.js"
],
"env": {},
"working_directory": "/Users/nocoo/Workspace/mcp-make-sound",
"start_on_launch": true
}
}Example tool calls:
{
"name": "play_info_sound",
"arguments": {}
}{
"name": "play_sound",
"arguments": {
"type": "system",
"name": "Hero"
}
}π οΈ Development
π Project Structure
mcp-make-sound/
βββ src/
β βββ index.ts # Main server implementation
β βββ __tests__/ # Unit tests
β βββ sound.test.ts # Sound system tests
βββ dist/ # Compiled JavaScript output
βββ eslint.config.js # ESLint configuration
βββ vitest.config.ts # Vitest test configuration
βββ package.json # Project configuration
βββ tsconfig.json # TypeScript configuration
βββ README.md # This fileπ Scripts
Build & Run
npm run build- π¨ Compile TypeScript to JavaScriptnpm start- βΆοΈ Run the compiled servernpm run dev- π Development mode with auto-rebuild and restartnpm run kill- π Stop all running MCP server instances
Code Quality & Testing
npm run lint- π Check code style and errorsnpm run lint:fix- π§ Fix auto-fixable linting issuesnpm run test- π§ͺ Run tests in watch modenpm run test:run- β Run tests oncenpm run test:ui- ποΈ Run tests with interactive UI
βοΈ How It Works
The server implements the MCP protocol using the official SDK
It exposes four tools for different sound capabilities
When a tool is called, it uses macOS commands:
afplayfor audio file playback (system sounds and custom files)sayfor text-to-speech synthesis
System sounds are located in
/System/Library/Sounds/The server communicates over stdio transport
π§ Technical Details
π Transport: Standard I/O (stdio)
π‘ Protocol: Model Context Protocol (MCP)
π§ Audio Backend: macOS
afplayandsaycommandsπ΅ Sound Files: System .aiff files, custom audio files, and synthesized speech
π¨ Error Handling
The server includes comprehensive error handling:
Validates tool names and parameters
Handles
afplayandsaycommand failuresValidates required parameters for each sound type
Returns appropriate error messages to clients
Graceful server shutdown on errors
π License
MIT License
π€ Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
πΌ Sound Capabilities
Simple Methods (Legacy)
π Info: Glass.aiff - A pleasant chime sound
β οΈ Warning: Purr.aiff - A gentle alert sound
β Error: Sosumi.aiff - A distinctive error sound
Advanced Method (play_sound)
π΅ 14 System Sounds: All built-in macOS sounds available
π£οΈ 50+ TTS Voices: Multiple languages and character voices
π Custom Files: Support for .aiff, .wav, .mp3, .m4a, and more
These capabilities provide rich audio feedback options for any application need.
Available Tools
4 toolsplay_error_soundB
Play an error system sound
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action ('Play') but doesn't describe what 'Play' entails (e.g., audio output, duration, system requirements) or any side effects (e.g., interruptions, permissions needed). This leaves significant gaps in understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It's front-loaded and appropriately sized for a simple tool, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is minimally adequate. It states what the tool does but lacks details on behavioral context (e.g., how the sound is played, any system dependencies). For such a straightforward tool, this is acceptable but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't mention parameters, which is correct for a parameterless tool, earning a high baseline score for not adding unnecessary information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Play') and resource ('error system sound'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'play_warning_sound' or 'play_info_sound' beyond the 'error' qualifier, which is why it doesn't reach a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'play_warning_sound' or 'play_sound'. It lacks explicit context, exclusions, or comparisons to sibling tools, leaving the agent to infer usage based on the 'error' keyword alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_info_soundB
Play an informational system sound
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the action ('play') but doesn't describe what happens when invoked (e.g., audible output, system behavior, permissions needed, or side effects). This leaves significant gaps for understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It's appropriately sized for a simple, parameterless tool and is front-loaded with essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no output schema, the description adequately covers the basic purpose. However, it lacks behavioral context (e.g., what 'play' entails, system requirements, or error conditions) that would be helpful given the absence of annotations and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, maintaining focus on the tool's purpose without unnecessary detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Play an informational system sound' clearly states the action (play) and resource (informational system sound). It distinguishes from sibling tools like 'play_error_sound' and 'play_warning_sound' by specifying the type of sound, though it doesn't explicitly differentiate from the generic 'play_sound' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'play_error_sound', 'play_warning_sound', or 'play_sound'. It doesn't specify appropriate contexts, exclusions, or prerequisites for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_soundC
Play various types of sounds with customizable parameters
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | Type of sound to play | |
| name | No | System sound name (required when type is "system") | |
| text | No | Text to speak (required when type is "tts") | |
| voice | No | Voice name (optional, used with type "tts", uses system default if not specified) | |
| path | No | Absolute path to audio file (required when type is "file") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'customizable parameters' but fails to explain key traits like whether playback is blocking or non-blocking, error handling (e.g., if a file path is invalid), or system dependencies (e.g., TTS availability). This leaves significant gaps in understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose ('play various types of sounds'). It avoids redundancy but could be more structured by explicitly mentioning the parameter types (e.g., system, TTS, file) to enhance clarity without adding unnecessary length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of handling multiple sound types (system, TTS, file) with no annotations or output schema, the description is incomplete. It lacks details on behavioral aspects like playback effects, error responses, or how outputs are handled, making it insufficient for safe and effective use by an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, providing clear details for all parameters, including enums and requirements. The description adds minimal value beyond this, as it only vaguely references 'customizable parameters' without elaborating on their semantics or interactions, aligning with the baseline score for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'play[s] various types of sounds with customizable parameters', which clarifies the action (play) and resource (sounds) but is vague about scope and differentiation. It does not specify what 'various types' means or how it differs from sibling tools like play_error_sound, leaving room for ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as the sibling tools (play_error_sound, play_info_sound, play_warning_sound). The description implies general sound playback but offers no context for choosing between this and more specific tools, leading to potential misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_warning_soundB
Play a warning system sound
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action ('Play') but doesn't describe what happensβe.g., whether it's audible, requires permissions, has side effects, or how it interacts with the system. For a tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It's front-loaded with the core action and resource, making it easy to parse quickly. Every word earns its place, achieving maximum clarity in minimal space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema), the description is minimal but adequate for basic understanding. However, with no annotations and no output schema, it lacks details on behavior, effects, or return values, which could be important for an agent to use it correctly in context. It's incomplete for a tool that might have auditory or system-level implications.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so there are no parameters to document. The description doesn't need to add parameter semantics, and it appropriately doesn't mention any. Baseline is 4 for zero parameters, as there's nothing to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Play') and the resource ('a warning system sound'), making the purpose immediately understandable. However, it doesn't differentiate this tool from its siblings (play_error_sound, play_info_sound, play_sound) beyond specifying 'warning' versus other sound types, which is somewhat implicit but not explicit about distinctions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings or alternatives. It implies usage for warning scenarios but doesn't specify contexts, exclusions, or comparisons with other sound-playing tools, leaving the agent to infer based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The tool set has significant ambiguity, as 'play_sound' appears to be a general-purpose tool that can handle what the three specific tools (error, info, warning) do, making it unclear when to choose one over the other. This overlap could easily lead to misselection by agents, as the specific tools seem redundant given the customizable 'play_sound'.
All tool names follow a consistent verb_noun pattern with 'play_' as the prefix, using snake_case throughout. This predictability makes it easy for agents to parse and understand the naming conventions without confusion.
With 4 tools, the count is borderline for the server's purpose of playing sounds. It feels slightly thin, as the domain might benefit from more specific or varied sound operations, but it's not extreme. The tools cover basic sound types, though the overlap reduces the effective utility of having four separate tools.
The tool set covers basic sound playback for error, info, warning, and customizable sounds, which aligns with the server's name 'Make Sound'. However, there are notable gaps, such as missing operations like listing available sounds, stopping sounds, or adjusting volume, which could limit agent workflows in more complex scenarios.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yoβ¦
A Model Context Protocol server for Wix AI tools
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to control Unreal Eβ¦
The Telnyx MCP server is an official implementation of the Model Context Protocol that enables AI clients (like Claude Desktop, Cursor, and OpenAI Agents) to interact with Telnyx's telephony, messaging, and AI assistant APIs. It provides comprehensive capabilities including making and managing phone calls, sending SMS/MMS messages, purchasing and configuring phone numbers, creating AI assistants with custom instructions, managing cloud storage buckets, scraping and embedding website content, and handling integration secrets. The server exists as both a local implementation and a remotely hosted version, allowing developers to integrate real-world communication infrastructure directly into AI applications.
Related MCP Servers
- AlicenseBqualityFmaintenanceA Model Context Protocol service that sends desktop notifications and alert sounds when AI agent tasks are completed, integrating with various LLM clients like Claude Desktop and Cursor.154MIT
- AlicenseAqualityFmaintenanceA Model Context Protocol implementation that plays sound effects (completion, error, notification) for Cursor AI and other MCP-compatible environments, providing audio feedback for a more interactive coding experience.31MIT
- AlicenseBqualityCmaintenanceA Model Context Protocol server that allows AI agents to play notification sounds when tasks are completed.12714Apache 2.0
- AlicenseBqualityCmaintenanceA Model Context Protocol server that allows AI assistants to communicate with the ChatGPT desktop app on macOS, enabling users to send prompts to ChatGPT from any MCP-compatible assistant.378MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nocoo/mcp-make-sound'
If you have feedback or need assistance with the MCP directory API, please join our Discord server