MCP Sound Tool
MCP 声音工具
模型上下文协议 (MCP) 实现,可为 Cursor AI 和其他兼容 MCP 的环境播放音效。此 Python 实现提供音频反馈,以实现更具互动性的编码体验。
特征
播放各种事件(完成、错误、通知)的音效
使用模型上下文协议 (MCP) 与 Cursor 和其他 IDE 进行标准化集成
跨平台支持(Windows、macOS、Linux)
可配置的音效
Related MCP server: MCP Notify Server
安装
Python 版本兼容性
此软件包已使用 Python 3.8-3.11 测试。如果您在使用 Python 3.12+ 时遇到错误(尤其是BrokenResourceError或TaskGroup异常),请尝试使用早期版本的 Python。
推荐:使用 pipx 安装
安装 mcp-sound-tool 的推荐方法是使用pipx ,它会在隔离的环境中安装包,同时使命令在全球范围内可用:
# Install pipx if you don't have it
python -m pip install --user pipx
python -m pipx ensurepath
# Install mcp-sound-tool
pipx install mcp-sound-tool这种方法确保工具拥有自己独立的环境,避免与其他包发生冲突。
替代方案:使用 pip 安装
也可以直接用pip安装:
pip install mcp-sound-tool来自源
克隆此存储库:
git clone https://github.com/yourusername/mcp-sound-tool cd mcp-sound-tool直接从源目录使用 pipx 安装:
pipx install .或者使用 pip:
pip install -e .
用法
添加声音文件
将声音文件放入sounds目录中。预期的声音文件如下:
completion.mp3- 代码生成后播放error.mp3- 发生错误时播放notification.mp3- 用于一般通知
您可以在 freesound.org 等网站上找到免费的音效。
运行 MCP 服务器
运行 MCP 服务器:
mcp-sound-tool服务器将启动并通过 stdio 传输监听来自 Cursor 或其他 MCP 兼容客户端的事件。
光标中的配置
要将此服务器与 Cursor 一起使用,请将其添加到您的 MCP 配置文件中:
在 macOS 上:
// ~/Library/Application Support/Cursor/mcp.json
{
"mcpServers": {
"sound": {
"command": "mcp-sound-tool",
"args": [],
"type": "stdio",
"pollingInterval": 5000,
"startupTimeout": 10000,
"restartOnFailure": true
}
}
}在 Windows 上:
// %APPDATA%/Cursor/mcp.json
{
"mcpServers": {
"sound": {
"command": "mcp-sound-tool",
"args": [],
"type": "stdio",
"pollingInterval": 5000,
"startupTimeout": 10000,
"restartOnFailure": true
}
}
}当与pipx一起安装时, mcp-sound-tool命令将在您的 PATH 中可用,因此 Cursor 将能够找到并执行它而无需指定完整路径。
适用于 AI 模型的合理 MCP 使用指南
此 MCP 服务器为 AI 交互提供音频反馈功能。它旨在通过提供清晰的音频提示来指示操作状态,而无需用户阅读文本,从而提升用户体验。
何时使用声音反馈
AI代理应该在适当的时候主动使用声音工具:
成功声音(
completion) :任务或命令成功完成后
当重要操作成功完成时
当确认用户的请求已得到满足时
错误声音(
error) :当命令失败或遇到错误时
当警告用户出现问题时
当操作无法按要求完成时
通知声音(
notification) :提醒用户重要信息时
当提示用户注意或输入时
用于长期运行操作的状态更新
示例用法
# When a command completes successfully
@mcp.tool()
def execute_command(command):
result = run_command(command)
if result.success:
play_sound("completion") # Indicate success with audio
return "Command executed successfully"
else:
play_sound("error") # Indicate failure with audio
return f"Error: {result.error_message}"可用工具
play_sound(sound_type="completion", custom_sound_path=None):播放音效list_available_sounds():列出所有可用的声音文件install_to_user_dir():将声音文件安装到用户的配置目录
有关更多详细信息,请连接到 MCP 服务器并查看工具说明。
发展
对于开发:
# Install development dependencies
pip install -e ".[dev]"
# Run tests
pytest致谢
SIAM-TheLegend创建了原始的sound-mcp JavaScript 实现,并启发了此 Python 版本
MCP 协议开发人员为 AI 工具交互创建了强大的标准
测试和文档的贡献者
执照
该项目根据 MIT 许可证获得许可 - 有关详细信息,请参阅 LICENSE 文件。
Available Tools
3 toolsinstall_to_user_dirA
Install sound files to user's config directory.
WHEN TO USE THIS TOOL:
- When the user wants to customize the sound files
- When setting up the sound tool for the first time
- When troubleshooting missing sound files
This tool copies the default sound files to the user's configuration directory
where they can be modified or replaced with custom sounds.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must cover behavioral traits. It states the tool copies default sound files to the config directory for modification, but lacks details on whether files are overwritten, directory creation, or error conditions. This provides basic but incomplete behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, uses bullet points for clarity, and front-loads the core action in the first line. Every sentence adds value with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description adequately covers purpose and usage scenarios. It does not elaborate on return values (not required due to output schema) or side effects like overwriting, but the simplicity of the tool makes this likely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, making schema coverage trivially 100%. Per the rubric, a baseline of 4 applies. No parameter information is needed, and the description does not need to add any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the primary action: 'Install sound files to user's config directory.' This is a specific verb-resource combination that distinguishes it from sibling tools list_available_sounds and play_sound, which are read and playback operations respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes explicit 'WHEN TO USE THIS TOOL' bullets, covering customization, first-time setup, and troubleshooting. While it does not specify when not to use or explicitly name alternatives, the use cases are clear and distinct from siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_available_soundsA
List all available notification sounds.
WHEN TO USE THIS TOOL:
- When you need to check what sound options are available
- When determining if a specific sound file exists
- Before using a custom sound to verify available options
This tool helps you discover what sounds are available for providing audio feedback.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies read-only fetching of available sounds, but does not explicitly state that it is safe, nondestructive, or what the output format is. Since it's a simple list tool, this is adequate but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with only two sentences plus a bulleted usage section. It front-loads the purpose and uses clear formatting, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool complexity (no parameters), the presence of an output schema, and the absence of annotations, the description fully covers the tool's purpose and usage. No additional details are needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, and schema coverage is 100%. The description adds value by explaining the purpose, aligning with the baseline score of 4 for parameterless tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists all available notification sounds. It uses specific verb 'list' and resource 'notification sounds', distinguishing it from sibling tools like 'play_sound' and 'install_to_user_dir'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides three scenarios for when to use the tool: checking available options, verifying existence of a specific sound, and before using a custom sound. This gives clear guidance without needing to mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_soundA
Play a notification sound on the user's device.
WHEN TO USE THIS TOOL:
- Use 'completion' sound when a task or command has SUCCESSFULLY completed
- Use 'error' sound when a command has FAILED or an error has occurred
- Use 'notification' sound for important alerts or information that needs attention
- Use 'custom' sound only when you need a specific sound not covered by the standard types
AI agents SHOULD proactively use these sounds to provide audio feedback based on
the outcome of commands or operations, enhancing the user experience with
non-visual status indicators.
Example usage: After executing a terminal command, play a 'completion' sound if
successful or an 'error' sound if it failed.
| Name | Required | Description | Default |
|---|---|---|---|
| sound_type | No | completion | |
| custom_sound_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether audio playback is synchronous, permission requirements, or error handling. The output schema exists but is not described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear introduction and bullet-point guidelines. The example adds context. It is slightly verbose but efficiently conveys necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no required ones), the description covers the purpose, usage scenarios, and parameter semantics adequately. It does not discuss return values, but that is acceptable for a straightforward action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description adds meaning to the sound_type parameter by explaining when to use each value. The custom_sound_path parameter is implied but not detailed. This compensates partially for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Play a notification sound on the user's device' and differentiates between sound types with specific use cases. It clearly distinguishes from sibling tools like install_to_user_dir and list_available_sounds.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides detailed WHEN TO USE guidelines for each sound type (completion, error, notification, custom) and an example. This gives clear context for when the tool should be invoked.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
install_to_user_dir - First observed
list_available_sounds - First observed
play_sound
TDQS
Scored across 3 tools
Each tool has a distinct purpose: installing, listing, and playing sounds, with no overlap in functionality.
All tools use a consistent verb_noun pattern in snake_case, with clear and descriptive names.
Three tools is a reasonable number for a sound tool, covering setup, exploration, and core usage.
The tool set covers basic lifecycle: install, list, play. Minor gap: no tool for direct volume control or custom sound management beyond defaults.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
CC0 sound effects API for AI agents — search, preview, and download via MCP.
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
Shared memory and actions for Claude, Kiro, OpenAI, Cursor, and other MCP-compatible AI clients.
Related MCP Servers
- FlicenseDqualityDmaintenanceProvides audio feedback by playing sound effects when Cursor AI completes code generation, creating a more interactive coding experience.119-
- AlicenseBqualityFmaintenanceA Model Context Protocol service that sends desktop notifications and alert sounds when AI agent tasks are completed, integrating with various LLM clients like Claude Desktop and Cursor.154MIT
- FlicenseBqualityDmaintenancePlays sound effects when Cursor AI completes code generation, providing audio feedback for a more interactive coding experience.12-
- AlicenseBqualityCmaintenanceA Model Context Protocol server that allows AI agents to play notification sounds when tasks are completed.146 npm14Apache 2.0