discord-music-mcp
Lets the server play music in Discord voice channels. It lists the servers and voice channels the companion bot can join, resolves channels by name (partial names work) or ID, joins the requested voice channel, and can stop playback and leave the channel. Capabilities include playing a track in a specific guild/voice channel and reporting what is currently playing and where.
Provides YouTube music search and playback: plays the top YouTube result for a text search or a direct URL, and can search YouTube and return a list of results (e.g. several versions of a song) without playing them, so the user can pick one.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@discord-music-mcpplay some lo-fi in the Lobby"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
discord-music-mcp
An MCP server that lets Claude (or any MCP client) play music in your Discord voice channel.
"Play some lo-fi in the Lobby" → 🎶
📚 Learning MCP? This repo is also a commented, step-by-step tutorial (Turkish): docs/TUTORIAL.tr.md.
How it relates to Discord YouTube DJ
This repo is not a Discord bot. It's a small adapter in front of Discord YouTube DJ, the self-hosted bot that does the actual work.
┌──────────────────┐ stdio (JSON-RPC) ┌──────────────────┐ HTTP 127.0.0.1:1231 ┌──────────────────────┐ ┌─────────┐
│ Claude Desktop / │ ─────────────────► │ discord-music-mcp│ ────────────────────► │ Discord YouTube DJ │ ─────► │ Discord │
│ Claude Code │ ◄───────────────── │ (this repo) │ ◄──────────────────── │ bot (other repo) │ voice │ │
└──────────────────┘ └──────────────────┘ └──────────────────────┘ └─────────┘Claude starts this server as a child process when it needs it. You don't run it yourself.
This server turns tool calls (
play_song,stop_music, …) into requests to the bot's local API.The bot holds the Discord token and joins voice channels. This repo never sees or needs your Discord token.
Because the bot only accepts connections from
127.0.0.1, this server must run on the same computer as the bot.
discord-music-mcp (this repo) | ||
Talks to Discord | ✅ | ❌ |
Needs the bot token | ✅ (entered once in its dashboard) | ❌ |
Runs | all the time (starts with Windows) | only while Claude uses it |
Required | ✅ | optional |
Related MCP server: YouTube Search & Download MCP Server
What you need
Item | Needed? | Secret? | Where you get it |
Discord YouTube DJ installed, set up, and running | ✅ | — | Follow its README. The bot must be invited to a server, and the dashboard at http://localhost:1231 should show "Online". |
Node.js 20+ | ✅ | — | nodejs.org. The Windows installer of the bot bundles its own Node, but this repo needs a normal install. |
✅ | — | git-scm.com (or download this repo as a ZIP) | |
An MCP client | ✅ | — | Claude Desktop, Claude Code, or any other MCP client |
Bot address | only if not default | no |
|
Not needed: no Discord token, no client secret, no API keys, no server or channel IDs. You refer to channels by name ("Lobby"). Channel IDs also work: enable Discord Settings → Advanced → Developer Mode, then right-click a channel → Copy Channel ID.
Setup
1. Get the code
git clone https://github.com/cengizhanpece/discord-music-mcp
cd discord-music-mcp
npm installNote the full path of src/index.js. You'll need it below, e.g. C:\Users\you\discord-music-mcp\src\index.js or /home/you/discord-music-mcp/src/index.js.
2. Check that it can reach the bot
npm testYou should see the list of tools followed by your voice channels. If you see "Could not reach the Discord YouTube DJ bot", start the bot first.
3. Connect it to your MCP client
Claude Code:
claude mcp add discord-music -- node "C:\full\path\to\discord-music-mcp\src\index.js"Then run /mcp inside Claude Code to confirm it's connected.
Claude Desktop: open Settings → Developer → Edit Config (this opens claude_desktop_config.json) and add:
{
"mcpServers": {
"discord-music": {
"command": "node",
"args": ["C:\\full\\path\\to\\discord-music-mcp\\src\\index.js"]
}
}
}On Windows, write the path with double backslashes (\\) as shown. Fully quit Claude Desktop (including the tray icon) and reopen it.
Non-default bot port: if you changed the bot's port, add the address to the config. MCP clients don't pass your terminal's environment variables to the server, so it has to go in the config:
claude mcp add discord-music -e MUSIC_BOT_URL=http://localhost:4000 -- node "C:\full\path\to\discord-music-mcp\src\index.js""discord-music": {
"command": "node",
"args": ["C:\\full\\path\\to\\discord-music-mcp\\src\\index.js"],
"env": { "MUSIC_BOT_URL": "http://localhost:4000" }
}4. Try it
"Which voice channels can you play in?" "Play Daft Punk – Get Lucky in the Lobby." "Search for three versions of Kuzu Kuzu and let me pick." "What's playing?" / "Stop the music."
Updating
cd discord-music-mcp
git pull
npm installRestart Claude Desktop, or run /mcp → reconnect in Claude Code.
What it can do
Tool | Description |
| Play a URL or the top YouTube result for a search. Optional |
| Search YouTube and list results without playing. |
| List servers and voice channels the bot can join. |
| What's playing, and where. |
| Stop and leave the voice channel. |
Also: resource discord://voice-channels and prompt dj ("pick something that fits a mood and play it").
Troubleshooting
Symptom | Fix |
Could not reach the Discord YouTube DJ bot | Start the bot (Start menu → Discord YouTube DJ). Check that http://localhost:1231 opens. |
The bot is not connected to Discord yet | Finish the bot's setup in its dashboard (token + invite). |
🔒 channel / "The bot can't play in …" | The bot has no access to that private channel. See Private voice channels. |
No voice channel selected | Name a channel in your request, or pick one in the bot's dashboard. |
Server doesn't show up in Claude | Check the path to |
Wrong port / 404 errors | Set |
Development
node scripts/test-client.js search "daft punk" # a ~40-line MCP client, see how a host talks to the server
node scripts/test-client.js play "daft punk get lucky"
npm run inspect # MCP Inspector UILogs go to stderr. stdout is reserved for the JSON-RPC stream.
License
MIT
Available Tools
5 toolslist_voice_channelsSes kanallarını listeleARead-only
Botun erişebildiği tüm Discord sunucularını ve ses kanallarını listeler. Kullanıcı bir kanaldan bahsettiğinde ve hangisi olduğu belirsizse kullan.
| Name | Required | Description | Default |
|---|---|---|---|
| guild | No | Sadece bu sunucunun kanallarını göster (isim ya da ID) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a safe read operation. The description adds useful scope by saying it covers only servers and channels the bot can access, but it does not discuss return format, pagination, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, with the action stated first and the usage condition second. No redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with one fully documented parameter and an annotation covering safety, the description gives enough context to invoke it correctly. Return-value details are not needed because no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single guild parameter has 100% schema description coverage, including name-or-ID semantics, so the schema already carries the parameter meaning. The description does not mention the parameter or add any filtering details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specifies a clear verb and resource: listing all Discord servers and voice channels the bot can access. This is immediately distinguishable from the music-focused sibling tools such as search_songs and play_song.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: when the user mentions a channel and it is unclear which one. It does not name alternatives or when-not-to-use conditions, but no sibling tool performs this channel-listing role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
now_playingŞu an ne çalıyor?BRead-only
Botun bağlı olduğu kanalı ve çalan şarkıyı gösterir.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes this as a safe, non-mutating operation, so the description doesn't need to carry that burden. It adds value by disclosing the two pieces of state returned (connected channel and current track), but says nothing about what happens when nothing is playing or no channel is joined.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the two resources it reveals. It is appropriately sized, with no filler, though it offers no additional structure beyond that one clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only query with no output schema, the description adequately conveys what the call returns. The only real gap is the empty-state behavior when the bot is idle or not in a channel.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema imposes no semantics to explain and the description correctly offers none. Baseline 4 applies for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('gösterir' / shows) paired with two concrete resources: the voice channel the bot is connected to and the currently playing song. This is unambiguous and clearly distinct from action-oriented siblings like play_song or stop_music, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives such as search_songs, nor any stated prerequisites (e.g., that a bot must be active in a channel). The agent must infer the use case purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
play_songŞarkı çalA
Discord ses kanalında bir şarkı çalmaya başlar (o an çalanı keser). Ya doğrudan bir URL (YouTube, SoundCloud vb.) ya da arama metni ver; arama metni verilirse ilk YouTube sonucu çalınır. channel verilmezse bot en son kullanılan kanalda çalar.
| Name | Required | Description | Default |
|---|---|---|---|
| song | Yes | Şarkının URL'si veya arama metni | |
| guild | No | Aynı isimde kanal birden fazla sunucuda varsa sunucu adı ya da ID | |
| channel | No | Hedef ses kanalı: isim (kısmi olabilir) ya da kanal ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide a safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true). The description adds valuable behavioral context beyond annotations: it interrupts the current song and falls back to the last used channel if channel is omitted. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and its main side effect, then input modes, then channel default. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and annotations covering safety, the description supplies key runtime behavior: interruption, input modes, and channel default. It omits discussion of the guild parameter, but the schema fully documents it, so the definition is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds behavioral meaning for the song parameter (search text causes the first YouTube result to play) and for the channel parameter (defaults to last used channel if omitted). It does not further explain guild, but the schema already documents that parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: starts playing a song in a Discord voice channel, and notes that it interrupts the currently playing song. The purpose is clear, but it does not explicitly differentiate itself from sibling tools like search_songs or now_playing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful input-mode guidance (URL vs search text, first YouTube result fallback, default channel when omitted), but gives no explicit when-to-use-this-vs-alternatives guidance or exclusions relative to siblings such as search_songs or stop_music.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_songsYouTube'da şarkı araARead-only
YouTube'da arama yapar ve sonuçları (başlık, kanal, süre, URL) döndürür. Bir şey ÇALMAZ. Kullanıcı seçenek görmek istediğinde ya da doğru versiyondan emin olmak istediğinde kullan; sonra play_song'a seçilen URL'yi ver.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Kaç sonuç dönsün | |
| query | Yes | Arama metni, ör. "tarkan kuzu kuzu" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, covering safety and scope. The description adds valuable behavioral context beyond that: it lists the exact return fields and emphasizes that no playback occurs, which prevents a common misuse. It does not cover limits, error handling, or pagination, so a 4 rather than a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and return fields, then the critical 'does not play' clarification, then usage and follow-up. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description compensates by naming the returned fields. Annotations cover read-only and open-world behavior, and the description adds the no-play guarantee plus the play_song handoff. For a simple search tool, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters (query, limit) are documented in the schema with examples and bounds. The description adds no parameter syntax or meaning beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('YouTube'da arama yapar') and lists the return fields (başlık, kanal, süre, URL). It also explicitly distinguishes itself from siblings by saying 'Bir şey ÇALMAZ' and naming play_song as the alternative. An agent can tell exactly what this tool does versus play_song.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('Kullanıcı seçenek görmek istediğinde ya da doğru versiyondan emin olmak istediğinde') and when not to use it for playback ('Bir şey ÇALMAZ'). It also names the follow-up alternative: pass the selected URL to play_song.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_musicMüziği durdurAIdempotent
Çalan müziği durdurur ve botu ses kanalından çıkarır.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context beyond annotations by stating the bot is removed from the voice channel, which is a key side effect an agent should know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states both the primary action and its secondary effect without any wasted words. It is appropriately sized for a simple command.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (no parameters, no output schema, annotations present), the description is nearly complete for an agent to call the tool. It could optionally state behavior when no music is playing, but the idempotentHint annotation covers safe repeated calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics to document. Per the rubric, a zero-parameter tool receives a baseline score of 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('durdurur' / stops) and resource ('çalan müziği' / the playing music), and adds the side effect of removing the bot from the voice channel. It inherently distinguishes from siblings like play_song and now_playing, since it is the only stop/disconnect tool in the list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives or prerequisites. It does not mention conditions like 'use when music is playing' or 'call this to end playback and leave the channel', leaving usage context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
list_voice_channels - First observed
now_playing - First observed
play_song - First observed
search_songs - First observed
stop_music
TDQS
Scored across 5 tools
Most tools target distinct actions: list channels, search, play, stop, and status. The main overlap is that play_song can accept a search text and play the first result, which could be confused with search_songs, but descriptions clearly guide the agent to use search_songs for selecting options.
Names mostly follow a snake_case verb_noun pattern (list_voice_channels, search_songs, play_song), but now_playing is not verb-first and stop_music uses 'music' while others use 'song'/'songs'. Minor deviations keep it readable but slightly inconsistent.
Five tools is well within the ideal range and each tool has a clear, non-redundant role in the music playback workflow. No obvious bloat or missing count-driven concerns.
The surface covers basic play, stop, search, and status, but omits common music bot operations like queue management, skip, pause/resume, and volume control. These gaps will cause agent failures for typical user requests, such as skipping a track or adding to a queue.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channel browsing, and playlists for AI agents via MCP.
Generate AI music via the Lacuna Music API from MCP clients like Claude Desktop & Code.
MCP server for Producer/Riffusion AI music generation
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables LLMs to search YouTube for music, download videos as MP3 files, and play audio with playback controls.8-
- AlicenseAqualityCmaintenanceEnables searching, retrieving metadata, and downloading YouTube videos or audio without requiring an API key. It utilizes yt-dlp to support media retrieval and playlist management within MCP-compliant clients like Claude and Cursor.81MIT
- AlicenseAqualityBmaintenanceMCP server for the Spotify Web API — gives Claude and other AI assistants tools to search music, control playback, manage playlists, library, and podcasts.472MIT
- FlicenseNot gradedqualityDmaintenanceAn MCP server for managing YouTube Music playlists via Claude. Add and remove songs, create playlists, and ask Claude to suggest music — all through conversation.-