Skip to main content
Glama

applemusic-mcp

CI License: MIT

English | 简体中文

An MCP server that lets Claude (or any MCP client) control the Music app on macOS and analyze your library. Just talk to it:

"Play some Leehom Wang" "Make a playlist of the K-pop songs in my library" "What genres do I listen to most?" "Favorite every song in playlists 3, 2, 1, last track first, so Favorite Songs shows them in order"

Everything runs locally on your Mac through JavaScript for Automation (JXA). No Apple Developer account, no API keys, and nothing leaves your machine.

Features

Area

Tools

Playback

now_playing, playback (play / pause / toggle / next / previous / stop), set_player_options (volume, shuffle, repeat)

Search

search_library, search_and_play, play_track

Playlists

list_playlists, get_playlist_tracks, play_playlist, create_playlist, add_to_playlist, remove_from_playlist, compare_playlists

Favorites

set_favorite, favorite_playlist (in order, in batches, optional gap so "sort by Date Favorited" keeps your order)

Stats

listening_stats (overview, top tracks / artists / albums / genres, recently added / played), artist_stats

Limits: search covers your library, not the full Apple Music catalog. Play counts are the lifetime totals the Music app keeps on this Mac, so there's no per-week history. If you mostly listen on your phone, counts may be low.

Related MCP server: Apple Music MCP Server

Requirements

  • macOS with the Music app (tested on a recent macOS, Apple Silicon)

  • Claude Desktop or another MCP client

  • uv. The installer sets it up for you.

Install

One-click

  1. Download or clone this repo, e.g. into ~/Documents/applemusic-mcp.

  2. Double-click install.command. If macOS blocks it, right-click > Open. It installs uv, runs a read-only self-test, and adds the server to Claude Desktop's config (it backs up the old config first).

  3. When macOS asks whether the app may control Music, click OK.

  4. Quit Claude Desktop with Cmd+Q, reopen it, and ask "What's playing?"

Manual

git clone https://github.com/bwu109-netizen/applemusic-mcp.git
cd applemusic-mcp
uv sync
uv run python scripts/smoke_test.py   # read-only check

Then add this to ~/Library/Application Support/Claude/claude_desktop_config.json, using your own paths (which uv):

{
  "mcpServers": {
    "apple-music": {
      "command": "/Users/you/.local/bin/uv",
      "args": ["run", "--directory", "/Users/you/applemusic-mcp", "python", "-m", "applemusic_mcp.server"],
      "env": { "PYTHONPATH": "/Users/you/applemusic-mcp/src" }
    }
  }
}

Troubleshooting

  • "Not authorized to send Apple events": open System Settings > Privacy & Security > Automation and turn on Music for Claude (or your terminal).

  • Server keeps disconnecting / No module named applemusic_mcp: if the project lives in an iCloud-synced folder like Documents, iCloud can break a .venv inside it. install.command keeps the venv in ~/.local/share/applemusic-mcp/venv instead. Re-run it.

  • Changes to the code don't show up: Claude Desktop starts the server once. Quit it with Cmd+Q and reopen.

How it works

Claude ──MCP (stdio)──▶ server.py ──osascript -l JavaScript──▶ Music.app
  • src/applemusic_mcp/jxa.py runs JXA via osascript. User input is passed as JSON in argv, never pasted into script source, so song titles with quotes can't break or inject code. Errors come back as readable MCP tool errors.

  • src/applemusic_mcp/server.py defines the tools. Library reads use bulk property fetches (one Apple Event per property, not per track), so even large libraries stay fast.

  • src/applemusic_mcp/stats.py is plain Python aggregation over a 5-minute library cache.

Development

uv sync
uv run pytest

The tests mock the Music app, so they run on Linux too (CI uses Ubuntu). They also run node --check on every generated JXA script to catch syntax errors before they reach a Mac.

PRs welcome. Ideas: Apple Music catalog search via MusicKit, AirPlay device selection, ratings, smart-playlist creation.

License

MIT

Available Tools

17 tools
add_to_playlistB

Add tracks (by id) to a playlist. Tracks already in the playlist are skipped.

ParametersJSON Schema
NameRequiredDescriptionDefault
playlistYes
track_idsYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses an idempotency-adjacent behavior ('tracks already in the playlist are skipped'), but omits whether the playlist must already exist, required permissions, and whether missing track ids cause failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the core action front-loaded and the dedup caveat following. Appropriately sized for the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and 0% parameter coverage, the description is thinner than ideal. It covers the action and one behavior but leaves error handling, playlist identification, and permissions unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that track_ids are IDs, but leaves the 'playlist' parameter ambiguous (name vs. id vs. URI) and says nothing about ID format or playlist resolution.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (add) and resource (tracks to a playlist), clearly distinguishing it from siblings like remove_from_playlist and create_playlist. It doesn't explicitly reference those siblings, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as create_playlist or remove_from_playlist, nor any prerequisite conditions. The agent must infer usage entirely from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

artist_statsC

Everything about one artist in the library: track count, total plays, albums, top songs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
artistYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Everything about one artist' is vague and doesn't specify whether it returns all tracks if no limit is given, what the top songs are based on, or any behavioral details like pagination or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the resource and lists key outputs. It's concise but could be more structured with separate when-to-use information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% schema coverage for parameters, the description is incomplete. It should explain the limit parameter and perhaps what 'everything' includes. For a stats tool, it lacks crucial details like whether results are paginated or what the response format is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description does not explain the 'limit' parameter at all. The word 'top songs' hints at a limit but doesn't clarify if 'limit' controls the number of top songs returned. The 'artist' parameter is obvious, but the other parameter is left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (artist) and enumerates the data returned (track count, total plays, albums, top songs). It's clear what the tool does, though it doesn't contrast with the sibling listening_stats, which likely also provides stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives like listening_stats or search_library. An agent has to infer that this is for artist-specific statistics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_playlistsB

Tracks in playlist that are NOT in any of the against playlists (read-only).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
againstYes
playlistYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose '(read-only)', which is a useful safety signal beyond the bare schema, but says nothing about pagination behavior driven by `limit`, result ordering, or the size/shape of the returned set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the core semantics appear immediately and the read-only marker is efficiently parenthesized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter description coverage, the description covers the essential comparison semantics but leaves gaps around the `limit` parameter, return shape, and result ordering that an agent calling this tool would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully clarifies the two required parameters by tying `playlist` and `against` to the set-difference semantics, but the `limit` parameter (default 200) is never mentioned, leaving one of three params undocumented in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific operation and precise semantics: it returns tracks in `playlist` that are NOT in any of the `against` playlists, which is a set-difference. An agent can distinguish this from sibling tools like get_playlist_tracks or compare-style operations. It stops short of naming a sibling it displaces, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no when-not-to-use, and no mention of alternatives such as get_playlist_tracks. The purpose is implied by the semantics but the agent is left to infer when comparison is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_playlistB

Create a new empty playlist. Fails if one with that name already exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
descriptionNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose two genuine behavioral traits: the created playlist is empty, and duplicate names cause failure. It does not state required permissions, whether the new playlist is returned (e.g., its ID), or what happens to the optional description field, leaving notable gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler: the primary action first, then the failure condition. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter creation tool with no output schema and no annotations, the description covers the core action and one failure mode. It is still thin on parameter meaning and on what the caller receives back, which matters given there is no output schema to fall back on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does not mention either parameter. "Empty playlist" indirectly hints that no track list is accepted, but nothing explains the required `name` (uniqueness, length) or the optional `description` (its effect or null default).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Create a new empty playlist") and adds the meaningful scope qualifier "empty," which distinguishes it from tools that add tracks. Siblings like add_to_playlist and play_playlist are clearly different verbs, so explicit differentiation isn't required, but the description doesn't reference any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Fails if one with that name already exists" is a useful precondition, implying the tool is for net-new playlists and that callers should check for duplicates first. However, there is no explicit when-to-use guidance and no alternative named (e.g., how to add tracks afterward or how to handle the duplicate case).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

favorite_playlistA

Mark a playlist's tracks as Favorite one by one, in batches.

With reverse=true it walks from the LAST track toward the first. start is the position in that walking order; call again with the returned next_start until done=true. Tracks already favorited are skipped (keeps their earlier timestamp). gap_sec (default 0) optionally spaces out timestamps if the Music app's timestamps turn out to be too coarse to keep the order.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNo
startNo
gap_secNo
reverseNo
playlistYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description does heavy lifting: it discloses the skip-already-favorited behavior with timestamp preservation, the reverse traversal, and the continuation protocol via next_start/done. It omits permission requirements, failure handling, and any rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the paging contract, then the gap_sec caveat. Minor redundancy in the trailing 'if the Music app's timestamps turn out to be too coarse' clause, but nothing egregious.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param mutation tool with no annotations and no output schema, the description covers the traversal, continuation, and idempotency behavior an agent needs. Remaining gaps are count semantics and error/partial-failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it explains start (position in walking order), reverse (direction), and gap_sec (timestamp spacing with rationale). Only count and playlist go unexplained, but playlist is self-evident from the name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Mark a playlist's tracks as Favorite') and clarifies the batched, one-by-one scope, which an agent can distinguish from single-track favorites. It does not explicitly contrast with the sibling set_favorite, so it falls short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage mechanics: call again with the returned next_start until done=true, reverse walks from the last track, and gap_sec is for coarse timestamps. It never says when to prefer this over set_favorite or play_playlist, so no explicit alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_playlist_tracksC

List tracks in a playlist (by exact name).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
playlistYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but says nothing about whether this is a read-only listing, whether it paginates (the limit param defaults to 100), or what happens when the exact name is not found. 'List' implies a read, but nothing confirms it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource front-loaded and the matching constraint appended. No waste, though it is terse to the point of under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and 0% schema coverage, so the description should be doing more work. It omits the return shape (track fields, ordering), pagination behavior, and the failure mode for a non-matching name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters. The description clarifies the 'playlist' argument must match an exact name, which adds real value, but the 'limit' parameter (default 100) is completely unaddressed in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List tracks in a playlist') that is clearly distinct from sibling list_* tools like list_playlists. The parenthetical narrows the lookup semantics, but no sibling is named to sharpen the boundary against list_playlists or play_playlist.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(by exact name)' note is the only usage signal, and it constrains matching rather than telling the agent when to pick this tool over list_playlists, play_playlist, or compare_playlists. No prerequisites, no when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

listening_statsA

Analyze the library: play counts, favorite artists/albums/genres, recent activity.

Play counts are lifetime totals tracked by the Music app (no per-period history). Library data is cached for 5 minutes; pass refresh=true to re-read it.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNooverview
limitNo
refreshNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does deliver real facts: play counts are lifetime-only with no per-period history, and library data is cached for 5 minutes with refresh=true forcing a re-read. It still omits that this is read-only and any rate/permission constraints, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, tightly front-loaded: what it analyzes first, then the lifetime-totals caveat, then the caching/refresh behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, so the description bears the load; it covers scope, data semantics, and caching well. It leaves the view modes and return shape undefined, a modest gap for a read-only stats tool with an enum selector.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains refresh=true behavior clearly and implies what the stats cover, but says nothing about the 'view' enum values (the primary selector) or how 'limit' applies to each view.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Analyze the library' with enumerated dimensions (play counts, favorite artists/albums/genres, recent activity). This distinguishes it from artist_stats, which is artist-scoped, though the description never explicitly differentiates from that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the library-wide scope, but there is no explicit statement of when to use this versus artist_stats or search_library, and no exclusions. The agent must infer the routing from the sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_playlistsA

List the user's own playlists (regular and smart), with track counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden; it does disclose the scope limit ('the user's own', excluding followed/shared playlists) and the payload shape (regular and smart, with track counts). It omits pagination, ordering, and whether unauthenticated access fails, so it is helpful but not thorough for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the action and resource and qualifying scope in parentheses. No filler, nothing wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param list tool with an output schema that already documents the return shape, the description covers scope and content sufficiently. Minor gaps around ordering/pagination keep it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema has nothing to document and the baseline is 4. The description correctly implies no filtering inputs are available.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('List ... playlists') and adds scope detail: the user's OWN playlists, both regular and smart, with track counts. This distinguishes it from siblings like get_playlist_tracks and search_library implicitly, but it never names an alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance and no mention of alternatives such as search_library or get_playlist_tracks. Usage is only inferable from the verb itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

now_playingA

What's playing right now, plus player state, volume, shuffle and repeat.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully enumerates the returned state (player state, volume, shuffle, repeat), which is valuable given there is no output schema, but it never states that the call is read-only with no side effects or what happens when nothing is playing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the primary payload first and appends the secondary state fields. No filler, no restating of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description's main obligation is to say what comes back, and it does list the returned fields. It stops short of covering the empty/idle state or confirming read-only behavior, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. The baseline of 4 applies; the sentence is not misleading about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific query and its payload: 'What's playing right now, plus player state, volume, shuffle and repeat.' That distinguishes it from the control-oriented siblings (playback, set_player_options) without naming any of them explicitly, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the phrase 'right now' signals a read of current state, and the sibling 'set_player_options' implies this is not the tool for changing volume/shuffle/repeat. There is no explicit when-to-use, when-not-to-use, or named alternative, so it sits at minimum viable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playbackC

Control playback. Returns the player state afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose that player state is returned, which is a small plus, but says nothing about side effects (e.g., whether stop clears the queue), permission requirements, error behavior, or what "player state" actually contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler; the tool's purpose and its return behavior come first. It is efficient, though the brevity is partly the result of under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a player-control tool with no annotations and no output schema, the definition is too thin: it never describes the returned player state, session requirements, or the semantics of the destructive-ish actions (stop/next). An agent has enough to guess a call but not enough to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single action parameter, and it adds nothing. The enum makes the legal values self-evident, but the description never explains behavioral differences between them (e.g., toggle vs. play, or what stop does versus pause).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Control playback" states a resource and a general verb, but it does not distinguish this tool from siblings like play_track, now_playing, or set_player_options. The specific actions (play/pause/next/stop) live only in the schema enum, so the description alone reads as a vague catch-all for playback manipulation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over play_track, play_playlist, or now_playing, nor any stated preconditions (e.g., an active player session). The agent must infer usage entirely from the name and enum values.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

play_playlistB

Start playing a playlist by name, optionally setting shuffle first.

ParametersJSON Schema
NameRequiredDescriptionDefault
shuffleNo
playlistYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose one ordering trait ('setting shuffle first'), but says nothing about what happens to the current playback/queue, whether the change is reversible, error behavior on an unknown playlist, or whether playback is local or remote.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the primary action and appending the optional modifier. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the essentials are present, but with zero annotation coverage and zero schema descriptions the description should say more about side effects on existing playback and failure modes. It is minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it partially does: 'by name' clarifies how the required playlist parameter is matched, and 'optionally setting shuffle first' conveys that shuffle is a pre-playback toggle. It still omits the expected name format (exact vs fuzzy) and the meaning of the null default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start playing a playlist by name'), which is clearly distinguishable from siblings like play_track or list_playlists on inspection. However, it never names an alternative tool, so the agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this over search_and_play, play_track, or playback, nor any prerequisite such as needing an exact playlist name. The 'by name' phrasing hints at an identifier requirement but stops short of stating what happens if the name is ambiguous or absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

play_trackC

Play a specific track by its id (from search_library or get_playlist_tracks).

ParametersJSON Schema
NameRequiredDescriptionDefault
track_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but only says 'Play a specific track'. It doesn't disclose whether playback is immediate, requires authentication, or what happens on invalid track IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is concise and front-loaded with the core action and required input. No unnecessary detail, though the parenthetical source hint could be seen as slightly cluttered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with no annotations or output schema, the description provides basic purpose but leaves out key behavioral details like auth requirements, playback behavior, or error handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with one parameter. The description clarifies the track_id source ('from search_library or get_playlist_tracks'), which adds some value, but it doesn't describe the parameter's expected format or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Play) and resource (a specific track), clearly distinguishing it from siblings like play_playlist or search_and_play. It doesn't fully differentiate from search_and_play, but the 'by its id' framing implies a direct play versus a search-based one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the track id should come from search_library or get_playlist_tracks, giving some retrieval context. However, it doesn't state when to use this tool over alternatives like play_playlist or search_and_play.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_from_playlistA

Remove tracks (by id) from a playlist. Does not delete them from the library.

ParametersJSON Schema
NameRequiredDescriptionDefault
playlistYes
track_idsYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It explicitly states that tracks are not deleted from the library, which is critical non-obvious behavior. It does not mention whether the operation is idempotent, requires specific permissions, or what happens if track IDs are invalid, but the key side-effect distinction is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the main action and immediately followed by the crucial side-effect clarification. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a simple two-parameter tool with no output schema and no annotations, the description covers the essential action and the key distinction from library deletion. It could be improved with brief info on failure modes or return values, but it is mostly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It specifies that tracks are removed 'by id', which clarifies the expected format of track_ids, but does not elaborate on the playlist parameter (which is required but presumably a playlist identifier). Adequate but not rich.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (remove) and resource (tracks) and scope (from a playlist) with a clear contrast clause. It distinguishes itself from the likely sibling add_to_playlist and clarifies it is not a library deletion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it (removing tracks from a playlist) and clarifies the inverse of add_to_playlist, which is present in the sibling list. It lacks explicit guidance on when not to use it or alternative tools like delete from library, but the scope is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_and_playB

Search the library and immediately play the best match. Also returns other candidates.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoall
queryYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden, and it does surface one important side effect: the tool immediately starts playback and also returns other candidates. However, it says nothing about what happens to currently playing audio, how a 'best match' is chosen, or what happens when no match exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero padding: the action is stated first and the secondary return behavior second. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description partially covers returns ('also returns other candidates'), but it omits the kind filter semantics and the effect on existing playback state. For a tool that both reads and mutates playback, that is a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate, and it does not. The 'kind' enum (all/songs/artists/albums) is never mentioned, and 'query' is only obliquely implied by 'Search the library', leaving the filter parameter undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific compound verb (search + play) and resource (library), and the phrase 'immediately play the best match' distinguishes it from the sibling search_library, which presumably only searches. It stops short of naming the alternatives explicitly, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to use this combined tool versus search_library plus play_track, and no prerequisites or exclusions are given. The combined search-and-play intent is only implied by the sentence structure.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_libraryC

Search the user's Music library. Returns tracks with ids you can play or add to playlists.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoall
limitNo
queryYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that tracks with ids are returned, but says nothing about pagination despite the limit parameter, result ordering, or whether library mutations are involved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight, front-loaded sentences with no padding. The brevity is efficient, though it borders on under-specification rather than true conciseness for a three-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage, the description is the only source of behavioral detail and it delivers very little. An agent cannot tell how to constrain or interpret results beyond the bare purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain query, kind, and limit, but it mentions none of them. The kind enum values (all/songs/artists/albums) are self-describing in the schema, but the meaning of limit and the expected query format are left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Search the user's Music library') and even sketches the downstream use of results ('ids you can play or add to playlists'). It does not, however, explicitly distinguish itself from the sibling search_and_play, so an agent must infer which search tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the mention that results can be played or added to playlists, which hints at a discovery-then-act flow. There is no explicit when-to-use, when-not-to-use, or named alternative such as search_and_play or play_track.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_favoriteA

Mark (or unmark) tracks as Favorite, one at a time in the given order.

gap_sec waits between tracks so each gets a distinct timestamp. Keep calls under ~40 tracks when using a gap so the call doesn't time out.

ParametersJSON Schema
NameRequiredDescriptionDefault
gap_secNo
favoriteNo
track_idsYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load and does disclose real behavioral traits: sequential application, gap_sec producing distinct timestamps, and a concrete timeout threshold (~40 tracks). It omits things like auth/permission requirements or what happens to invalid track ids, but the timeout and ordering disclosure is genuine value beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action before the gap/timeout caveat. Slightly awkward line break before 'gap_sec waits...' but no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param mutation with no annotations and no output schema, the description covers order semantics, the gap mechanism, and the timeout limit, which is most of what an agent needs to call it safely. Only the success/return behavior and parameter typing remain unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three parameters. It explains favorite (mark/unmark), gap_sec (inter-track wait for distinct timestamps), and implies track_ids and ordering, but adds no format/type detail for the ids and never states the default polarity of the favorite flag.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Mark (or unmark) tracks as Favorite') and makes the dual polarity explicit. It implicitly separates itself from the sibling favorite_playlist by scoping to tracks, but never names the sibling, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives operational guidance on how to call it (one at a time, in order, keep under ~40 tracks when using a gap), which implies usage context. It does not, however, say when to prefer this over favorite_playlist or how it relates to add_to_playlist/remove_from_playlist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_player_optionsA

Set volume (0-100), shuffle on/off, and/or repeat mode. Omit what you don't want to change.

ParametersJSON Schema
NameRequiredDescriptionDefault
repeatNo
volumeNo
shuffleNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses PATCH-style merge semantics (omitted fields are left untouched), which prevents an agent from zeroing out settings it did not intend to change. It says nothing about scope (current session vs. persisted preference), required playback state, error behavior, or return value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both load-bearing: the first specifies the fields and a constraint, the second states the partial-update rule. Scope and fields are front-loaded with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-optional-parameter setter with no output schema, the description covers the essential call contract. It omits preconditions and the effect surface (does volume apply to the current playback session or persist?), which an agent would want before invoking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description documents all three parameters by name and adds the volume range 0-100, which the schema does not express. Shuffle on/off and the repeat mode's three states are conveyed, though the concrete repeat values (off/one/all) are left to the schema enum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (set) and enumerates exactly what gets changed: volume, shuffle, repeat. An agent can immediately tell this is a playback-option mutator, distinct from siblings like playback, now_playing or play_track. It stops short of explicitly naming an alternative tool, so it doesn't reach the top of the scale.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Omit what you don't want to change" gives real partial-invocation guidance, which is the main usage question for a tool whose parameters are all optional. However, it never says when to reach for this versus playback or now_playing, nor whether a player/session must already be active, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.1.0
    • First observedadd_to_playlist
    • First observedartist_stats
    • First observedcompare_playlists
    • First observedcreate_playlist
    • First observedfavorite_playlist
    • First observedget_playlist_tracks
    • First observedlist_playlists
    • First observedlistening_stats
    • First observednow_playing
    • First observedplay_playlist
    • First observedplay_track
    • First observedplayback
    • First observedremove_from_playlist
    • First observedsearch_and_play
    • First observedsearch_library
    • First observedset_favorite
    • First observedset_player_options

TDQS

B3.2/5.0

Scored across 17 tools

Disambiguation4/5

Most tools target distinct resources/actions (now_playing vs playback vs set_player_options; search_library vs search_and_play). Minor overlap exists between set_favorite and favorite_playlist (single vs batch) and between playback and set_player_options for player state, but descriptions clarify the boundaries well.

Naming Consistency4/5

Consistent snake_case throughout, but the convention mixes verb_noun (play_track, create_playlist, add_to_playlist) with noun-first/state names (now_playing, playback, listening_stats, artist_stats). Readable and predictable despite the minor stylistic split.

Tool Count4/5

17 tools is slightly above the ideal 3-15 sweet spot, but each covers a distinct facet of music playback, library search, playlists, favorites, and stats, so it earns its place rather than feeling padded.

Completeness4/5

Strong coverage of playback, search, playlist lifecycle (create/list/get/add/remove/play), favorites, and analytics. Minor gaps: no delete/rename playlist and no explicit queue management, but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Integrates Apple Music with MCP clients to search the global catalog, manage personal playlists, and access library data. It enables users to perform actions like creating playlists, adding tracks, and viewing recommendations through natural language commands.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Apple Music playback control, library search, playlist management, and queue operations via MCP.
    2
    MIT