Skip to main content
Glama

spotify-mcp

A dependency-free Model Context Protocol server for Spotify. It speaks stdio JSON-RPC 2.0, uses only the Python standard library, and gives an MCP client (Claude Code, Claude Desktop, anything that speaks MCP) 36 tools for playback, queue, search, library and bulk playlist management. A small CLI is included.

Built by Robert Lingoes with AI coding agents (Claude Code / Codex); Robert owns the architecture, requirements and review.

Features

  • Zero dependencies. Python 3.9+ and nothing to pip install to run it.

  • Bulk playlist work. Add any number of tracks by name or URI; names are auto-resolved, de-duplicated against the playlist, and chunked to Spotify's 100-per-call limit (10,000-track playlist ceiling is enforced up front).

  • Safe by default. Destructive tools (remove, empty, delete, unsave, unfollow) return a dry-run preview unless you pass confirm=true.

  • Robust auth. Access tokens are minted from your refresh token, cached with 0600 permissions, refreshed under a cross-process lock, and retried once on 401. 429 responses honor Retry-After.

  • Flexible inputs. Anywhere an id is expected you can pass a raw id, a spotify: URI, or an open.spotify.com URL.

Related MCP server: YouTube MCP Server

Tools

Group

Tools

Playback

now_playing, play, pause, skip, previous, volume, seek, queue, devices, transfer

Search and discovery

search, resolve, artist_catalog, recommendations, audio_features, recent_tracks, top_tracks

Playlists

playlist_list, playlist_get, playlist_create, playlist_add, playlist_remove, playlist_reorder, playlist_rename, playlist_set_details, playlist_empty, playlist_delete

Library and follows

saved_tracks, save_tracks, unsave_tracks, saved_albums, save_albums, unsave_albums, followed_artists, follow_artists, unfollow_artists

Playback tools need Spotify Premium and an active device.

Setup

1. Create a Spotify developer app

  1. Go to the Spotify developer dashboard and create an app.

  2. Under the app's settings add this Redirect URI (exact match):

    http://127.0.0.1:8888/callback
  3. Copy the app's Client ID and Client Secret.

2. Get a refresh token

export SPOTIFY_CLIENT_ID="your-client-id"
export SPOTIFY_CLIENT_SECRET="your-client-secret"
python3 auth_setup.py            # use --no-browser to just print the URL

auth_setup.py runs the standard authorization-code flow (with PKCE) against a temporary server on 127.0.0.1:8888, then prints a refresh token once. It stores nothing on disk. Use a different port by setting SPOTIFY_REDIRECT_URI (and registering it in the dashboard); it must be a loopback address.

Scopes requested:

user-read-playback-state        user-modify-playback-state
user-read-currently-playing     user-read-recently-played
user-read-playback-position     user-top-read
playlist-read-private           playlist-read-collaborative
playlist-modify-public          playlist-modify-private
user-library-read               user-library-modify
user-follow-read                user-follow-modify

3. Configure credentials

The server reads three environment variables:

Variable

Meaning

SPOTIFY_CLIENT_ID

your app's client id

SPOTIFY_CLIENT_SECRET

your app's client secret

SPOTIFY_REFRESH_TOKEN

the token printed by auth_setup.py

Optional:

Variable

Default

Meaning

SPOTIFY_MCP_CONFIG_DIR

~/.config/spotify-mcp

where the token cache (tokens.json) lives

SPOTIFY_MCP_TOKEN_CACHE

on

set to off to keep tokens in memory only

Spotify may rotate the refresh token on refresh; the cache keeps the newest one. If you paste a new SPOTIFY_REFRESH_TOKEN into your config, the stale cache is ignored automatically.

Never commit these values. Keep them in your MCP client config or a secrets manager.

Use with Claude Code

claude mcp add spotify \
  --env SPOTIFY_CLIENT_ID=your-client-id \
  --env SPOTIFY_CLIENT_SECRET=your-client-secret \
  --env SPOTIFY_REFRESH_TOKEN=your-refresh-token \
  -- python3 -m spotify_mcp

Run it from a checkout of this repo (so python3 -m spotify_mcp resolves), or pip install . first, which installs a spotify-mcp command; then the last line is simply -- spotify-mcp.

Equivalent .mcp.json:

{
  "mcpServers": {
    "spotify": {
      "command": "python3",
      "args": ["-m", "spotify_mcp"],
      "cwd": "/path/to/spotify-mcp",
      "env": {
        "SPOTIFY_CLIENT_ID": "your-client-id",
        "SPOTIFY_CLIENT_SECRET": "your-client-secret",
        "SPOTIFY_REFRESH_TOKEN": "your-refresh-token"
      }
    }
  }
}

Use with Claude Desktop

Add to claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\) and restart the app:

{
  "mcpServers": {
    "spotify": {
      "command": "python3",
      "args": ["-m", "spotify_mcp"],
      "cwd": "/path/to/spotify-mcp",
      "env": {
        "SPOTIFY_CLIENT_ID": "your-client-id",
        "SPOTIFY_CLIENT_SECRET": "your-client-secret",
        "SPOTIFY_REFRESH_TOKEN": "your-refresh-token"
      }
    }
  }
}

After pip install . you can replace command/args/cwd with "command": "spotify-mcp".

CLI

The same operations are available from the shell (python3 -m spotify_mcp.cli ..., or spotify-mcp-cli once installed):

spotify-mcp-cli now
spotify-mcp-cli search "daft punk" --limit 5
spotify-mcp-cli create "Road Trip" --desc "windows down"
spotify-mcp-cli add <playlist_id> "Blinding Lights" "Mr Brightside by The Killers"
spotify-mcp-cli add <playlist_id> --file songs.txt        # one song per line
cat songs.txt | spotify-mcp-cli add <playlist_id> --stdin
spotify-mcp-cli remove <playlist_id> "some song" --confirm

Development Mode caveats

Spotify limits apps in "Development Mode", and its Web API has been changing. Depending on your app's tier, some endpoints may return 403 (for example certain playlist-write, library-save or follow calls), and recommendations / audio_features have been restricted for newer apps; those two tools return an explanatory error instead of failing hard. A 403 is passed through unchanged so you can see exactly what Spotify said. This project does not work around Spotify's access policies.

Tests

No network, no credentials: all HTTP is mocked.

python3 -m pytest          # or: python3 -m unittest discover -s tests -v

License

MIT. See LICENSE.

Available Tools

36 tools
artist_catalogC

An artist's top tracks or album catalog (name/URI/URL accepted).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
limitNo
artistYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only implies a read operation. It says nothing about auth requirements, rate limits, pagination, whether limit has a default or cap, or what the response looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the dual mode and accepted input formats come first. It is efficient, though terseness contributes to the documentation gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 3-parameter tool with no annotations, no output schema, and 0% schema description coverage needs more than one line. Missing limit semantics, kind default, and return shape leave an agent unable to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must document all three parameters. It adds useful detail for artist (name/URI/URL accepted) and implies the kind values via 'top tracks or album catalog', but the limit parameter and any default for kind are left entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Identifies the resource (an artist's catalog) and its two modes — top tracks or albums — which map directly to the kind enum. It is clear what is fetched, but it never states a verb and does not differentiate itself from the sibling top_tracks tool, so an agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as top_tracks, search, or resolve. The parenthetical about accepted input formats is parameter detail, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_featuresB

Audio features (tempo/energy/danceability) for a track (Spotify has restricted this endpoint for newer apps; may return an explanatory error).

ParametersJSON Schema
NameRequiredDescriptionDefault
trackYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it does disclose one important trait: Spotify has restricted the endpoint for newer apps and it may return an explanatory error. That is genuinely useful operational context. However, it omits read-only nature, auth/scope requirements, and rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the returned fields and the caveat front-loaded in parentheses. Every clause earns its place; only the absent parameter format is a gap.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with no output schema, the description covers the value and the deprecation caveat reasonably well. The unresolved input format for 'track' and the absence of any return-shape hint leave it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single required parameter 'track' is undocumented in both schema and description. The description says 'for a track' but never clarifies whether the value is an ID, URI, or name, which is a real ambiguity for the agent. It falls well short of compensating for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (audio features) and enumerates the returned attributes (tempo/energy/danceability) for a track. No sibling tool covers this ground, so the agent can distinguish it without opening the schema. It lacks an explicit verb like 'get', but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus recommendations, top_tracks, or other track-metadata tools, and no prerequisites stated. The only contextual note is about endpoint restriction, not about tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

devicesA

List available playback devices (speakers, desktop apps, phone).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only list, but does not state side effects, auth/permission requirements, or whether it identifies the active device. Adequate for a simple zero-param read tool, but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with useful parenthetical examples. It is appropriately sized and every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-param list tool with no output schema, the description covers what is returned and gives device examples. It could mention return structure or active-device status, but it is adequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is an empty object, so there are no parameter semantics to add. The baseline for a zero-param tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('playback devices'), with examples of device types. Clear enough for selection, but does not explicitly distinguish this from any sibling tool beyond the obvious device-list purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance. The phrase 'available playback devices' implies it is the lookup step before device-targeted actions like transfer or play, but that is left to inference rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

follow_artistsB

Follow artists (needs follow scope).

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one real behavioral trait: the required OAuth scope. However, for a mutation it omits effect on existing data, idempotency on duplicate follows, and failure behavior, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the action and immediately followed by the auth constraint. No filler, though it is arguably terse to the point of under-specification for a write operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one array param, no output schema) so the minimal description is close to sufficient, and the auth scope is covered. But ambiguity remains about whether 'items' truly accepts artist names/URIs given the schema text says 'Track names', and nothing describes the result of the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single required parameter, so the schema already documents 'items' fully. The description adds nothing about batching or limits. Note a schema mismatch worth flagging: the items field says 'Track names OR URIs' while the tool follows artists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (follow) and resource (artists), which an agent can distinguish from unfollow_artists and followed_artists by verb alone. It stops short of explicitly contrasting with those siblings, so it is clear but not sibling-differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is the parenthetical 'needs follow scope', which is a prerequisite rather than a when-to-use condition. It never says when to prefer this over followed_artists (checking) or unfollow_artists (reversing), nor what happens on repeat calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

followed_artistsC

Artists you follow (needs follow scope — needs the user-follow-* scopes).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, but it does disclose a genuine behavioral trait: the required user-follow-* scopes. Beyond that it says nothing about pagination, ordering, or result shape for what is presumably a paged list endpoint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, which is good, but the parenthetical repeats the scope requirement verbatim ('needs follow scope — needs the user-follow-* scopes'), wasting the only sentence of substance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list-style tool with one optional parameter and no output schema, the definition is still too thin: no indication of what the agent gets back, how limit affects results, or how this differs from the sibling follow_artists. Only the auth scope is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (limit) with 0% schema description coverage, so the description must compensate and does not — it never mentions limit, its default, or its accepted range. The parenthetical 'needs follow scope — needs the user-follow-* scopes' restates the same requirement twice instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource ('artists you follow') and implies a read/list operation, but it is a noun fragment with no verb. It is separable from follow_artists/unfollow_artists only by name inference, not by anything stated in the text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance or mention of alternatives such as follow_artists (which adds follows) or artist_catalog (which looks up artist data). The only condition given is an auth prerequisite, not a usage condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

now_playingA

What's currently playing on Spotify (track, artist, album, progress).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It helpfully discloses the return fields, which matters because there is no output schema, but it does not state what happens when nothing is playing (null, error, empty) or any auth/premium constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the resource stated first and the return fields in a compact parenthetical. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param read tool the description is nearly complete, and listing the return fields compensates for the absent output schema. The only real gap is undefined behavior when no track is playing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is no parameter syntax for the description to explain or omit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-and-resource ('What's currently playing on Spotify') and enumerates the returned fields (track, artist, album, progress). It is inherently distinct from siblings like recent_tracks or saved_tracks, though it never names an alternative to sharpen that boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrasing implies the use case — check the current playback state — but gives no explicit when-to-use guidance or contrast with siblings such as recent_tracks or top_tracks. For a zero-param read tool the usage is fairly self-evident, so implied guidance is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pauseB

Pause playback.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses nothing beyond the action itself: no auth requirements, no effect on playback state or queue, and no failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with zero waste. This is appropriately sized for a zero-parameter command.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter tool, the description is minimally adequate but leaves usage context and behavioral effects unspecified. Given no annotations or output schema, it should do more to be fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so per the rubric the baseline is 4. The description adds no parameter information because there are none to describe.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (pause) and resource (playback), making the tool's action clear. However, it does not differentiate from siblings such as play, skip, or previous, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use guidance, no alternatives, and no prerequisites. The description only restates the action, leaving the agent to infer appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playB

Resume playback, or play a specific track/album/playlist URI.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full burden. It doesn't disclose that playback requires an active device, whether it interrupts current playback, error behavior for invalid URIs, or any auth/premium constraints. For a playback control tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is efficient, though the compact 'or' phrasing is what leaves the optional/required semantics implicit.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one optional param, no output schema), so the bar is modest, and the description covers both modes at a high level. However, it omits device targeting and the resume-vs-play distinction tied to the optional parameter, which an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only types 'uri' as a string, so the description must compensate. It does clarify that the uri refers to a track, album, or playlist, which adds real meaning, but it never states the parameter is optional or what omitting it does (resume), leaving the most important semantic ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Resume playback', 'play') and resource (track/album/playlist URI), making the two modes of operation clear. It doesn't name the sibling it pairs with (pause) or contrast with skip/previous, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'or' construction implies two usage modes: call with no uri to resume, call with a uri to start specific content. That is useful implied guidance, but there is no explicit statement of when to prefer this over queue/skip/previous, nor any prerequisites (e.g. active device, premium account).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_addB

Bulk-add songs (names or URIs) to a playlist. Auto-resolves names, dedupes, and chunks any number of songs up to Spotify's 10k/playlist ceiling.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
dedupeNo
playlist_idYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses auto-resolution of names, implicit deduping, and internal chunking up to Spotify's 10k ceiling, but omits key mutation behavior: whether tracks are appended to the end, ordering guarantees, partial-failure handling, and required auth scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences that front-load the action, then the resolution/dedupe/chunking behavior and the capacity limit. Every clause earns its place and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the batch mechanics reasonably well but leaves gaps: no return/confirmation semantics, no error behavior on unresolvable names, and no permission requirements. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% and the description does not compensate. The items semantics it describes are already in the schema's parameter description; 'dedupes' is ambiguous against the boolean dedupe parameter (is dedupe always on, or opt-in, and what is its default?), and playlist_id goes unexplained anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Bulk-add songs ... to a playlist') and adds meaningful scope detail (name/URI auto-resolution, dedupe, 10k ceiling, chunking). It is easily distinguished from mutation siblings like playlist_remove, playlist_reorder, or playlist_empty, though it never explicitly contrasts an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. 'Bulk-add' and the 10k ceiling imply this is the tool for large batch insertions, but the description never tells the agent when to prefer this over, say, searching/resolving first, nor what happens if the playlist is at capacity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_createC

Create a new playlist.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
publicNo
descriptionNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It implies a mutation ('create') but does not state permissions required, whether the playlist name must be unique, default visibility, or what happens on conflict—significant gaps for a creation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and wastes no words. It is efficient, though its extreme brevity leaves the tool under-specified for a three-parameter creation operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and zero schema description coverage, the description does almost nothing to fill the gap. It omits parameter semantics, behavioral details, and any context an agent needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for three parameters (name, public, description), and the description adds no meaning or format details. An agent cannot tell whether 'public' defaults to false or how 'name' is validated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource ('Create a new playlist'), making the tool's basic action unambiguous. However, it does not differentiate from sibling playlist operations (e.g., playlist_rename, playlist_set_details) or specify scope, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. The description merely restates the action without helping an agent choose between playlist_create and related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_deleteA

Delete (unfollow) a playlist. DESTRUCTIVE — needs confirm=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo
playlist_idYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the two most important traits — the operation is DESTRUCTIVE and requires confirm=true — but omits what exactly is destroyed (library entry vs. playlist object), whether it is reversible, and what happens if confirm is absent or false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse sentences with zero filler; the destructive warning and the confirm requirement are front-loaded where an agent will read them before invoking.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no annotations, no output schema, and 0% parameter coverage, the description covers the essential safety gate but leaves the blast radius and reversibility unstated. Adequate to prevent an accidental destructive call, incomplete for understanding the outcome.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for two undocumented parameters. It explains the semantics of confirm ("needs confirm=true"), which is the non-obvious one, but adds nothing about playlist_id's expected format or source. Partial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Delete a playlist") and the parenthetical "(unfollow)" clarifies that this removes the playlist from the user's library rather than emptying its tracks, which usefully separates it from the sibling playlist_empty. It does not name that sibling explicitly, so differentiation is implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a hard precondition ("needs confirm=true"), which is genuine usage guidance, but says nothing about when to choose this over playlist_empty or playlist_remove, nor about irreversibility or follow-up steps. Usage context is implied by the destructive framing rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_emptyA

Remove ALL tracks from a playlist. DESTRUCTIVE — needs confirm=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo
playlist_idYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It clearly discloses the destructive nature and the confirmation requirement, which are essential behavioral traits beyond the schema. It does not state irreversibility or whether the playlist container itself is preserved, but the core safety information is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by the critical warning. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive two-parameter tool with no annotations and no output schema, the description covers the essential action and safety requirement. It could go further by clarifying that the playlist itself remains or describing irreversibility, but it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It adds meaningful semantics for 'confirm' (must be true) but says nothing about 'playlist_id', though that name is largely self-explanatory. Partial compensation yields a baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Remove ALL tracks from a playlist') and the 'ALL' scope implicitly distinguishes it from playlist_remove (partial removal) and playlist_delete (deletes the playlist itself). It does not explicitly name those siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a critical usage prerequisite ('needs confirm=true') but provides no guidance on when to use this tool versus alternatives like playlist_remove or playlist_delete. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_getC

Get a playlist's details + tracks (paginated).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
playlist_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses only that results are paginated; it says nothing about authentication requirements, the shape of the returned details, embedding/ownership nuances, or error behavior for invalid playlist IDs. That is thin for a read tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler; the parenthetical flags pagination without wasting words. Every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage, the description should compensate but does not. It omits pagination mechanics, ID format, and any behavioral caveats needed to call this correctly on the first try.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so all three parameters (playlist_id, limit, offset) are undocumented in both schema and description. The single word 'paginated' hints that limit/offset exist but defines nothing about defaults, ranges, or the required playlist_id format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (a playlist's details + tracks), so the agent knows this returns a single playlist with its track listing rather than a list of playlists. It does not explicitly distinguish itself from the sibling playlist_list, but the singular 'a playlist's' scope makes the distinction inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and never references the obvious alternatives (playlist_list for enumeration, playlist_create/add/remove for mutation). The agent must infer usage entirely from the name and the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_listC

List your playlists (with ids + track counts).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'List' implies a read, but nothing is said about authentication requirements, pagination through large playlist collections, ordering, or the effect of omitting limit. Only the return shape is hinted at.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. It is appropriately sized for a simple list tool, though the parenthetical return note is the only content beyond the verb+resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-arg, no-output-schema read tool, the description tells the agent what it gets back (ids + track counts), which is genuinely useful. But with no annotations and an undocumented limit parameter, it stops short of being complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'limit' is undocumented in both schema and description. The description mentions ids and track counts, which is return-value information rather than parameter semantics, so it does not compensate for the missing 'limit' meaning, bounds, or default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('playlists') and adds the returned shape (ids + track counts). It is distinguishable from playlist_get (single playlist) and playlist_create, but it never names those siblings to route the agent explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus playlist_get or saved_tracks, no prerequisites (e.g. auth/scope required for a user's playlists), and no mention of what to do if the list is empty or large.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_removeA

Remove songs from a playlist. DESTRUCTIVE — needs confirm=true, else returns a dry-run preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
confirmNo
playlist_idYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full safety burden and does disclose the critical trait: the operation is destructive and is gated behind confirm=true, with a dry-run preview as the fallback. It stops short of covering auth requirements, reversibility, or what happens when a listed track isn't in the playlist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the destructive warning and confirm requirement front-loaded in the second sentence. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-required-parameter mutation tool with no output schema and no annotations, the definition covers the essential behavior an agent needs: what it does and the safety gate. Minor gaps remain around permissions and the behavior on non-matching items, but nothing blocking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only items is documented in the schema), so the description must compensate, and it does for confirm — explaining that omitting it yields a dry-run preview rather than a deletion. playlist_id remains undocumented but is self-evident; items' name/URI/URL flexibility is already handled by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Remove songs from a playlist"), which is distinguishable from siblings like playlist_empty, playlist_delete, and unsave_tracks without opening the schema. It does not, however, explicitly contrast itself with those siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (removing individual songs from a playlist) but never states when to prefer this over playlist_empty (clear all) or playlist_delete. The confirm=true note gives operational context but is not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_renameC

Rename a playlist.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
playlist_idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, but it only implies a mutation. It says nothing about permissions required, ownership constraints, whether the change is reversible, or how duplicates/conflicts are handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no wasted words. It is appropriately sized, though the terseness reflects under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and 0% schema description coverage on two required parameters, the description is far too thin. An agent lacks the information needed to invoke it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is documented in the schema. The word 'rename' weakly implies that 'name' is the new value, but playlist_id's format and the name's constraints (length, uniqueness) are entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Rename) and resource (playlist), which is clear on its own. However, it does not distinguish itself from the sibling playlist_set_details, which plausibly also edits playlist metadata, so the boundary is ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus playlist_set_details or other playlist mutation tools. Nothing about preconditions or context is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_reorderC

Move a track (or range) to a new position in a playlist.

ParametersJSON Schema
NameRequiredDescriptionDefault
playlist_idYes
range_startYes
range_lengthNo
insert_beforeYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Move' implies a mutation, but the description omits whether this requires auth/scopes, how positions are indexed (0-based vs 1-based), whether insert_before is absolute or relative, and what happens to surrounding tracks. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is well-sized for the literal content, though the content itself is thin rather than the structure being at fault.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and four parameters at 0% schema coverage, the description should clarify indexing, permissions, and ordering semantics. It does not, leaving an agent unable to call the tool confidently beyond the basic intent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are four parameters, three required. The description vaguely gestures at 'range' and 'new position' but never names or explains playlist_id, range_start, range_length (optional), or insert_before. It does not compensate for the schema's lack of descriptions, leaving insert_before and range semantics opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Move') and resource ('a track (or range)') and the effect ('to a new position in a playlist'). This clearly distinguishes it from siblings like playlist_add, playlist_remove, and playlist_empty. It stops short of naming an alternative, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of when an agent should prefer playlist_add/remove or vice versa. The description only says what the tool does, not when to select it over the many sibling playlist_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

playlist_set_detailsC

Update a playlist's name / description / public flag.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
publicNo
descriptionNo
playlist_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Update' implies mutation, but nothing is said about permissions required, reversibility, whether omitted fields are left untouched or cleared, or the effect of flipping the public flag. Significant gaps for a write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding; the slash-separated field list is compact. Slightly terse given how much remains unexplained, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with no annotations, no output schema, 0% parameter coverage, and an overlapping sibling (playlist_rename) needs more than one sentence. Neither the ownership requirement nor the partial-update behavior is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully enumerates the three mutable fields, but omits the required playlist_id and gives no semantics such as what boolean public values mean or string format expectations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Update') plus resource ('a playlist') plus the fields it touches (name/description/public flag) makes the purpose immediately legible. It does not differentiate from the sibling playlist_rename, which plausibly overlaps on the name field.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus playlist_rename, no prerequisites (e.g., must own the playlist or be a collaborator), and no statement about whether fields are patched or replaced. The agent must guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

previousB

Go to the previous track.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and does almost nothing. It does not say what happens at the start of a queue (restart current track vs. seek to prior track), whether active playback is required, or whether a device must be selected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler. It is concise but so terse that it lacks any of the routing or behavioral detail a playback control tool typically benefits from.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, no-annotation, no-output-schema playback command, the description leaves key edge cases unexplained (queue boundaries, required playback/device state) and offers no sibling disambiguation. It is barely enough to identify the tool, not enough to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to document beyond the target of the action. The baseline for a no-parameter tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and resource ('Go to the previous track'), so the core action is unambiguous. However, it does not differentiate itself from the closely related sibling 'skip' (likely next track), which an agent must infer from name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus 'skip', 'play', 'pause', or 'seek', which are adjacent playback controls in the sibling list. Usage is only weakly implied by the phrase 'previous track'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queueC

Add a track URI to the playback queue.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it discloses almost nothing. It doesn't say where in the queue the track is inserted, whether it requires an active playback session or device, whether it requires an auth scope, or what happens if no device is active.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, which is structurally clean. It is arguably under-specified rather than concise, but every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and an undocumented parameter, the description is too thin. It omits the operational context (active device/session, insertion position, error/return behavior) an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate, and it does only minimally by labeling the parameter as a 'track URI' (implying a Spotify track URI rather than a plain ID or name). It supplies no format example or constraint, so the agent gets marginal added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Add a track URI to the playback queue'), so the agent immediately knows the operation. However, it does not differentiate from siblings like 'play' (which presumably starts playback) or 'playlist_add', which matters a lot for a queueing tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use 'queue' versus 'play', nor any prerequisites such as needing an active device or playback session. The agent must infer the distinction entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recent_tracksC

Recently played tracks.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. 'Recently played' implies a read-only, recency-ordered list, but there is no mention of auth requirements, ordering, or whether results are scoped to the current user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short but not concise in a useful sense: it is a bare fragment rather than a front-loaded sentence with an action and scope. It under-specifies rather than trimming waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A one-parameter tool with no annotations and no output schema needs the description to explain operation, scoping, and limit behavior. A four-word fragment leaves every one of those gaps open.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single limit parameter is undocumented in both the schema and the description. No default, maximum, or meaning for limit is given, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment names the resource (recently played tracks) but supplies no verb, so the agent must infer that it lists them. It is distinguishable from siblings like top_tracks or now_playing, but only because of the noun phrase, not because the description states the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this rather than top_tracks, saved_tracks, or now_playing, nor any prerequisites or exclusions. The description provides no usage context at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommendationsB

Recommendations from seed tracks/artists/genres (Spotify has restricted this endpoint for newer apps; may return an explanatory error).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
seed_genresNoTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
seed_tracksNoTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
seed_artistsNoTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so description carries full burden. It discloses a key behavioral trait: the endpoint may be restricted and return an explanatory error. However, it omits other important behaviors like authentication requirements, rate limits, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the core function and immediately followed by a critical caveat. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description is somewhat complete by mentioning the endpoint restriction, but it misses usage context and detailed behavior. Adequate as a minimum viable description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with seed parameters documented as accepting names OR URIs OR URLs. The description adds no further parameter semantics beyond what the schema already states. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb-less but specific resource: recommendations generated from seed tracks/artists/genres. Distinct from siblings like search or artist_catalog. Lacks an explicit action verb but the purpose is understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. Does not mention prerequisites like at least one seed required, nor how it compares to search or top_tracks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolveB

Resolve a track name (e.g. 'Blinding Lights by The Weeknd') to its URI.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It implies a read-only lookup returning a URI but says nothing about ambiguity handling (multiple matches), no-match failure, or whether it selects the best/first result — significant gaps for a name-to-ID resolver.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the verb and output front-loaded and an inline example that earns its place. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter lookup with no output schema, the description covers the input shape and the return type (URI), which is most of what is needed. It is still missing failure/ambiguity behavior, which matters for a resolution tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter is a bare 'text' string, so the description must compensate. It does add real meaning with the format hint and example ('Blinding Lights by The Weeknd'), clarifying the expected 'title by artist' input, but it does not state accepted variations or matching rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('resolve') and resource ('track name') plus the output form ('its URI'), with a concrete example. It does not distinguish itself from the sibling 'search', which could plausibly be used for the same task, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus alternatives such as 'search', nor any prerequisite or precondition. Usage is only weakly implied by the purpose sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_albumsC

Save albums (ids/URIs/URLs).

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether the save is idempotent, what happens on duplicate saves, whether authentication is required, or that this is a mutating operation — all of which matter for a 'save' tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short parenthetical, so there is no padding, but it is arguably under-specified rather than appropriately concise. The one clause is front-loaded but does not earn enough informational value for a mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with full schema coverage this could be adequate, but with no annotations and no output schema the description should at least disclose the mutation effect and duplicate-handling behavior. That behavioral context is entirely absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single 'items' parameter thoroughly, including auto-resolution of names. The description's '(ids/URIs/URLs)' largely repeats the schema and actually diverges from it, which says 'Track names OR URIs OR open.spotify.com URLs' — so it adds no meaning and introduces a minor inconsistency.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Save albums'), which is clear and actionable. However, it does nothing to distinguish itself from the closely related siblings save_tracks, saved_albums, and unsave_albums, so an agent gets no help telling them apart beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, no prerequisites, and no reference to alternatives like save_tracks for individual tracks. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

saved_albumsC

Your saved albums.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Your saved albums' implies a read operation but does not confirm it is read-only, nor does it describe pagination behavior for the limit parameter or any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, but it is under-specified rather than concise. It consists of a single fragment that does not earn its place by conveying actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a query tool with no annotations, no output schema, and an undocumented parameter, the description is inadequate. An agent cannot reliably determine the tool's behavior or the meaning of its parameter from this text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'limit' parameter is undocumented in both the schema and the description. The description adds no meaning about what limit controls or its default/range.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Your saved albums' merely restates the tool name rather than stating a verb+resource action. It does not clarify that it retrieves/lists the user's saved albums, and it fails to distinguish itself from siblings like saved_tracks, save_albums, or unsave_albums.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as save_albums or saved_tracks. The agent is left to infer context entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

saved_tracksD

Your Liked Songs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses nothing. It does not state ordering, pagination across the liked-songs collection, authentication requirements, or what a response contains, leaving the agent blind on a retrieval operation over a potentially large list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short, but this is under-specification rather than conciseness - a four-word fragment that earns no information. There is no front-loaded statement of what the tool does, so brevity here provides no benefit to the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, an undocumented parameter, and no usage guidance, the definition is completely inadequate for a retrieval tool inside a large, cluttered sibling set. An agent cannot reliably distinguish it from saved_albums or decide how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter (limit) has 0% schema description coverage, and the description adds no meaning whatsoever about it. Nothing explains whether limit caps returned tracks, defaults, or its maximum, so the schema gap is entirely uncompensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Your Liked Songs.' restates the tool name (saved_tracks) as a noun fragment without a verb, so it never says what operation is performed (list? fetch? save?). It implies retrieval of the user's liked tracks, but it is essentially a tautological label rather than a statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all: no mention of when this differs from sibling tools such as saved_albums, top_tracks, recent_tracks, or save_tracks, and no prerequisites or exclusions. The agent must infer everything from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_tracksB

Add songs (names or URIs) to Liked Songs.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses little beyond the mutation itself. It never states whether duplicates are ignored or rejected, whether there is a per-call item limit, whether authentication is required, or what the call returns. For a library-mutating tool with zero annotation coverage this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, front-loading the action and the destination collection. Nothing needs to be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with full schema coverage and no output schema, the definition is adequate to invoke correctly. It falls short on the mutation-specific context an agent would want (duplicate handling, batch limits), which the absent annotations leave uncovered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameter is documented as accepting track names, URIs, or open.spotify.com URLs with auto-resolution. The description repeats only the names/URIs portion and adds no syntax, batch-size, or resolution-failure guidance beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Add) and target collection (Liked Songs) with the accepted item forms, so an agent can distinguish it from unsave_tracks and saved_tracks. It doesn't explicitly contrast itself with those siblings, but the direction of the operation (add vs. remove vs. list) is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent must infer that this is the tool for adding to the liked library rather than unsave_tracks or saved_tracks. There is no explicit when-to-use, prerequisite, or routing statement naming an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seekB

Seek to a position (milliseconds) in the current track.

ParametersJSON Schema
NameRequiredDescriptionDefault
msYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not say whether a track must be actively playing, whether seeking works while paused, what happens with no active device/track, or whether the position is clamped to track length. For a playback-mutating control, that is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action and including the unit inline. Nothing is wasted, though it is minimal enough that it reads as under-specified rather than deliberately tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers the core action and unit, but omits the operational preconditions (active playback, error behavior) that an agent would need to invoke it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the sole parameter 'ms' is undocumented in the schema, so the description must compensate. It does the essential job by naming the unit (milliseconds) and the reference point (position in the current track), which is the key semantic an agent needs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Seek to a position') plus resource ('the current track'), and it implicitly distinguishes itself from sibling transport controls like skip, previous, and play by acting on a position within the current track rather than changing tracks. It stops short of naming those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus skip/previous/play, nor any prerequisite such as an active playback session. Usage is only inferable from the word 'current track'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skipA

Skip to the next track.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the entire disclosure burden. It says nothing about whether playback must be active, what happens at the end of a queue, whether the skip is reversible (aside from 'previous'), or any side effects, which is a notable gap for a mutation of playback state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero waste. Nothing could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial zero-parameter playback control with no output schema and no annotations, the description is nearly sufficient. The only omission is behavioral context such as what happens at the end of the queue.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify beyond the schema. Baseline 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (skip to next track) that clearly distinguishes it from the adjacent sibling 'previous', which moves backward. It is not differentiated from other playback siblings like 'play' or 'pause', but the action itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no alternatives are named. However, usage is strongly implied by the verb in a well-known playback-control domain, so an agent can infer it applies during active playback.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

top_tracksC

Your top tracks over short/medium/long term.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
time_rangeNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and discloses almost nothing: no indication that it is a read-only personalization call, no defaults for limit or time_range, no pagination or result-shape context. It adds only the existence of three time windows, which the schema enum already implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no wasted words and the resource front-loaded, but brevity here reflects under-specification rather than efficient communication; nothing about defaults or usage is stated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool with no output schema, the description should at minimum state defaults (limit, time_range) and that it returns ranked tracks for the authenticated user. None of that is present, leaving the agent guessing at call semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It gestures at the time_range enum values ('short/medium/long term') but adds no default, format, or meaning, and completely omits the limit parameter. Half the parameter surface is undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the resource ('your top tracks') and scope ('short/medium/long term'), which implicitly distinguishes it from siblings like saved_tracks (library) and recent_tracks (history). However, no explicit verb and no direct sibling comparison, so an agent must infer that this returns ranked listening stats rather than a saved collection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g., user auth), and no routing to alternatives such as recent_tracks for history or recommendations for discovery. The agent is left to infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transferC

Move playback to a specific device by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
device_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states the action but omits key details such as required permissions, whether current playback state is preserved, and what happens if the device is unavailable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It communicates the core action efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and zero schema description coverage, the description should provide more context. It lacks usage guidance, parameter sourcing, and behavioral details needed for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the only parameter is 'device_id'. The description says 'by id', which minimally clarifies the parameter, but it does not specify the ID format or how to obtain a valid device ID (e.g., via the 'devices' tool).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Move playback') and resource ('to a specific device by id'), making the action clear. However, it does not explicitly distinguish itself from siblings like 'play' or 'devices', so an agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as 'play' or 'devices'. The description implies moving playback but does not state prerequisites, conditions, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unfollow_artistsA

Unfollow artists. DESTRUCTIVE — needs confirm=true (needs follow scope).

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
confirmNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the key traits: it is DESTRUCTIVE, it is gated behind confirm=true, and it requires the follow scope. It omits what is actually removed (unfollow vs. remove from library) and any return/confirmation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short fragments with the destructive warning front-loaded; every clause earns its place. Slightly terse phrasing ('needs confirm=true (needs follow scope)') costs it polish but not clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter destructive tool with no annotations and no output schema, the description covers the critical safety and auth facts an agent needs. The main remaining gap is the meaning/resolution behavior of 'items'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'items' is documented in the schema while 'confirm' is not. The description compensates for 'confirm' by explaining it must be true and why, but says nothing about 'items' (and the schema's 'track names' wording for an artist tool is unaddressed).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Unfollow artists'), which is immediately distinguishable from the sibling 'follow_artists' by polarity. It does not explicitly name the sibling or scope beyond that, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete prerequisite: confirm=true must be supplied and the follow scope is required, which tells the agent when the call will succeed. It stops short of stating when to prefer this over other library-mutating siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unsave_albumsB

Remove saved albums. DESTRUCTIVE — needs confirm=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
confirmNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It usefully flags 'DESTRUCTIVE' and the confirm=true requirement, but omits auth/permission needs, reversibility details, and what actually happens to the saved entries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler: the action first, then the destructive warning. Appropriately sized for a simple two-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no annotations and no output schema, the description discloses the key risk and the confirm requirement but leaves auth needs, reversibility, and return behavior unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'items' is well documented in the schema (names/URIs/URLs, auto-resolved) but 'confirm' is not. The description partially compensates by stating confirm=true is needed, adding meaning beyond the bare boolean.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Remove saved albums'), which cleanly distinguishes it from save_albums and unsave_tracks by name. It does not explicitly name the sibling alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case (removing saved albums) and states the confirm=true requirement, but gives no when-to-use/when-not guidance or explicit routing to alternatives like unsave_tracks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unsave_tracksA

Remove songs from Liked Songs. DESTRUCTIVE — needs confirm=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesTrack names OR URIs OR open.spotify.com URLs — names are auto-resolved.
confirmNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose the two most important traits: the operation is DESTRUCTIVE and requires confirm=true. It stops short of stating that removal is reversible (tracks can be re-saved), what happens on an unresolvable name, or any auth scope requirements, so the disclosure is valuable but partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and immediately followed by the destructive warning and gate. Every clause earns its place with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter destructive tool with no annotations and no output schema, the description covers the essentials: what is removed, that it is destructive, and that confirm is required. Minor gaps remain around confirm's default value and error behavior on unresolved track names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% — items is well documented in the schema (names, URIs, or URLs, auto-resolved), while confirm has no schema description. The description compensates for confirm by stating confirm=true is needed, but does not give a default, accepted false behavior, or any detail beyond the schema's own items description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "Remove songs from Liked Songs." The phrase "Liked Songs" and "songs" cleanly separates it from unsave_albums, though it does not name that sibling explicitly. Clear enough for an agent to select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies one real usage condition — confirm=true is required — but gives no guidance on when to prefer this over siblings like playlist_remove or unsave_albums, and no note on what happens if confirm is omitted. Usage is implied by the purpose statement rather than explained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

volumeB

Set playback volume 0-100.

ParametersJSON Schema
NameRequiredDescriptionDefault
pctYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the 0-100 scale but says nothing about whether playback/device must be active, how out-of-range values are handled, whether 0 mutes, or whether the change is persisted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the range constraint front-loaded; nothing is redundant or padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter playback control with no annotations and no output schema, the definition covers the core action but omits the operational context an agent needs: an active device/session requirement, error behavior, and mute semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate: it supplies the meaning and valid range (0-100) for the single 'pct' parameter, which is the key semantic gap. However, it adds nothing about units beyond percentage, integer rounding, or behavior at the boundaries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Set') and resource ('playback volume') with the valid range, which cleanly separates it from siblings like seek, play, and pause. An agent can identify this tool's job without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives, nor any precondition such as an active playback session or device. The only implied usage is the obvious 'change volume' reading of the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 36 tool updatesv1.0.0
    • First observedartist_catalog
    • First observedaudio_features
    • First observeddevices
    • First observedfollow_artists
    • First observedfollowed_artists
    • First observednow_playing
    • First observedpause
    • First observedplay
    • First observedplaylist_add
    • First observedplaylist_create
    • First observedplaylist_delete
    • First observedplaylist_empty
    • First observedplaylist_get
    • First observedplaylist_list
    • First observedplaylist_remove
    • First observedplaylist_rename
    • First observedplaylist_reorder
    • First observedplaylist_set_details
    • First observedprevious
    • First observedqueue
    • First observedrecent_tracks
    • First observedrecommendations
    • First observedresolve
    • First observedsave_albums
    • First observedsave_tracks
    • First observedsaved_albums
    • First observedsaved_tracks
    • First observedsearch
    • First observedseek
    • First observedskip
    • First observedtop_tracks
    • First observedtransfer
    • First observedunfollow_artists
    • First observedunsave_albums
    • First observedunsave_tracks
    • First observedvolume

TDQS

C2.7/5.0

Scored across 36 tools

Disambiguation4/5

Most tools target a clearly distinct resource or action, and the playback/library/playlist clusters are cleanly separated. The main overlap is playlist_rename versus playlist_set_details (both can change a playlist's name), and resolve versus search have related but distinguishable purposes.

Naming Consistency3/5

Naming mixes conventions: single verbs (play, pause, skip), noun-first (saved_tracks, now_playing, top_tracks), verb_noun (save_albums, follow_artists), and noun_verb for the playlist_* family. The playlist_ prefix is internally consistent and everything stays readable, but there is no single predictable pattern across the server.

Tool Count3/5

36 tools is on the heavy side, but the domain (playback control, library, discovery, and full playlist CRUD) genuinely warrants many operations. It sits at the borderline where a few tools (e.g. separate rename and set_details) could be consolidated.

Completeness4/5

Coverage is broad and coherent: playback, library management, search/discovery, and playlist lifecycle all have create/read/update/delete paths with sensible destructive guards. Minor gaps remain, e.g. shuffle/repeat mode toggles and dedicated album/artist detail retrieval.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to search and read YouTube data and safely manage playlists, including creating private playlists from ranked song matches with preview-before-commit playlist mutations.
    1
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables MCP clients to build and edit Spotify playlists from natural language descriptions.
    7
    81 npm
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables MCP clients to search SoundCloud and retrieve tracks, playlists, and users, as well as perform authenticated actions such as likes, reposts, follows, comments, playlist editing, and uploads with multi-account OAuth support.
    36
    MIT