SpotifyScraper MCP Server
# SpotifyScraper
[](https://aliakhtari.com/spotify/)
[](https://pypi.org/project/spotifyscraper/)
[](https://pypi.org/project/spotifyscraper/)
[](https://pepy.tech/project/spotifyscraper)
[](https://github.com/AliAkhtari78/SpotifyScraper/actions/workflows/ci.yml)
[](https://spotifyscraper.readthedocs.io)
[](https://github.com/AliAkhtari78/SpotifyScraper/pkgs/container/spotifyscraper)
[](https://github.com/AliAkhtari78/SpotifyScraper#reliability--maintenance)
[](https://github.com/AliAkhtari78/SpotifyScraper/blob/master/LICENSE)
[](https://github.com/AliAkhtari78/SpotifyScraper/stargazers)
**Extract public Spotify data โ tracks, albums, artists, playlists, and podcasts โ without the official API or an API key.**
> ๐ง **[Try it live in your browser โ](https://aliakhtari.com/spotify/)** โ paste any Spotify link and watch SpotifyScraper pull typed data, cover art, and a preview, with the exact Python that does it. ([How it was built](https://aliakhtari.com/work/spotify-scraper/).)
SpotifyScraper bootstraps an anonymous token from Spotify's own public embed
pages and reads the same JSON endpoints the web player uses, returning typed,
immutable models. v3 is a ground-up rewrite focused on reliability and a clean,
modern API. Public data needs no login; the opt-in **logged-in** features
(lyrics, podcast transcripts, and account info) add your own Spotify `sp_dc`
cookie โ never a password or an API key.
> **Upgrading from v2?** See the [migration guide](https://spotifyscraper.readthedocs.io/en/latest/migration/). The previous line lives on the [`v2.x` branch](https://github.com/AliAkhtari78/SpotifyScraper/tree/v2.x).
## SpotifyScraper vs. the official API ([spotipy](https://github.com/spotipy-dev/spotipy))
`spotipy` wraps Spotify's **official Web API** โ the right choice when you need to
write to a user's account or read private/library data. SpotifyScraper reads the
**public** data the web player already exposes, so it skips the setup entirely.
| | **SpotifyScraper** | **spotipy** (official API) |
| --------------------------------- | :----------------: | :------------------------: |
| API key / app registration | โ not needed | โ
required |
| OAuth flow | โ not needed | โ
required for most data |
| Rate-limit quota / billing | none | Spotify quota |
| Sync **and** async | โ
| sync only |
| Fully typed, immutable models | โ
| partial |
| Lyrics & podcast transcripts | โ
(cookie) | โ |
| MCP server for Claude / LLM agents | โ
| โ |
| Write / playback / private data | โ (read-only public) | โ
|
Use **spotipy** for authenticated writes and private, market-accurate data; use
**SpotifyScraper** for fast, key-free access to public metadata, lyrics, and
previews โ plus a drop-in **MCP server** for AI agents.
> **Hit by the official-API deprecations?** Spotify's `audio-features`,
> `recommendations`, and `related-artists` endpoints have returned `403` for new
> apps since Nov 2024. SpotifyScraper still returns **related artists and
> recommendations** with no API key. (It can't bring back `audio-features` โ
> Spotify removed that data entirely, from every tool.)
## Install
```bash
pip install spotifyscraper # core (only depends on httpx)
pip install "spotifyscraper[media]" # + cover/preview embedding (mutagen)
pip install "spotifyscraper[browser]" # + Playwright browser fallback & login
pip install "spotifyscraper[cli]" # + the spotifyscraper command-line tool
pip install "spotifyscraper[keyring]" # + store the login cookie in the OS keyring
pip install "spotifyscraper[mcp]" # + the spotifyscraper-mcp MCP server for LLM hosts
pip install "spotifyscraper[all]" # everything
```
Python 3.10+.
## Quickstart
```python
from spotify_scraper import SpotifyClient
with SpotifyClient() as client:
track = client.get_track("https://open.spotify.com/track/4uLU6hMCjMI75M1A2tKUQC")
print(track.name, "โ", track.artists[0].name)
print(track.duration_ms, "ms |", track.preview_url)
print(track.to_dict()) # JSON-safe dict, if you prefer dicts
```
Every entity has its own method โ `get_track`, `get_album`, `get_artist`,
`get_playlist`, `get_episode`, `get_show` โ each accepting a URL, URI, or bare
ID.
### Async
```python
import asyncio
from spotify_scraper import AsyncSpotifyClient
async def main():
async with AsyncSpotifyClient() as client:
track, album = await asyncio.gather(
client.get_track("4uLU6hMCjMI75M1A2tKUQC"),
client.get_album("6N9PS4QXF1D0OWPk0Sxtb4"),
)
print(track.name, "|", album.name)
asyncio.run(main())
```
### Download a cover and preview
```python
from spotify_scraper import SpotifyClient
with SpotifyClient() as client:
track = client.get_track("4uLU6hMCjMI75M1A2tKUQC")
client.download_cover(track, dest="covers/")
client.download_preview(track, dest="previews/", embed_cover=True) # needs [media]
```
### Localized display names
Pass `locale` โ a BCP-47 **language** tag: a bare language subtag (`"de"`,
`"ja"`) or a language-region tag (`"ja-JP"`) โ to localize the *language* of
display names. Set it per client or override it per call:
```python
with SpotifyClient(locale="ja-JP") as client: # default for every call
track = client.get_track("4uLU6hMCjMI75M1A2tKUQC")
other = client.get_track("4uLU6hMCjMI75M1A2tKUQC", locale="de-DE") # per-call wins
```
It is sent as the `Accept-Language` header and changes only how names are
*spelled*. It is **not** a country/market code โ a bare `"US"` is meaningless as
a language and is ignored โ and it does **not** filter regional **availability**
or vary **preview URLs**: anonymous Spotify resolves country from the request IP,
and its pathfinder silently ignores a `market` variable. True market/availability
filtering requires the authenticated Web API, which this library does not
implement; for region-specific results, point the client's `proxy` at the target
region. See the
[localization guide](https://spotifyscraper.readthedocs.io/en/latest/guides/localization/).
## Features
- **All core entities + podcasts** โ tracks, albums, artists, playlists, shows, episodes.
- **Search** across every entity type, returning one typed `SearchResults`.
- **Charts & discovery** โ editorial charts, related artists, full paginated discography, and album recommendations.
- **Cover colors & Canvas** โ extract an artwork's theming palette and download a track's looping Canvas video.
- **Credits & concerts** โ performers/writers/producers and an artist's upcoming live events.
- **Public user profiles** โ `get_user()` (name, follower counts, public playlists).
- **MCP server** โ expose everything to Claude/LLMs via `spotifyscraper-mcp` (batch tools + a one-call `get_track_visuals`); also ships as a container on ghcr.io.
- **Localized display names** โ pass a BCP-47 language tag (`locale`) to set the language of names.
- **Lyrics & podcast transcripts** โ cookie-authenticated, time-synced, one token for both.
- **Browser-assisted login + session persistence** โ log in once, then run headless (no stored passwords).
- **Account-aware** โ `get_account()` / `is_premium()`, plus cookie-free `session_info()`.
- **Batch helpers** โ plural `get_*s([...])` with partial-failure-safe results and managed concurrency.
- **Sync & async** clients sharing one sans-io core.
- **Typed, frozen models** with JSON-safe `to_dict()` / `from_dict()`.
- **Two-tier resilience** โ Spotify's GraphQL API with automatic fallback to the embed page.
- **One core dependency** (`httpx`); media and browser support are optional extras.
- **Optional response cache** โ opt-in, persistent, token-safe (only token-free pathfinder GETs).
- **Anti-ban built in** โ per-host rate limiting, retries with backoff, UA rotation, proxies.
- **Browser fallback** via Playwright when you need a real browser.
## Command line
With the `cli` extra installed, a `spotifyscraper` command is available:
```bash
spotifyscraper track 4uLU6hMCjMI75M1A2tKUQC # entity metadata as JSON
spotifyscraper playlist <id> --max-tracks 50 --pretty
spotifyscraper download preview <id> -o ./previews --embed-cover
```
Every command emits JSON, so it composes with tools like `jq`. See the
[CLI guide](https://spotifyscraper.readthedocs.io/en/latest/guides/cli/).
## Batch helpers
Each getter has a plural sibling (`get_tracks`, `get_albums`, โฆ) that fetches
many inputs and returns one `BatchItem` per input โ index-aligned, and a dead
input never aborts the rest:
```python
items = client.get_tracks(["4uLU6hMCjMI75M1A2tKUQC", "bad-id"])
ok = [i.result for i in items if i.ok]
failed = {i.value: i.error for i in items if not i.ok}
```
The async client runs them concurrently, bounded by `max_concurrency` (default
5). See the [batch guide](https://spotifyscraper.readthedocs.io/en/latest/guides/batch/).
## Response caching
For repeated lookups, enable an opt-in persistent cache. It only stores
**token-free** pathfinder responses โ never the embed pages that carry the
anonymous token โ so no credential is ever written to disk:
```python
from spotify_scraper import SpotifyClient, CacheConfig, FileCache
with SpotifyClient(cache=CacheConfig(store=FileCache())) as client:
client.get_track("4uLU6hMCjMI75M1A2tKUQC") # first call hits the network
client.get_track("4uLU6hMCjMI75M1A2tKUQC") # served from the cache
```
Default TTL is 24h; the `FileCache` is stdlib-only and the backend is pluggable.
See the [caching guide](https://spotifyscraper.readthedocs.io/en/latest/guides/caching/).
## Search
`search()` runs one anonymous, aggregate query across every entity type and
returns a typed `SearchResults`:
```python
from spotify_scraper import SpotifyClient
with SpotifyClient() as client:
results = client.search("daft punk", types=("track", "artist"), limit=5)
print(results.total, "track matches")
for track in results.tracks:
print(track.name, "โ", track.artists[0].name)
```
Hits are sparse (pass an `id` to `get_album()`/`get_show()` for the full entity);
`total` is the track-match count. See the
[search guide](https://spotifyscraper.readthedocs.io/en/latest/guides/search/).
## Lyrics & transcripts
Lyrics and podcast transcripts need a Spotify account cookie (`sp_dc`); the
library handles the token handshake for you, and one cookie powers both:
```python
from spotify_scraper import SpotifyClient
with SpotifyClient(cookies="cookies.txt") as client: # or cookies={"sp_dc": "..."}
lyrics = client.get_lyrics("4uLU6hMCjMI75M1A2tKUQC")
for line in lyrics.lines:
print(line.start_ms, line.text)
transcript = client.get_transcript("07gKzPFkbvGF0cHoeG7ARS") # a podcast episode
for line in transcript.lines:
print(line.start_ms, line.text)
```
Your cookie is sent only to Spotify and never logged. An episode with no
transcript raises `NotFoundError`. See the
[lyrics & cookies guide](https://spotifyscraper.readthedocs.io/en/latest/guides/lyrics-and-cookies/).
## Browser-assisted login
Don't want to copy a cookie by hand? `login()` opens a real browser, you sign in
once, and the captured `sp_dc` is persisted (no password is ever collected or
stored). Later runs reconnect headlessly โ ideal for servers:
```python
from spotify_scraper import SpotifyClient
with SpotifyClient() as client:
client.login() # reuse a valid session, else open a browser
print(client.get_lyrics("4uLU6hMCjMI75M1A2tKUQC").sync_type)
# A later, headless run โ no browser needed:
with SpotifyClient.from_saved_session() as client:
account = client.get_account() # who am I?
print(account.product, account.country, client.is_premium())
transcript = client.get_transcript("07gKzPFkbvGF0cHoeG7ARS")
```
`login()` reuses a valid saved session by default (browser only the first time);
`from_saved_session()` never needs the `browser` extra. The cookie is stored in
an owner-only file, or the OS keyring with `store="keyring"` (the `keyring`
extra). `get_account()`/`is_premium()` report the logged-in account, and
`SpotifyClient.session_info()` checks a saved session without exposing the
cookie. See the
[authenticated sessions guide](https://spotifyscraper.readthedocs.io/en/latest/guides/authentication/).
## Roadmap
**Shipped**
| Version | Scope |
|---------|-------|
| **3.0** | The library: all entities, pagination, media downloads, browser fallback, docs |
| **3.1** | Command-line interface |
| **3.2** | Cookie-authenticated lyrics |
| **3.3** | Cookie-authenticated podcast transcripts (`get_transcript`); browser-assisted login, session persistence & account-awareness (`get_account`/`is_premium`) |
| **3.4** | [Search](https://github.com/AliAkhtari78/SpotifyScraper/issues/129) across every entity type (`search()`) ยท display-language [localization](https://github.com/AliAkhtari78/SpotifyScraper/issues/130) (`locale`) |
| **3.5** | Optional [response cache](https://github.com/AliAkhtari78/SpotifyScraper/issues/131) (`cache=CacheConfig(...)`) ยท [batch helpers](https://github.com/AliAkhtari78/SpotifyScraper/issues/132) with managed concurrency |
| **3.6** | **Visual & discovery**: cover colors, Canvas videos, charts, related artists, paginated discography, recommendations, public profiles, track credits, concerts ยท a best-in-class **MCP server** + container image |
| **3.7** | MCP **batch tools** (`get_tracks`/`get_albums`/โฆ) ยท `get_track_visuals` convenience tool for visual front-ends |
| **3.8** | Maintenance: dependency, toolchain & CI modernization (all Actions on current majors, SHA-pinned) ยท docs & PyPI backlinks |
| **3.9** | Official **MCP registry** publishing (+ Glama/mcp.so/PulseMCP/Smithery discovery) ยท "vs spotipy" comparison ยท one-time, opt-out CLI star hint |
**What's next** โ future ideas are tracked in the GitHub
[milestones](https://github.com/AliAkhtari78/SpotifyScraper/milestones) and
[issues](https://github.com/AliAkhtari78/SpotifyScraper/issues) โ ๐ or weigh in
on the ones that matter most to you. Scope is subject to change.
## Reliability & maintenance
This library rides Spotify's own public endpoints, so it can break when Spotify
changes them. To keep it dependable:
- A **daily canary** runs the live test suite against Spotify. When an endpoint
shifts, it automatically opens a `spotify-breakage` issue (and closes it on
recovery), so regressions surface before they reach you.
- Breakages are triaged and fixed promptly with the help of **[Claude Code](https://claude.com/claude-code)**
(Anthropic's coding agent), under the maintainer's review โ the same
agent-assisted workflow that keeps this project moving. Persisted-query hashes
live in a single file (`api/pathfinder.py`), so a Spotify rotation is a
one-line update.
- Every change runs through `ruff` + `mypy --strict` + a hermetic test suite
(85% coverage floor) across Python 3.10โ3.13 on Linux, macOS, and Windows.
If something is broken for you, please
[open an issue](https://github.com/AliAkhtari78/SpotifyScraper/issues) โ the
monitoring has often caught it already.
## Documentation
Full docs, guides, and the API reference: **<https://spotifyscraper.readthedocs.io>**
The MCP server also ships as a container:
`docker run -p 8000:8000 ghcr.io/aliakhtari78/spotifyscraper` (set `SPOTIFY_SP_DC`
to enable the authenticated tools).
## Legal
SpotifyScraper is an unofficial, independent project, not affiliated with
Spotify. It reads publicly available data and the ~30-second previews Spotify
publishes; it does not download full tracks or circumvent DRM. Use it for
educational and personal purposes, and in line with Spotify's Terms of Service.
See the [legal notice](https://spotifyscraper.readthedocs.io/en/latest/legal/).
## Contributing
Contributions are welcome โ see [CONTRIBUTING.md](https://github.com/AliAkhtari78/SpotifyScraper/blob/master/CONTRIBUTING.md). The project
is developed spec-first with [OpenSpec](https://github.com/Fission-AI/OpenSpec);
specs live in [`openspec/`](https://github.com/AliAkhtari78/SpotifyScraper/tree/master/openspec).
## Star history
If SpotifyScraper saved you the official-API OAuth dance, a โญ helps other
developers find it โ and tells me which features to keep building.
<a href="https://star-history.com/#AliAkhtari78/SpotifyScraper&Date">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/svg?repos=AliAkhtari78/SpotifyScraper&type=Date&theme=dark" />
<img alt="Star history of AliAkhtari78/SpotifyScraper" src="https://api.star-history.com/svg?repos=AliAkhtari78/SpotifyScraper&type=Date" width="600" />
</picture>
</a>
## License
[MIT](https://github.com/AliAkhtari78/SpotifyScraper/blob/master/LICENSE) ยฉ [Ali Akhtari](https://aliakhtari.com) โ full-stack AI engineer ([aliakhtari.com](https://aliakhtari.com)).
<!-- mcp-name: io.github.AliAkhtari78/spotifyscraper -->
TDQS
Scored across 28 tools
Each tool has a clearly distinct purpose: single-entity fetchers, batch fetchers, search, charts, related data, visuals, lyrics, etc. No two tools overlap in functionality; even the composite get_track_visuals is unique.
The predominant pattern is 'get_<entity>' (singular or plural), with a few exceptions like 'list_charts' and 'search'. This is mostly consistent, but the mix of 'get_' and 'list_' and the absence of a prefix for 'search' are minor deviations.
28 tools is on the high side, but each tool serves a specific data retrieval need for Spotify, including batch versions for efficiency. The count is justified given the breadth of Spotify's API, though it could be slightly trimmed.
The tool set covers nearly all read operations for Spotify entities, including batch fetches, charts, visuals, lyrics, and events. Missing are recommendation endpoints and playlist creation/modification, but for a scraper this is acceptable.