Discover Weekly: What Spotify Teaches Recommenders
Executive TL;DR
- Discover Weekly was a Hack Week project, not a research breakthrough. It combined three ordinary model families — collaborative filtering on listening behaviour, text analysis of what the internet writes about music, and models that listen to the raw audio — and shipped them as a weekly ritual.
- The product wrapper did as much work as the models: a fixed arrival time, a finite list, and a personal framing. It rolled out to 1% before 100 million users, and by month ten Spotify reported 40 million unique listeners and 5 billion tracks streamed.
- For most teams, start with collaborative filtering on implicit signals plus approximate nearest-neighbour search, delivered on a schedule. Add content embeddings for cold start and exploration for diversity — in that order, not all at once.
Every few months a founder asks for “a Discover Weekly for our product." What they really mean is our catalog is bigger than anyone can browse, engagement is flat, and personalization seems like the answer. Instinct is right. The mental model of how Spotify got there is often wrong.
Discover Weekly didn’t start life in a research lab. The service came about as a concept during the company’s annual Hack Week, Spotify said, and the engineers who made it rolled it out to employees first, then to 1% of users, before getting to a base of about 100 million. The modeling below was a mishmash of well-documented approaches: collaborative filtering over listening behavior, natural-language analysis of blogs, reviews, and playlist titles, and models that analyze the audio itself—capabilities that Spotify reinforced by acquiring The Echo Nest in 2014.
It was the packaging that made it land. 30 songs, not an endless feed. Monday morning, same time, every week. A tone that said, You were born for this. The product was the batch schedule that sounds like a technical limitation: Spotify engineer Edward Newett has described refreshing 100 million playlists every Sunday night on about a terabyte of new data. Spotify’s reported results were not subtle—40 million unique users and 5 billion tracks streamed in the first ten months, and more than 8,000 artists got over half their monthly listeners from the playlist.
So the useful question for your team is not "How do we build Spotifys model?" It is which recommender approach costs at your data volume and what ritual you wrap around it.
Collaborative filtering, content embeddings, a hybrid with exploration, or a managed service — which recommender should you build first?
| Collaborative filtering | Content embeddings | Hybrid + exploration | Managed service | |
|---|---|---|---|---|
| Cost to first version | Lowest — days to a few weeks with an open-source ALS implementation and an ANN index | Moderate — embedding generation plus a vector store; higher if you process audio or video | Highest — two candidate sources, a ranker, and an experimentation layer to tune | Low upfront, rising with usage and catalogue size |
| Cold start | Poor — a new item with no interactions is invisible until someone finds it | Strong — a brand-new item has an embedding the moment it exists | Strong — content covers what behaviour cannot | Varies by vendor; usually needs your metadata to be clean |
| Data you must already have | Meaningful implicit signals: plays, saves, repeat views, purchases | Rich item content: text, audio, images, structured attributes | Both, plus enough traffic to run experiments | Whatever the vendor's schema expects, in their format |
| Ops burden | Low — a scheduled batch job and an index rebuild | Moderate — embedding pipelines and a vector index to keep fresh | High — multiple models, feature freshness, online serving, monitoring | Lowest day to day; you inherit their incidents instead |
| Lock-in | None — the artefacts are your vectors | Low, unless you depend on one embedding provider | Low, but high internal complexity to unwind | High — your ranking quality and roadmap sit inside someone else's product |
| Explainability and control | Good — 'people who played this also played that' is defensible to users and to legal | Good — similarity reasons are inspectable | Hardest to explain, easiest to tune once instrumented | Weakest — limited insight into why an item ranked |
| Fails when | The catalogue turns over fast or the long tail never gets seen | Behaviour, not content, drives taste — lookalike items that nobody actually wants together | You do not yet have the traffic to tell two rankers apart | Personalisation is your core product rather than a feature |
When to pick each option
If you have interaction data and a catalog that is not changing on a daily basis, start with collaborative filtering. Using an implicit-feedback matrix factorization on a plays-or-views matrix, plus an approximate nearest-neighbour index for retrieval, this system is a genuine weekend-to-fortnight build using mature open-source libraries. And it is the approach whose output you can describe to a skeptical stakeholder in a sentence. This is where the heart of Discover Weekly’s signal lives—behavior, at scale.
If your real problem is cold start, choose content embeddings: a marketplace where sellers add items hourly, a news product where yesterday’s article is worthless, or a video library where new uploads need to be discoverable immediately. Text and audio models put it in the space before anyone touches a new item. Which is why Spotify marries behavior with text and audio analysis—a track nobody has streamed yet still has to be recommendable.
If you have both signals and sufficient traffic to measure, go hybrid with exploration. The shape is candidate generation from multiple sources, one ranker that scores them together, and deliberate exploration so the system doesn’t collapse into the same twenty items. Don't begin here. A ranker with nothing to rank is a month's work that improves nothing.
Use a managed service when recommendations are a secondary feature and you want to buy the pipeline—an example is an e-commerce “you may also like” row. Trade-off: you get speed now, but control later is limited. If the product you are selling is personalization then the extra weeks to own the vectors are worth itpersonalization,.
Rendering diagram…
# pip install implicit annoy scipy
# Inputs: one row per interaction (user_id, item_id, weight).
# Weight is confidence, not rating: a save counts more than a view.
import numpy as np, scipy.sparse as sp
from implicit.als import AlternatingLeastSquares
from annoy import AnnoyIndex
FACTORS = 64
def build(interactions, n_users, n_items):
rows = [u for u, _, _ in interactions]
cols = [i for _, i, _ in interactions]
vals = [w for _, _, w in interactions]
matrix = sp.csr_matrix((vals, (rows, cols)), shape=(n_users, n_items))
model = AlternatingLeastSquares(
factors=FACTORS, regularization=0.05, iterations=20
)
model.fit(matrix) # user_factors, item_factors
index = AnnoyIndex(FACTORS, 'dot') # dot product matches the ALS objective
for item_id, vec in enumerate(model.item_factors):
index.add_item(item_id, vec)
index.build(20)
return model, matrix, index
def weekly_list(user_id, model, matrix, index, seen, k=30, explore=3):
# Over-fetch, then filter: the recommendation you must not show is
# the one the user already consumed.
raw = index.get_nns_by_vector(model.user_factors[user_id], k * 5)
picks = [i for i in raw if i not in seen][: k - explore]
# Reserve a few slots for deliberate novelty, or the list converges
# on the same catalogue corner every week.
tail = [i for i in raw[k * 2 :] if i not in seen and i not in picks]
picks += list(np.random.choice(tail, size=min(explore, len(tail)), replace=False))
return picks
# Run it on a schedule, write the result per user, and serve the stored list.
# Precomputing is not a compromise here — it is what makes a fixed weekly slot possible.We had to refresh 100 million playlists every Sunday night— Edward Newett
MVP development planner
Scope, budget, and launch plan
Choose the product, platforms, and modules. The range, phases, and team update as you go. It is a planning figure, not a quote.
Newsletter
One architecture teardown or evaluation, every other week
Practical write-ups on framework choices, migrations, and scaling decisions – written for people who have to justify the call in a meeting.
Next step
Book a 15-min Architecture Review
Walk through your system with a lead architect and leave with a scoped recommendation.
Schedule on Cal.com