Case Study•
October 7, 2026
•
7 min read
•
0 views

Discover Weekly: What Spotify Teaches Recommenders

karmakoders Team
Design & Engineering
Diagram of a recommendation architecture with collaborative filtering, content embeddings, candidate generation, ranking and exploration

Executive TL;DR

  • Discover Weekly was a Hack Week project, not a research breakthrough. It combined three ordinary model families — collaborative filtering on listening behaviour, text analysis of what the internet writes about music, and models that listen to the raw audio — and shipped them as a weekly ritual.
  • The product wrapper did as much work as the models: a fixed arrival time, a finite list, and a personal framing. It rolled out to 1% before 100 million users, and by month ten Spotify reported 40 million unique listeners and 5 billion tracks streamed.
  • For most teams, start with collaborative filtering on implicit signals plus approximate nearest-neighbour search, delivered on a schedule. Add content embeddings for cold start and exploration for diversity — in that order, not all at once.

Every few months a founder asks for “a Discover Weekly for our product." What they really mean is our catalog is bigger than anyone can browse, engagement is flat, and personalization seems like the answer. Instinct is right. The mental model of how Spotify got there is often wrong.

Discover Weekly didn’t start life in a research lab. The service came about as a concept during the company’s annual Hack Week, Spotify said, and the engineers who made it rolled it out to employees first, then to 1% of users, before getting to a base of about 100 million. The modeling below was a mishmash of well-documented approaches: collaborative filtering over listening behavior, natural-language analysis of blogs, reviews, and playlist titles, and models that analyze the audio itself—capabilities that Spotify reinforced by acquiring The Echo Nest in 2014.

It was the packaging that made it land. 30 songs, not an endless feed. Monday morning, same time, every week. A tone that said, You were born for this. The product was the batch schedule that sounds like a technical limitation: Spotify engineer Edward Newett has described refreshing 100 million playlists every Sunday night on about a terabyte of new data. Spotify’s reported results were not subtle—40 million unique users and 5 billion tracks streamed in the first ten months, and more than 8,000 artists got over half their monthly listeners from the playlist.

So the useful question for your team is not "How do we build Spotifys model?" It is which recommender approach costs at your data volume and what ritual you wrap around it.

Collaborative filtering, content embeddings, a hybrid with exploration, or a managed service — which recommender should you build first?

Collaborative filteringContent embeddingsHybrid + explorationManaged service
Cost to first versionLowest — days to a few weeks with an open-source ALS implementation and an ANN indexModerate — embedding generation plus a vector store; higher if you process audio or videoHighest — two candidate sources, a ranker, and an experimentation layer to tuneLow upfront, rising with usage and catalogue size
Cold startPoor — a new item with no interactions is invisible until someone finds itStrong — a brand-new item has an embedding the moment it existsStrong — content covers what behaviour cannotVaries by vendor; usually needs your metadata to be clean
Data you must already haveMeaningful implicit signals: plays, saves, repeat views, purchasesRich item content: text, audio, images, structured attributesBoth, plus enough traffic to run experimentsWhatever the vendor's schema expects, in their format
Ops burdenLow — a scheduled batch job and an index rebuildModerate — embedding pipelines and a vector index to keep freshHigh — multiple models, feature freshness, online serving, monitoringLowest day to day; you inherit their incidents instead
Lock-inNone — the artefacts are your vectorsLow, unless you depend on one embedding providerLow, but high internal complexity to unwindHigh — your ranking quality and roadmap sit inside someone else's product
Explainability and controlGood — 'people who played this also played that' is defensible to users and to legalGood — similarity reasons are inspectableHardest to explain, easiest to tune once instrumentedWeakest — limited insight into why an item ranked
Fails whenThe catalogue turns over fast or the long tail never gets seenBehaviour, not content, drives taste — lookalike items that nobody actually wants togetherYou do not yet have the traffic to tell two rankers apartPersonalisation is your core product rather than a feature

When to pick each option

If you have interaction data and a catalog that is not changing on a daily basis, start with collaborative filtering. Using an implicit-feedback matrix factorization on a plays-or-views matrix, plus an approximate nearest-neighbour index for retrieval, this system is a genuine weekend-to-fortnight build using mature open-source libraries. And it is the approach whose output you can describe to a skeptical stakeholder in a sentence. This is where the heart of Discover Weekly’s signal lives—behavior, at scale.

If your real problem is cold start, choose content embeddings: a marketplace where sellers add items hourly, a news product where yesterday’s article is worthless, or a video library where new uploads need to be discoverable immediately. Text and audio models put it in the space before anyone touches a new item. Which is why Spotify marries behavior with text and audio analysis—a track nobody has streamed yet still has to be recommendable.

If you have both signals and sufficient traffic to measure, go hybrid with exploration. The shape is candidate generation from multiple sources, one ranker that scores them together, and deliberate exploration so the system doesn’t collapse into the same twenty items. Don't begin here. A ranker with nothing to rank is a month's work that improves nothing.

Use a managed service when recommendations are a secondary feature and you want to buy the pipeline—an example is an e-commerce “you may also like” row. Trade-off: you get speed now, but control later is limited. If the product you are selling is personalization then the extra weeks to own the vectors are worth itpersonalization,.

Recommended architecture for a first personalised feed

Rendering diagram…

# pip install implicit annoy scipy
# Inputs: one row per interaction (user_id, item_id, weight).
# Weight is confidence, not rating: a save counts more than a view.

import numpy as np, scipy.sparse as sp
from implicit.als import AlternatingLeastSquares
from annoy import AnnoyIndex

FACTORS = 64

def build(interactions, n_users, n_items):
    rows = [u for u, _, _ in interactions]
    cols = [i for _, i, _ in interactions]
    vals = [w for _, _, w in interactions]
    matrix = sp.csr_matrix((vals, (rows, cols)), shape=(n_users, n_items))

    model = AlternatingLeastSquares(
        factors=FACTORS, regularization=0.05, iterations=20
    )
    model.fit(matrix)                      # user_factors, item_factors

    index = AnnoyIndex(FACTORS, 'dot')     # dot product matches the ALS objective
    for item_id, vec in enumerate(model.item_factors):
        index.add_item(item_id, vec)
    index.build(20)
    return model, matrix, index

def weekly_list(user_id, model, matrix, index, seen, k=30, explore=3):
    # Over-fetch, then filter: the recommendation you must not show is
    # the one the user already consumed.
    raw = index.get_nns_by_vector(model.user_factors[user_id], k * 5)
    picks = [i for i in raw if i not in seen][: k - explore]

    # Reserve a few slots for deliberate novelty, or the list converges
    # on the same catalogue corner every week.
    tail = [i for i in raw[k * 2 :] if i not in seen and i not in picks]
    picks += list(np.random.choice(tail, size=min(explore, len(tail)), replace=False))
    return picks

# Run it on a schedule, write the result per user, and serve the stored list.
# Precomputing is not a compromise here — it is what makes a fixed weekly slot possible.
We had to refresh 100 million playlists every Sunday night— Edward Newett

MVP development planner

Scope, budget, and launch plan

Choose the product, platforms, and modules. The range, phases, and team update as you go. It is a planning figure, not a quote.

Product
Platforms
Modules
Design
Pace
Launch scale
Compliance

Newsletter

One architecture teardown or evaluation, every other week

Practical write-ups on framework choices, migrations, and scaling decisions – written for people who have to justify the call in a meeting.

No spam. Unsubscribe anytime.

Next step

Book a 15-min Architecture Review

Walk through your system with a lead architect and leave with a scoped recommendation.

Schedule on Cal.com
Book a call WhatsApp