Home

ML · Analytics

Netflix Recommender: Keywords vs Embeddings

GitHub ↗

First built in 2023, revisited in 2026 with a newer method.

The idea

A recommender with no viewing history to learn from has only the catalogue itself: titles, cast, genres and one-line descriptions. I built one on about 7,800 Netflix films and shows, then came back three years later and rebuilt it a different way. Putting the two side by side turned out to be more useful than either on its own.

Two versions

2023

Keyword matching

scikit-learnword countscosine similarity

Each title becomes a bag of words built from its cast, country, rating, genres and description. Two titles are similar when they share words. I also clustered the catalogue with K-Means and hierarchical clustering to explore how content groups.

2026

Sentence embeddings

all-MiniLM-L6-v2384 dimensionsFAISS

A pretrained language model turns each title's text into a vector that captures meaning, not exact wording. A FAISS index then finds the nearest vectors, so two titles can match without sharing a single word.

Same titles, two answers

  • If you liked The Crown

    Keywords · 2023

    1. 1Hinterland
    2. 2Kiss Me First
    3. 3Collateral
    4. 4Flowers
    5. 5Dracula

    Embeddings · 2026

    1. 1Reign
    2. 2The Windsors
    3. 3The Royal House of Windsor
    4. 4The English Game
    5. 5W1A

    Embeddings win. Keyword matching found British dramas; the embedding model understood the show is about royalty.

  • If you liked Breaking Bad

    Keywords · 2023

    1. 1Better Call Saul
    2. 2The School Nurse Files
    3. 3Re:Mind
    4. 4Hormones
    5. 5Marvel's The Punisher

    Embeddings · 2026

    1. 1The Break
    2. 2Reckoning
    3. 3The Show
    4. 4Medical Police
    5. 5Somewhere Between

    Keywords win. Shared cast puts Better Call Saul first. The embedding model latched onto the word “Break” in the title.

What I learned

  • Newer is not automatically better. Each method has a typical failure. Keywords miss a theme when the wording differs; embeddings can be pulled off course by one prominent word.
  • What goes into the text matters as much as the model. I included the title in the text the model reads, which is why it matched on title words. Leaving the title out, or combining both methods, is the next thing to try.
  • Reading lists is not evaluation. Without viewing data there is no score to say which version is better overall, only examples. A real test would need to know what people went on to watch.