ML · Analytics
Netflix Recommender: Keywords vs Embeddings
First built in 2023, revisited in 2026 with a newer method.
The idea
A recommender with no viewing history to learn from has only the catalogue itself: titles, cast, genres and one-line descriptions. I built one on about 7,800 Netflix films and shows, then came back three years later and rebuilt it a different way. Putting the two side by side turned out to be more useful than either on its own.
Two versions
2023
Keyword matching
scikit-learnword countscosine similarity
Each title becomes a bag of words built from its cast, country, rating, genres and description. Two titles are similar when they share words. I also clustered the catalogue with K-Means and hierarchical clustering to explore how content groups.
2026
Sentence embeddings
all-MiniLM-L6-v2384 dimensionsFAISS
A pretrained language model turns each title's text into a vector that captures meaning, not exact wording. A FAISS index then finds the nearest vectors, so two titles can match without sharing a single word.
Same titles, two answers
If you liked The Crown
Keywords · 2023
- 1Hinterland
- 2Kiss Me First
- 3Collateral
- 4Flowers
- 5Dracula
Embeddings · 2026
- 1Reign
- 2The Windsors
- 3The Royal House of Windsor
- 4The English Game
- 5W1A
Embeddings win. Keyword matching found British dramas; the embedding model understood the show is about royalty.
If you liked Breaking Bad
Keywords · 2023
- 1Better Call Saul
- 2The School Nurse Files
- 3Re:Mind
- 4Hormones
- 5Marvel's The Punisher
Embeddings · 2026
- 1The Break
- 2Reckoning
- 3The Show
- 4Medical Police
- 5Somewhere Between
Keywords win. Shared cast puts Better Call Saul first. The embedding model latched onto the word “Break” in the title.
What I learned
- Newer is not automatically better. Each method has a typical failure. Keywords miss a theme when the wording differs; embeddings can be pulled off course by one prominent word.
- What goes into the text matters as much as the model. I included the title in the text the model reads, which is why it matched on title words. Leaving the title out, or combining both methods, is the next thing to try.
- Reading lists is not evaluation. Without viewing data there is no score to say which version is better overall, only examples. A real test would need to know what people went on to watch.
- Python
- scikit-learn
- Sentence Transformers
- FAISS
- K-Means
- Hierarchical clustering
- Streamlit