A movie recommendation system built with PySpark ALS collaborative filtering for personalised user recommendations and FAISS-powered item similarity search for "find movies like this" queries. Served through a clean Streamlit web app.
Live demo: Deployed on Streamlit Community Cloud
- Collaborative Filtering (ALS) — Predicts top-N movies for each user based on learned latent taste factors
- Item–Item Similarity — Finds movies similar to a given title using ALS item vectors + FAISS cosine search
- Model Metrics dashboard — RMSE, MAE, catalogue coverage, and hyperparameter summary
- SQLite backend — Recommendations and model metrics persisted in a local database (not a flat CSV)
- EDA notebook — Data exploration covering rating distributions, user activity, sparsity, and genre landscape
├── movie_app.py # Streamlit app (3 tabs: User, Movie, Metrics)
├── Training_ALS_model.py # Train ALS; writes to SQLite + saves item_factors.pkl
├── faiss_index.py # Build FAISS similarity index from item_factors.pkl
├── inference.py # Item–item similarity search (FAISS cosine lookup)
├── notebooks/
│ └── EDA.ipynb # Exploratory data analysis
├── data/
│ ├── movies.csv # movieId, title, genres
│ └── ratings.csv # userId, movieId, rating[, timestamp]
├── artifacts/
│ └── recommendations.db # SQLite: movies + recommendations + model_metrics
├── item_factors.pkl # ALS item latent vectors (dict: movieId → np.ndarray)
├── item_factors.faiss # Pre-built FAISS index
├── requirements.txt
└── Dockerfile
ALS (Alternating Least Squares) factorises the user–movie rating matrix into two lower-dimensional matrices:
- User factors — each user's taste profile in latent space
- Item factors — each movie's profile in the same space
The predicted rating for a (user, movie) pair is the dot product of their respective latent vectors. Top-N predictions per user are stored in SQLite for fast retrieval.
Hyperparameters used:
| Parameter | Value | Rationale |
|---|---|---|
rank |
50 | Richer latent space — rank=8 produced near-identical vectors for obscure movies, causing nonsensical similarity results |
maxIter |
15 | Additional iterations for better convergence at higher rank |
regParam |
0.1 | Moderate regularisation to prevent power-user overfitting |
nonnegative |
True | Ratings are non-negative; constrain vectors accordingly |
coldStartStrategy |
drop | Exclude unseen users/movies from RMSE to avoid NaN pollution |
min_movie_ratings |
20 | Pre-filter: movies with <20 ratings are removed before training; cold-start items produce degenerate vectors |
min_user_ratings |
5 | Pre-filter: users with <5 ratings contribute too little signal |
Note: this is NOT genre-based cosine similarity. It uses the item latent vectors learned by ALS. Movies with similar audience taste profiles (regardless of genre) end up close in latent space.
The vectors are L2-normalised and indexed with faiss.IndexFlatIP — inner product on normalised vectors equals cosine similarity.
| Metric | Description |
|---|---|
| RMSE | Root Mean Squared Error on held-out 20% test set. Below 1.0 is generally good for 1–5 star ratings. |
| MAE | Mean Absolute Error — more interpretable; average absolute deviation between predicted and actual ratings. |
| Coverage | % of catalogue movies that appear in at least one recommendation list. Low = cold-start problem. |
- Python 3.11+
- Java 8+ (required by PySpark)
pip install -r requirements.txtpython Training_ALS_model.pyThis reads data/ratings.csv and data/movies.csv, trains ALS, evaluates on a held-out test set, and writes:
artifacts/recommendations.db— SQLite with all recommendations + metricsitem_factors.pkl— ALS item latent vectors
python faiss_index.pyReads item_factors.pkl and writes item_factors.faiss.
jupyter notebook notebooks/EDA.ipynbstreamlit run movie_app.pyApp runs at http://localhost:8501.
docker build -t cinematch .
docker run -p 8501:8501 cinematch| Layer | Technology |
|---|---|
| Distributed ML | PySpark 3.5 / Spark MLlib ALS |
| Vector search | FAISS (faiss-cpu) |
| Storage | SQLite |
| Web UI | Streamlit |
| Data | Pandas, NumPy |
| EDA | Matplotlib, Seaborn, Jupyter |
- FAISS for similarity, not genre strings — Latent vectors capture taste patterns that pure genre labels miss entirely. Two action movies can be very dissimilar in latent space if their audiences don't overlap.
- SQLite over CSV for recommendations — A pre-generated CSV is brittle and slow to query. SQLite adds indexed lookups, pagination, and the ability to store metadata (metrics, timestamps) alongside recommendations with zero infrastructure overhead.
- Coverage matters as much as RMSE — A model with great RMSE on popular movies might have poor coverage of the long tail. Both should be reported.
- If starting again: explore
implicitlibrary as a lighter ALS alternative that doesn't require a JVM, making Docker deployment much simpler.
Uses the MovieLens dataset (GroupLens Research). The small variant (100K ratings) works for local development; the full dataset (25M ratings) requires the Spark memory settings in Training_ALS_model.py.