Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

CineMatch — Movie Recommendation System

A movie recommendation system built with PySpark ALS collaborative filtering for personalised user recommendations and FAISS-powered item similarity search for "find movies like this" queries. Served through a clean Streamlit web app.

Live demo: Deployed on Streamlit Community Cloud


Features

  • Collaborative Filtering (ALS) — Predicts top-N movies for each user based on learned latent taste factors
  • Item–Item Similarity — Finds movies similar to a given title using ALS item vectors + FAISS cosine search
  • Model Metrics dashboard — RMSE, MAE, catalogue coverage, and hyperparameter summary
  • SQLite backend — Recommendations and model metrics persisted in a local database (not a flat CSV)
  • EDA notebook — Data exploration covering rating distributions, user activity, sparsity, and genre landscape

Project Structure

├── movie_app.py              # Streamlit app (3 tabs: User, Movie, Metrics)
├── Training_ALS_model.py     # Train ALS; writes to SQLite + saves item_factors.pkl
├── faiss_index.py            # Build FAISS similarity index from item_factors.pkl
├── inference.py              # Item–item similarity search (FAISS cosine lookup)
├── notebooks/
│   └── EDA.ipynb             # Exploratory data analysis
├── data/
│   ├── movies.csv            # movieId, title, genres
│   └── ratings.csv           # userId, movieId, rating[, timestamp]
├── artifacts/
│   └── recommendations.db    # SQLite: movies + recommendations + model_metrics
├── item_factors.pkl          # ALS item latent vectors (dict: movieId → np.ndarray)
├── item_factors.faiss        # Pre-built FAISS index
├── requirements.txt
└── Dockerfile

Recommendation Approaches

1. Collaborative Filtering — ALS

ALS (Alternating Least Squares) factorises the user–movie rating matrix into two lower-dimensional matrices:

  • User factors — each user's taste profile in latent space
  • Item factors — each movie's profile in the same space

The predicted rating for a (user, movie) pair is the dot product of their respective latent vectors. Top-N predictions per user are stored in SQLite for fast retrieval.

Hyperparameters used:

Parameter Value Rationale
rank 50 Richer latent space — rank=8 produced near-identical vectors for obscure movies, causing nonsensical similarity results
maxIter 15 Additional iterations for better convergence at higher rank
regParam 0.1 Moderate regularisation to prevent power-user overfitting
nonnegative True Ratings are non-negative; constrain vectors accordingly
coldStartStrategy drop Exclude unseen users/movies from RMSE to avoid NaN pollution
min_movie_ratings 20 Pre-filter: movies with <20 ratings are removed before training; cold-start items produce degenerate vectors
min_user_ratings 5 Pre-filter: users with <5 ratings contribute too little signal

2. Item–Item Similarity — ALS Vectors + FAISS

Note: this is NOT genre-based cosine similarity. It uses the item latent vectors learned by ALS. Movies with similar audience taste profiles (regardless of genre) end up close in latent space.

The vectors are L2-normalised and indexed with faiss.IndexFlatIP — inner product on normalised vectors equals cosine similarity.


Evaluation Metrics

Metric Description
RMSE Root Mean Squared Error on held-out 20% test set. Below 1.0 is generally good for 1–5 star ratings.
MAE Mean Absolute Error — more interpretable; average absolute deviation between predicted and actual ratings.
Coverage % of catalogue movies that appear in at least one recommendation list. Low = cold-start problem.

Running the Project

Prerequisites

  • Python 3.11+
  • Java 8+ (required by PySpark)

Install dependencies

pip install -r requirements.txt

Step 1 — Train the model

python Training_ALS_model.py

This reads data/ratings.csv and data/movies.csv, trains ALS, evaluates on a held-out test set, and writes:

  • artifacts/recommendations.db — SQLite with all recommendations + metrics
  • item_factors.pkl — ALS item latent vectors

Step 2 — Build the FAISS index

python faiss_index.py

Reads item_factors.pkl and writes item_factors.faiss.

Step 3 — Run the EDA notebook (optional)

jupyter notebook notebooks/EDA.ipynb

Step 4 — Launch the app

streamlit run movie_app.py

App runs at http://localhost:8501.

Docker

docker build -t cinematch .
docker run -p 8501:8501 cinematch

Tech Stack

Layer Technology
Distributed ML PySpark 3.5 / Spark MLlib ALS
Vector search FAISS (faiss-cpu)
Storage SQLite
Web UI Streamlit
Data Pandas, NumPy
EDA Matplotlib, Seaborn, Jupyter

What I Learned / What I'd Do Differently

  • FAISS for similarity, not genre strings — Latent vectors capture taste patterns that pure genre labels miss entirely. Two action movies can be very dissimilar in latent space if their audiences don't overlap.
  • SQLite over CSV for recommendations — A pre-generated CSV is brittle and slow to query. SQLite adds indexed lookups, pagination, and the ability to store metadata (metrics, timestamps) alongside recommendations with zero infrastructure overhead.
  • Coverage matters as much as RMSE — A model with great RMSE on popular movies might have poor coverage of the long tail. Both should be reported.
  • If starting again: explore implicit library as a lighter ALS alternative that doesn't require a JVM, making Docker deployment much simpler.

Dataset

Uses the MovieLens dataset (GroupLens Research). The small variant (100K ratings) works for local development; the full dataset (25M ratings) requires the Spark memory settings in Training_ALS_model.py.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages