The data preprocessed in this file can be found at this Kaggle repository
Original data when unzipped will give a lyrics.zip file and we will preprocess it to be used for the project in this notebook.
Use Preprocess380kMusicData.pynb file to get the following two csvs
- lyrics_big.csv (266556 rows)
- lyrics_small.csv (26656 rows)
The output files have been uploaded to this link: bit.ly/lm-dir