YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- π¬ Movie Recommendation System
- π How It Works
- π§ Machine Learning Approach
- π οΈ Technologies Used
- π Project Structure
- π Dataset
- π€ Pre-trained Model Files
- βοΈ Installation
- π₯ Using the Pre-trained Model
- βΆοΈ Run the Recommendation System
- π Example
- π Similarity Technique
- π§© Recommendation Process
- π‘ Why Content-Based Filtering?
- β οΈ Important Notes
- π Training / Development
- π¨βπ» Author
- β If You Find This Project Useful
- π How It Works
π¬ Movie Recommendation System
A content-based Movie Recommendation System built using Machine Learning and Natural Language Processing (NLP).
The system recommends movies that are similar to a movie selected by the user. It uses movie information such as genres, keywords, cast, crew, and overview to determine the similarity between movies.
π How It Works
The recommendation system follows these main steps:
Movie Dataset
β
Data Preprocessing
β
Feature Selection
β
Feature Combination
β
Text Processing
β
Vectorization
β
Cosine Similarity
β
Movie Recommendations
The movie information is converted into numerical vectors using text vectorization. The system then calculates the similarity between movies using Cosine Similarity.
When a user enters a movie name, the system finds movies with similar feature vectors and returns the most similar movies.
π§ Machine Learning Approach
This project uses a Content-Based Filtering approach.
Instead of using ratings from multiple users, the system recommends movies based on the characteristics of the selected movie.
For example:
User selects:
The Dark Knight
β
System finds movies with similar features
β
Recommendations:
Batman Begins
The Dark Knight Rises
...
π οΈ Technologies Used
- Python
- Pandas
- NumPy
- Scikit-learn
- Natural Language Processing (NLP)
- CountVectorizer
- Cosine Similarity
- Joblib
- Jupyter Notebook
π Project Structure
The project contains the source code, dataset, trained model files, and the recommendation script.
Movie-Recommendation-System/
β
βββ model.ipynb
βββ recommend.py
βββ dataset/
βββ tmdb_5000_credits.csv
βββtmdb_5000_movies.csv
βββ vectors.pkl
βββ DF.pkl
βββ requirements.txt
βββ README.md
Files
| File | Description |
|---|---|
model.ipynb |
Notebook used for preprocessing, feature engineering, vectorization, and model creation |
recommend.py |
Python script used to generate movie recommendations |
tmdb_5000_credits.csv |
Movie credits dataset |
tmdb_5000_movies.csv |
Movie information dataset |
vectors.pkl |
Pre-generated movie feature vectors |
DF.pkl |
Pre-generated movie dataframe |
requirements.txt |
Required Python dependencies |
README.md |
Project documentation |
π Dataset
This project uses the TMDB 5000 Movie Dataset, which contains movie information including titles, genres, keywords, cast, crew, and movie overviews.
The dataset is available on Kaggle:
Dataset: https://www.kaggle.com/datasets/tmdb/tmdb-movie-metadata
The dataset contains two files:
tmdb_5000_movies.csvtmdb_5000_credits.csv
These datasets are used for data preprocessing, feature engineering, text processing, and building the movie recommendation system.
π€ Pre-trained Model Files
The pre-generated model files are available here:
Hugging Face Repository:
https://huggingface.co/nuthan-444/Movie_Recommendation_System
The repository contains the files required to run the recommendation system without recreating the model from the notebook.
Model Files
| File | Description |
|---|---|
DF.pkl |
Contains the movie dataset/dataframe |
vectors.pkl |
Contains the vectorized movie features |
requirements.txt |
Required Python dependencies |
recommend.py |
Python script for generating recommendations |
βοΈ Installation
1. Download the Project
Clone the repository:
git clone https://github.com/nuthan-444/Movie_Recommendation_System.git
Move into the project directory:
cd Movie_Recommendation_System
2. Create a Virtual Environment
python -m venv venv
3. Activate the Virtual Environment
Windows
venv\Scripts\activate
macOS / Linux
source venv/bin/activate
4. Install Dependencies
pip install -r requirements.txt
π₯ Using the Pre-trained Model
If you want to use the already-generated model files without running the complete notebook, follow these steps.
Step 1 β Download the Model Files
Download the following files:
DF.pkl
vectors.pkl
from:
https://huggingface.co/nuthan-444/Movie_Recommendation_System
Place both files in the same directory as your Python script.
Your folder should look like:
movie-recommendation/
β
βββ DF.pkl
βββ vectors.pkl
βββ requirements.txt
βββ recommend.py
Step 2 β Install Dependencies
Make sure Python is installed, then run:
pip install -r requirements.txt
Step 3 β Create recommend.py
Create a Python file named:
recommend.py
Add the following code:
import joblib
from sklearn.metrics.pairwise import cosine_similarity
DF = joblib.load('DF.pkl')
vectors = joblib.load('vectors.pkl')
similarity = cosine_similarity(vectors)
def recommend(movie):
movie_index = DF[DF['title'] == movie].index[0]
distances = similarity[movie_index]
movie_list = sorted(
list(enumerate(distances)),
reverse=True,
key=lambda x: x[1]
)[1:6]
for i in movie_list:
print(DF.iloc[i[0]].title)
recommend('Avatar')
βΆοΈ Run the Recommendation System
Run the Python script:
python recommend.py
For example:
recommend('Avatar')
The system will print the 5 most similar movies to Avatar.
You can change the movie name:
recommend('Titanic')
or:
recommend('The Dark Knight')
Make sure the movie title exactly matches a title available in the
titlecolumn ofDF.pkl.
π Example
Input
The Dark Knight
Output
Batman Begins
The Dark Knight Rises
Batman
...
The recommendations are generated based on the similarity between the selected movie and the other movies in the dataset.
π Similarity Technique
The project uses Cosine Similarity to measure how similar two movie vectors are.
Cosine Similarity measures the angle between two vectors.
A higher cosine similarity indicates that two movies have more similar features.
The similarity matrix is generated using:
similarity = cosine_similarity(vectors)
π§© Recommendation Process
When a movie is provided:
Movie Name
β
Find Movie in Dataset
β
Get Its Vector
β
Compare With All Movie Vectors
β
Calculate Cosine Similarity
β
Sort Similarity Scores
β
Select Top 5
β
Display Recommendations
π‘ Why Content-Based Filtering?
Content-based filtering is useful when recommendations need to be based on the actual characteristics of an item.
In this project, movie metadata is used instead of relying on user ratings or user-to-user behavior.
This allows the system to recommend movies based on the characteristics of the selected movie.
β οΈ Important Notes
- Keep
DF.pkl,vectors.pkl, andrecommend.pyin the same directory when using the pre-trained model files. - Do not rename
DF.pklorvectors.pklunless you also update the filenames in the Python code. - The movie name must exactly match a title available in the
titlecolumn ofDF.pkl. - The recommendation system returns the top 5 similar movies.
- The
.pklfiles are pre-generated and do not need to be recreated to use the recommendation script. - The complete model-building process can be reproduced using
model.ipynbto download goto below link. - https://github.com/nuthan-444/Movie_Recommendation_System.
π Training / Development
The complete model-building process is available in:
model.ipynb
The notebook covers:
- Data preprocessing
- Feature selection
- Feature engineering
- Feature combination
- Text processing
- Vectorization
- Cosine similarity
- Saving the generated model files
π¨βπ» Author
Nuthan Prasad K G
GitHub:
β If You Find This Project Useful
Feel free to explore the code, experiment with the recommendation system, and improve the project.