TubeCLIP: AI-Driven YouTube Performance Predictor
TubeCLIP is a fine-tuned variant of CLIP Large designed to analyze YouTube video thumbnails and titles to predict view performance classes.
Why?
Content creators and marketers constantly rely on guesswork, gut feelings, and tedious A/B testing to figure out which thumbnail and title combination will drive attention. TubeCLIP solves this by replacing intuition with data-driven prediction. By fine-tuning a CLIP model, it evaluates the complex relationship between a thumbnail's visual elements and its title to predict its potential view tier.
Prediction Demo
(Coming Soon)
Getting it Running
Prerequisites
Before installing, ensure your environment meets the following requirements:
- Python 3.x
pytorch-lightningtorchao==0.16.0
Installation Guide
You can install the required dependencies using pip:
pip install torch pytorch-lightning torchao==0.16.0 huggingface_hub transformers Pillow
Usage & API Examples
Here is a quick example of how to load the model weights from Hugging Face and run a prediction on your thumbnail and title.
import torch
from huggingface_hub import snapshot_download
from PIL import Image
# 1. Download The Repository
model_path = snapshot_download(repo_id="Krudev/TubeCLIP", local_dir="/TubeCLIP"
)
Now run the cli with predict.py
python TubeCLIP.predict.py --model_path "/TubeCLIP/TubeCLIP.ckpt" --input_path "path/to/your/thumbnail/or/directory/of/thumbnails" --title "Your YouTube Video Title!"
AI & Technical Specifics
Performance Metrics & Results
TubeCLIP achieves 66% accuracy in classifying video performance across three distinct view tiers:
- Tier 1: 10k - 100k views
- Tier 2: 100k - 1M views
- Tier 3: 1M+ views
Evaluation Results
- Test Loss: 1.8045
- Test Accuracy: 0.6747 (67.47%)
Classification Report
| Class | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| 10k-100k | 0.70 | 0.72 | 0.71 | 1,606 |
| 100k-1M | 0.61 | 0.59 | 0.60 | 1,608 |
| 1M+ | 0.72 | 0.71 | 0.71 | 1,511 |
| Accuracy | 0.67 | 4,725 | ||
| Macro Avg | 0.67 | 0.68 | 0.68 | 4,725 |
| Weighted Avg | 0.67 | 0.67 | 0.67 | 4,725 |
Limitations & Biases
While the model is highly effective at distinguishing generally "good" (high potential) versus "bad" (low potential) thumbnail/title combinations, it is not an exact view-count calculator. Viewership relies on external factors (channel size, algorithmic luck, trending topics, time of day) that the model cannot see. Expect it to serve as a strong directional compass for A/B testing rather than a perfect view predictor.
Datasets Used
The model was trained on a custom-built, highly filtered, and strictly balanced dataset of 30,000 YouTube videos. This curated dataset ensures the model learns pure visual-textual relationships without being overwhelmed by garbage data.
MORE
License
This project is open-source but restricted for commercial use. It is licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You are free to use, modify, and build upon this tool for personal and research purposes, but you may not use it for commercial gains without permission.
Support & Contact
If you encounter bugs, have questions, or want to discuss collaboration, feel free to reach out:
- Email: krishnenduk462@gmail.com
- X (Twitter): @krishnendw
- Instagram: @contentbykrishnendu