TubeCLIP: AI-Driven YouTube Performance Predictor

TubeCLIP is a fine-tuned variant of CLIP Large designed to analyze YouTube video thumbnails and titles to predict view performance classes.

Why?

Content creators and marketers constantly rely on guesswork, gut feelings, and tedious A/B testing to figure out which thumbnail and title combination will drive attention. TubeCLIP solves this by replacing intuition with data-driven prediction. By fine-tuning a CLIP model, it evaluates the complex relationship between a thumbnail's visual elements and its title to predict its potential view tier.

Prediction Demo

(Coming Soon)

Getting it Running

Prerequisites

Before installing, ensure your environment meets the following requirements:

  • Python 3.x
  • pytorch-lightning
  • torchao==0.16.0

Installation Guide

You can install the required dependencies using pip:

pip install torch pytorch-lightning torchao==0.16.0 huggingface_hub transformers Pillow

Usage & API Examples

Here is a quick example of how to load the model weights from Hugging Face and run a prediction on your thumbnail and title.

import torch
from huggingface_hub import snapshot_download
from PIL import Image

# 1. Download The Repository
model_path = snapshot_download(repo_id="Krudev/TubeCLIP", local_dir="/TubeCLIP"
)

Now run the cli with predict.py

python TubeCLIP.predict.py --model_path "/TubeCLIP/TubeCLIP.ckpt" --input_path "path/to/your/thumbnail/or/directory/of/thumbnails" --title "Your YouTube Video Title!"

AI & Technical Specifics

Performance Metrics & Results

TubeCLIP achieves 66% accuracy in classifying video performance across three distinct view tiers:

  • Tier 1: 10k - 100k views
  • Tier 2: 100k - 1M views
  • Tier 3: 1M+ views

Evaluation Results

  • Test Loss: 1.8045
  • Test Accuracy: 0.6747 (67.47%)

Classification Report

Class Precision Recall F1-Score Support
10k-100k 0.70 0.72 0.71 1,606
100k-1M 0.61 0.59 0.60 1,608
1M+ 0.72 0.71 0.71 1,511
Accuracy 0.67 4,725
Macro Avg 0.67 0.68 0.68 4,725
Weighted Avg 0.67 0.67 0.67 4,725
Confusion Matrix

Limitations & Biases

While the model is highly effective at distinguishing generally "good" (high potential) versus "bad" (low potential) thumbnail/title combinations, it is not an exact view-count calculator. Viewership relies on external factors (channel size, algorithmic luck, trending topics, time of day) that the model cannot see. Expect it to serve as a strong directional compass for A/B testing rather than a perfect view predictor.

Datasets Used

The model was trained on a custom-built, highly filtered, and strictly balanced dataset of 30,000 YouTube videos. This curated dataset ensures the model learns pure visual-textual relationships without being overwhelmed by garbage data.


MORE

License

This project is open-source but restricted for commercial use. It is licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You are free to use, modify, and build upon this tool for personal and research purposes, but you may not use it for commercial gains without permission.

Support & Contact

If you encounter bugs, have questions, or want to discuss collaboration, feel free to reach out:

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support