z03: Channel-Agnostic YouTube Performance Predictors
z03 is a suite of two isolated, standalone models (one for vision, one for text) designed to analyze YouTube thumbnails and titles independently, predicting whether a video will be a high-performing outlier relative to its specific channel.
Why?
Most view-prediction models fail (like TubeCLIP) because they are blind to channel size—a 100k view video is a massive hit for a small channel, but a flop for a massive one.
z03 solves this by removing popularity bias. Instead of predicting absolute view tiers, it acts as a binary classifier for outliers. By averaging the views of a channel's past 10 videos, these models independently evaluate the thumbnail or title to predict if the new video will achieve at least 2.5x the channel's baseline average.
Prediction Demo
Try the Live Demo Here!
Note: this is a pre-predicted demo, hosting a ml model is expensive.
Getting It Running!
Prerequisites
Before installing and running the ONNX models, ensure your environment meets the following requirements:
- Python 3.x
onnxruntimepillownumpytransformershuggingface_hub
Installation Guide
You can install the required dependencies using pip:
pip install onnxruntime pillow numpy transformers huggingface_hub
Usage & API Examples
To begin, download the ONNX model files programmatically using the Hugging Face Hub library.
Note: kindly stay away from cloning git it'll download hash instead of actual model weights.
from huggingface_hub import snapshot_download
REPO = "Krudev/z03"
snapshot_download(repo_id=REPO, local_dir="./z03") #i hate hidden files
Once downloaded, you can import the utility functions to run predictions within your Python environment.
Thumbnail Model Execution
from inference.utils import predict_thumbnail_onnx
thumbnail = "./TBxS0XhdfmU-HD.jpg"
model_path = "./z03/onnx/thumbnail_predictor.onnx"
prob, pred = predict_thumbnail_onnx(model_path, thumbnail, image_size=299)
print(prob, pred)
Title Model Execution
from inference.utils import predict_title_onnx
title = "hello"
model_path = "./z03/onnx/title_predictor.onnx"
prob, pred = predict_title_onnx(model_path, title)
print(prob, pred)
Command Line Interface (CLI)
If you want easiest inference! this is for you ;) (CLI!!)
For streamlined operations, you can pass arguments directly via the terminal using easy_inference.py.
Thumbnail Only
python ./z03/inference/easy_inference.py --image_path ./TBxS0XhdfmU-HD.jpg
--- Thumbnail Inference ---
File: ./TBxS0XhdfmU-HD.jpg
Probability: 0.8742
Prediction: 1
Title Only
python ./z03/inference/easy_inference.py --title "SHOCKING Secret About My New Setup!"
--- Title Inference ---
Text: 'SHOCKING Secret About My New Setup!'
Probability: 0.9315
Prediction: 1
Both Simultaneously
python ./z03/inference/easy_inference.py --image_path ./TBxS0XhdfmU-HD.jpg --title "SHOCKING Secret About My New Setup!"
--- Thumbnail Inference ---
File: ./TBxS0XhdfmU-HD.jpg
Probability: 0.8742
Prediction: 1
--- Title Inference ---
Text: 'SHOCKING Secret About My New Setup!'
Probability: 0.9315
Prediction: 1
AI & Technical Specifics
Architecture Strategy After experimenting with heavier multimodal models (like Donut and Pix2struct) which suffered from high inference costs and early plateauing, this project pivoted to a decoupled, lightweight approach for simplicity and speed. Rather than a complex joint-embedding pipeline, z03 consists of two completely isolated models that run predictions individually:
- Vision-Only Predictor: Fine-tuned
Inception-v3 - Text-Only Predictor: Fine-tuned
all-MiniLM-L6-v2
Performance Metrics & Results
Recent retraining iterations, specifically filtering out dead/corrupted data samples, have yielded the following independent baseline scores:
- z03 Vision (Inception-v3): 0.621 AUC (Improved after removing ~1,900 broken/black-square images).
- z03 Text (MiniLM): 0.640 AUC / 0.423 PR-AUC (Achieved at Epoch 1 before declining).
Limitations & Biases
This model relies on historical channel data to establish its baseline. If a channel's past 10 videos are highly erratic, the 2.5x outlier threshold becomes difficult to define. Furthermore, the model is trained explicitly on predictable niches; it actively excludes highly volatile categories like breaking news, trending topics, and music where virality is often detached from thumbnail/title quality.
Datasets Used
The model is trained on a highly curated dataset of 16.28k samples collected using official YouTube Data Api V3 and rigorously filtered from 250k unique samples, finalized after cleaning out approximately 1,900 missing or corrupted images.
To hint to the model that the YouTube algorithm is naturally biased toward standard performance over viral hits, the dataset intentionally maintains a 1:3 ratio (outlier : normal). To prevent the model from simply predicting "normal" every time due to this skew, higher class weights were applied to the positive (outlier) samples during training to balance the learning process. Additionally, to prevent any single creator from skewing the model's understanding of an "outlier," data scraping was strictly limited to a maximum of 5 samples per channel.
To comply with YouTube's API policies the dataset will not be published and kept more than 30 days.
MORE
License
This project is open-source but restricted for commercial use. It is licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You are free to use, modify, and build upon this tool for personal and research purposes.
Support & Contact
If you encounter bugs, have questions, or want to discuss collaboration, feel free to reach out:
- Email: krishnenduk462@gmail.com
- X (Twitter): @krishnendw
- Instagram: @contentbykrishnendu