z03: Channel-Agnostic YouTube Performance Predictors

z03 is a suite of two isolated, standalone models (one for vision, one for text) designed to analyze YouTube thumbnails and titles independently, predicting whether a video will be a high-performing outlier relative to its specific channel.

Why?

Most view-prediction models fail (like TubeCLIP) because they are blind to channel size—a 100k view video is a massive hit for a small channel, but a flop for a massive one.

z03 solves this by removing popularity bias. Instead of predicting absolute view tiers, it acts as a binary classifier for outliers. By averaging the views of a channel's past 10 videos, these models independently evaluate the thumbnail or title to predict if the new video will achieve at least 2.5x the channel's baseline average.

Prediction Demo Model Logo

Try the Live Demo Here!

Demo

Note: this is a pre-predicted demo, hosting a ml model is expensive.


Getting It Running!

Prerequisites

Before installing and running the ONNX models, ensure your environment meets the following requirements:

  • Python 3.x
  • onnxruntime
  • pillow
  • numpy
  • transformers
  • huggingface_hub

Installation Guide

You can install the required dependencies using pip:

pip install onnxruntime pillow numpy transformers huggingface_hub

Usage & API Examples

To begin, download the ONNX model files programmatically using the Hugging Face Hub library.

Note: kindly stay away from cloning git it'll download hash instead of actual model weights.

from huggingface_hub import snapshot_download

REPO = "Krudev/z03"
snapshot_download(repo_id=REPO, local_dir="./z03") #i hate hidden files

Once downloaded, you can import the utility functions to run predictions within your Python environment.

Thumbnail Model Execution

from inference.utils import predict_thumbnail_onnx

thumbnail = "./TBxS0XhdfmU-HD.jpg"
model_path = "./z03/onnx/thumbnail_predictor.onnx"
prob, pred = predict_thumbnail_onnx(model_path, thumbnail, image_size=299)
print(prob, pred)

Title Model Execution

from inference.utils import predict_title_onnx

title = "hello"
model_path = "./z03/onnx/title_predictor.onnx"
prob, pred = predict_title_onnx(model_path, title)
print(prob, pred)

Command Line Interface (CLI)

If you want easiest inference! this is for you ;) (CLI!!)

For streamlined operations, you can pass arguments directly via the terminal using easy_inference.py.

Thumbnail Only

python ./z03/inference/easy_inference.py --image_path ./TBxS0XhdfmU-HD.jpg 
--- Thumbnail Inference ---
File: ./TBxS0XhdfmU-HD.jpg
Probability: 0.8742
Prediction: 1

Title Only

python ./z03/inference/easy_inference.py --title "SHOCKING Secret About My New Setup!"
--- Title Inference ---
Text: 'SHOCKING Secret About My New Setup!'
Probability: 0.9315
Prediction: 1

Both Simultaneously

python ./z03/inference/easy_inference.py --image_path ./TBxS0XhdfmU-HD.jpg --title "SHOCKING Secret About My New Setup!"
--- Thumbnail Inference ---
File: ./TBxS0XhdfmU-HD.jpg
Probability: 0.8742
Prediction: 1

--- Title Inference ---
Text: 'SHOCKING Secret About My New Setup!'
Probability: 0.9315
Prediction: 1

AI & Technical Specifics

Architecture Strategy After experimenting with heavier multimodal models (like Donut and Pix2struct) which suffered from high inference costs and early plateauing, this project pivoted to a decoupled, lightweight approach for simplicity and speed. Rather than a complex joint-embedding pipeline, z03 consists of two completely isolated models that run predictions individually:

  • Vision-Only Predictor: Fine-tuned Inception-v3
  • Text-Only Predictor: Fine-tuned all-MiniLM-L6-v2

Performance Metrics & Results

Recent retraining iterations, specifically filtering out dead/corrupted data samples, have yielded the following independent baseline scores:

  • z03 Vision (Inception-v3): 0.621 AUC (Improved after removing ~1,900 broken/black-square images).
  • z03 Text (MiniLM): 0.640 AUC / 0.423 PR-AUC (Achieved at Epoch 1 before declining).

Limitations & Biases

This model relies on historical channel data to establish its baseline. If a channel's past 10 videos are highly erratic, the 2.5x outlier threshold becomes difficult to define. Furthermore, the model is trained explicitly on predictable niches; it actively excludes highly volatile categories like breaking news, trending topics, and music where virality is often detached from thumbnail/title quality.

Datasets Used

The model is trained on a highly curated dataset of 16.28k samples collected using official YouTube Data Api V3 and rigorously filtered from 250k unique samples, finalized after cleaning out approximately 1,900 missing or corrupted images.

To hint to the model that the YouTube algorithm is naturally biased toward standard performance over viral hits, the dataset intentionally maintains a 1:3 ratio (outlier : normal). To prevent the model from simply predicting "normal" every time due to this skew, higher class weights were applied to the positive (outlier) samples during training to balance the learning process. Additionally, to prevent any single creator from skewing the model's understanding of an "outlier," data scraping was strictly limited to a maximum of 5 samples per channel.

To comply with YouTube's API policies the dataset will not be published and kept more than 30 days.


MORE

License

This project is open-source but restricted for commercial use. It is licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You are free to use, modify, and build upon this tool for personal and research purposes.

Support & Contact

If you encounter bugs, have questions, or want to discuss collaboration, feel free to reach out:

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support