HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
Paper • 2503.08585 • Published
Checkpoint for HierarQ from the paper HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding (CVPR, 2025).
hierarq_msvd_cap.pth: This is an intermediate checkpoint of the model trained on the MSVD dataset on the captioning taskDownload the checkpoint and follow the instructions in the code repository:
hf download shehreen3764/HierarQ_Checkpoint --local-dir ./checkpoints
@inproceedings{azad2025hierarq,
title = {HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding},
author = {Azad, Shehreen and Vineet, Vibhav and Rawat, Yogesh Singh},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2025}
}