CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval

Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang

🤗 Model | 🤗 Data ｜ 📑 Paper

📝 Introduction

This is CaRe trained after Stage-I. It can only handle video captioning tasks. Refer to our paper for details.

Usage

Loading from the huggingface remote path is not tested. It is recommended to download this checkpoint to your local environment to prevent potential bugs.

For Captioning Tasks

from utils.video import read_frames_decord
from models.modeling_captioners import AutoCaptioner

captioner = AutoCaptioner.from_pretrained('path/to/checkpoints/CaRe-7B-Stage-1')
frames = read_frames_decord(video_path='assets/demo.mp4', num_frames=32)
description = captioner.describe(frames.unsqueeze(0))
print(description[0])

MCG-NJU
/

CaRe-7B-Stage-1

CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval

📝 Introduction

Usage

For Captioning Tasks

Collection including MCG-NJU/CaRe-7B-Stage-1

CaRe