JVC

Main checkpoint for JVC: Joint Vision Cross-Attention for Fine-Grained Badminton Stroke Recognition (PDF). Navneeth Dhamotharan and Bin Han, ECCV 2026 Workshop: Human Motion-Informed World Models and Socially Intelligent Action, 2026.

The released weight is the joint-vision cross-attention model.

A 16-frame hit clip goes through two encoders: R(2+1)D Conv3D for RGB patches, and a four-stream SkateFormer for MediaPipe joints. Joint tokens cross-attend to visual patches, a divided space–time transformer mixes the clip, and contact-weighted pooling feeds multitask heads. The paper metric is 9-way stroke_type.

Val stroke_type accuracy 80.61% (epoch 18)
Split Video-level 80/20, seed 42. No clip from a validation video is in training.
Dataset FineBadminton-20K
Frames 16, span_linspace over the hit span, 224×224, ImageNet normalization
Skeleton MediaPipe, 33 joints × (x, y, z), four-stream (joint, bone, joint-motion, bone-motion)
Paper JVC: Joint Vision Cross-Attention for Fine-Grained Badminton Stroke Recognition
Demo isocourt.fit

Stroke classes, in logit order: Serve, Clear, Smash, Drop, Drive, Net_Shot, Lob, Defensive_Shot, Other.

The checkpoint also emits logits for technique, placement, position, intent, and quality. stroke_type is the number reported above.

Try the demo

Upload a clip at isocourt.fit.

Weights

from huggingface_hub import hf_hub_download

path = hf_hub_download("navneethdg/JVC", "jvc.pth")

frames is (batch, 16, 3, 224, 224) float RGB after ImageNet normalization (mean = (0.485, 0.456, 0.406), std = (0.229, 0.224, 0.225)). pose is (batch, 16, 33, 3) in the same layout as the training MediaPipe cache. Heads return logits, not probabilities.

Files

File Role
jvc.pth Checkpoint. Constructor metadata sits beside the weights.
config.json Architecture, label names, and metric.

What this weight is

Vision backbone r2plus1d_18 (conv3d), embed dim 128, 2 cross-attention layers, 4 divided space–time blocks, four-stream skeleton. Shuttle features are off.

No-cross-attention JVC ablations are separate runs and are not this file.

Citation

@inproceedings{
dhamotharan2026jvc,
title={{JVC}: Joint Vision Cross-Attention for Fine-Grained Badminton Stroke Recognition},
author={Navneeth Dhamotharan and Bin Han},
booktitle={ECCV26 Workshop: Human Motion-Informed World Models and Socially Intelligent Action},
year={2026},
url={https://openreview.net/forum?id=XJEhcXfwEe}
}
Downloads last month
47
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train navneethdg/JVC

Evaluation results

  • Val stroke_type accuracy on FineBadminton-20K (video-level 80/20 split, seed 42)
    self-reported
    80.610