Wavy Transformer mitigates this issue by introducing:
- a novel attention layer based on second-order wavy dynamics,
- a feed-forward network and normalization layer that preserve the physical state–velocity relationship implied by the chain rule.
Across diverse NLP and CV benchmarks, Wavy Transformer consistently improves performance with minimal extra parameters and no additional hyper-parameter tuning.
Contents
This repository comprises two main components:
- NLP Tasks: Everything related to pretraining the BERT-base model, fine-tuning on downstream benchmarks (GLUE and SQ2AD), and analyzing oversmoothing behavior.
- CV Tasks: Scripts and examples for ImageNet object classification using Vision Transformers, including training, evaluation, and analyzing oversmoothing behavior.
Citation
A formal citation will be provided here as soon as our paper is publicly available (arXiv / conference proceedings, currently in preparation). In the meantime, if you find Wavy Transformer useful for your research or applications, please consider pointing to this repository.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support