inference.py
Model Overview
A giant-scale implementation of the mixer architecture, built for multitask tasks.
Architecture
- Architecture: mixer
- Scale: giant
- Attention: linear
- Fusion strategy: tucker
- Task head: multitask
- Activation: swish
- Normalization: layernorm
- Initialization: kaiming
Training
- Optimizer: lamb
- LR scheduler: constant warmup
Files
inference.py— main artifact of this repository
License
See the license field above.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support