finetune.py

Model Overview

A nano-scale implementation of the mocov3 architecture, built for retrieval tasks.

Architecture

  • Architecture: mocov3
  • Scale: nano
  • Attention: standard
  • Fusion strategy: co attention
  • Task head: retrieval
  • Activation: gelu
  • Normalization: layernorm
  • Initialization: trunc normal

Training

  • Optimizer: lamb
  • LR scheduler: linear warmup

Files

  • finetune.py — main artifact of this repository

License

See the license field above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support