llama3-7b-en2-en1-codeswitching-134k

LLaMA-3 7B pretrained on English bilingual data (en1 + en2) with code-switching, 134k steps. Uses the en1 embedding matrix (--extract_vocab 0) from the 2×vocabulary checkpoint. For the en2 embedding matrix version see The-CoLab/llama3-7b-en1-en2-codeswitching-134k.

Validation Results

Final checkpoint: step 133,600

Validation set Cross-entropy loss (nats) Perplexity
English 1 (en1) 2.2185 9.19
English 2 (en2) 2.2175 9.18

Training Curves

See The-CoLab/llama3-7b-en1-en2-codeswitching-134k for training curves (same checkpoint, different embedding extraction).

Citation

Part of the The-CoLab multilingual-transfer collection.

Downloads last month
7
Safetensors
Model size
6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including The-CoLab/llama3-7b-en2-en1-codeswitching-134k