Llama1BxFWx1024x0pct

This repository contains a raw GPT-NeoX training checkpoint from the Partial RoPE experiments accompanying Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE. The paper was accepted to EMNLP 2026.

Checkpoint details

Field Value
Architecture Llama 3.2 1B
Dataset FineWeb (FW)
Training sequence length 1,024 tokens
Partial RoPE 0%
Checkpoint Global step 12,000
Format Raw GPT-NeoX checkpoint (not Transformers format)

The files are preserved in their original GPT-NeoX checkpoint format and have not been converted to the Hugging Face Transformers format.

Resources

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including aflah/Llama1BxFWx1024x0pct

Paper for aflah/Llama1BxFWx1024x0pct