OLMo 3 3B Baseline โ€” Stage 2 Mid-training

This repository is the Hugging Face export of o3b3b-s2-s8192-g256-m1-ga1-tp2-cp1-dp256-h32-b8-mc2-lr2p071e4-512npu-share-0905081317-s2v1 at iteration 47684. This is the matched pure OLMo 3 baseline. It uses Transformers' official Olmo3ForCausalLM implementation and does not require remote code.

  • Training sequence length: 8,192
  • Model context capacity: 8,192
  • Sliding-window size: 4,096
  • Attention pattern: [SWA, SWA, SWA, Full]
  • Vocabulary: 100,278 real tokens; 74 Megatron padding-only rows removed

Stage 3/4 use the frozen 65,536-token configuration. YaRN applies to the Full Attention layers; SWA layers retain their original RoPE and 4,096-token local window.

Loading

Use transformers>=4.57.6,<5.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "ArchSpace-Collection/OLMo3-3B-stage2"
tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    use_fast=True,
    fix_mistral_regex=False,
)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    dtype=torch.bfloat16,
    attn_implementation="sdpa",
)

fix_mistral_regex=False preserves the exact tokenizer behavior used during training. Conversion provenance, per-tensor hashes, and CPU validation results are included in conversion_manifest.json, SHA256SUMS, and hf_validation_report.json.

Downloads last month
5
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support