OLMo 3 3B Baseline โ€” Stage 4 Instruct SFT

This repository is the Hugging Face export of o3b3b-instruct-sft-dolci-s32768-g32-m1-ga1-tp2-cp8-dp32-h32-b2-lr3e5-min0-wd5e2-wu10pct-2ep-512npu-share-20260906t095032z-s4v2 at iteration 3252. This is the matched pure OLMo 3 baseline. It uses Transformers' official Olmo3ForCausalLM implementation and does not require remote code.

  • Training sequence length: 32,768
  • Model context capacity: 65,536
  • Sliding-window size: 4,096
  • Attention pattern: [SWA, SWA, SWA, Full]
  • Vocabulary: 100,278 real tokens; 74 Megatron padding-only rows removed

Stage 3/4 use the frozen 65,536-token configuration. YaRN applies to the Full Attention layers; SWA layers retain their original RoPE and 4,096-token local window.

Loading

Use transformers>=4.57.6,<5.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "ArchSpace-Collection/OLMo3-3B-stage4-instruct"
tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    use_fast=True,
    fix_mistral_regex=False,
)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    dtype=torch.bfloat16,
    attn_implementation="sdpa",
)

fix_mistral_regex=False preserves the exact tokenizer behavior used during training. Conversion provenance, per-tensor hashes, and CPU validation results are included in conversion_manifest.json, SHA256SUMS, and hf_validation_report.json.

Downloads last month
5
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support