image

150M model with OLMo-3 style arch (with 2:1 SWA with a 512 token window), pretrained on ~150B tokens of Dolma 3 and further midtrained on ~25B tokens of the OLMo 3 midtraining mix. Compute sponsored by lium.io, thank you <3!

Stage Model Available Checkpoints Data
Pretrained allura-org/Rambley-150M-RealBase allura-forge/Rambley-150M-Base-Checkpoints 150B tokens of Dolma 3
Midtrained allura-org/Rambley-150M-Base (you are here!) allura-forge/Rambley-150M-Midtrain-Checkpoints +25B tokens of the OLMo 3 midtraining mix

Motivation

I swear to god, if I see one more of these random Huggingface posts made by a 5 year old using Claude to pretrain a 150M with some homebrewed datamix, I'm going to lose it. Why not make my own? :) (it was also a fun learning experience so oh well)

Ablations

This cost way too much so I didn't have time for many ablations, but the main thing I tried was XSA attention but I found it to not converge as quickly so we just use regular attention

Benchmarks

Tasks Version Filter n-shot Metric Value Stderr
arc_challenge 1 none 0 acc ↑ 0.2065 ± 0.0118
none 0 acc_norm ↑ 0.2406 ± 0.0125
arc_easy 1 none 0 acc ↑ 0.5072 ± 0.0103
none 0 acc_norm ↑ 0.4823 ± 0.0103
hellaswag 1 none 0 acc ↑ 0.2907 ± 0.0045
none 0 acc_norm ↑ 0.3149 ± 0.0046
piqa 1 none 0 acc ↑ 0.6246 ± 0.0113
none 0 acc_norm ↑ 0.6126 ± 0.0114

Code / Logs

Other random benchmarks
BananaMind Base Bench 1.1
Overall Elo: 991
Accuracy: 177/350 (50.57%)
Weighted accuracy: 46.79%
language_completion: Elo 1278 | 46/50 (92.00%) | weighted 90.45%
commonsense: Elo 1022 | 32/50 (64.00%) | weighted 59.31%
world_knowledge: Elo 1040 | 32/50 (64.00%) | weighted 61.76%
context_tracking: Elo 824 | 14/50 (28.00%) | weighted 27.39%
quantitative: Elo 828 | 11/50 (22.00%) | weighted 22.40%
logical_reasoning: Elo 965 | 18/50 (36.00%) | weighted 32.29%
code_completion: Elo 1076 | 24/50 (48.00%) | weighted 47.45%
====================================================================
  /root/step23201-hf (149,935,616 params) RESULTS
====================================================================
  Raw continuation accuracy        39.90%
  Length-normalized accuracy       39.90%
  Primary (acc_norm)           39.90%
====================================================================
Downloads last month
5
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for allura-org/Rambley-150M-Base

Finetuned
(1)
this model