150M model with OLMo-3 style arch (with 2:1 SWA with a 512 token window), pretrained on ~150B tokens of Dolma 3 and further midtrained on ~25B tokens of the OLMo 3 midtraining mix. Compute sponsored by lium.io, thank you <3!
| Stage | Model | Available Checkpoints | Data |
|---|---|---|---|
| Pretrained | allura-org/Rambley-150M-RealBase | allura-forge/Rambley-150M-Base-Checkpoints | 150B tokens of Dolma 3 |
| Midtrained | allura-org/Rambley-150M-Base (you are here!) | allura-forge/Rambley-150M-Midtrain-Checkpoints | +25B tokens of the OLMo 3 midtraining mix |
Motivation
I swear to god, if I see one more of these random Huggingface posts made by a 5 year old using Claude to pretrain a 150M with some homebrewed datamix, I'm going to lose it. Why not make my own? :) (it was also a fun learning experience so oh well)
Ablations
This cost way too much so I didn't have time for many ablations, but the main thing I tried was XSA attention but I found it to not converge as quickly so we just use regular attention
Benchmarks
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| arc_challenge | 1 | none | 0 | acc | ↑ | 0.2065 | ± | 0.0118 |
| none | 0 | acc_norm | ↑ | 0.2406 | ± | 0.0125 | ||
| arc_easy | 1 | none | 0 | acc | ↑ | 0.5072 | ± | 0.0103 |
| none | 0 | acc_norm | ↑ | 0.4823 | ± | 0.0103 | ||
| hellaswag | 1 | none | 0 | acc | ↑ | 0.2907 | ± | 0.0045 |
| none | 0 | acc_norm | ↑ | 0.3149 | ± | 0.0046 | ||
| piqa | 1 | none | 0 | acc | ↑ | 0.6246 | ± | 0.0113 |
| none | 0 | acc_norm | ↑ | 0.6126 | ± | 0.0114 |
Code / Logs
- Pretraining/midtraining codebase: https://code.allura.moe/fizz/OLMo-core
- Wandb pretraining logs
- Wandb midtraining logs
Other random benchmarks
BananaMind Base Bench 1.1
Overall Elo: 991
Accuracy: 177/350 (50.57%)
Weighted accuracy: 46.79%
language_completion: Elo 1278 | 46/50 (92.00%) | weighted 90.45%
commonsense: Elo 1022 | 32/50 (64.00%) | weighted 59.31%
world_knowledge: Elo 1040 | 32/50 (64.00%) | weighted 61.76%
context_tracking: Elo 824 | 14/50 (28.00%) | weighted 27.39%
quantitative: Elo 828 | 11/50 (22.00%) | weighted 22.40%
logical_reasoning: Elo 965 | 18/50 (36.00%) | weighted 32.29%
code_completion: Elo 1076 | 24/50 (48.00%) | weighted 47.45%
====================================================================
/root/step23201-hf (149,935,616 params) RESULTS
====================================================================
Raw continuation accuracy 39.90%
Length-normalized accuracy 39.90%
Primary (acc_norm) 39.90%
====================================================================
- Downloads last month
- 5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for allura-org/Rambley-150M-Base
Base model
allura-org/Rambley-150M-RealBase