OpenSML-150M

Author: William Zebrowski · Checkpoint: OpenSML-150M Instruct · Status: research preview

Technical report · Results and provenance · Tokenizer

An English-first language model trained from scratch with Apple's MLX framework. Pretraining processed 7.800B tokens across Stage A and Stage B on a four-Mac cluster, later expanded to five Macs using Thunderbolt RDMA. Three supervised fine-tuning stages produced the selected OpenSML-150M checkpoint.

Availability and Inference

The selected OpenSML-150M weights, frozen tokenizer, and standalone native MLX inference code are included. The OpenSML-150M Instruct checkpoint contains the selected fine-tuned FP32 weights; optimizer state is excluded. Its training-step identity is retained in the checkpoint provenance.

Use Python 3.11+ on Apple Silicon. Install the Hugging Face CLI first and authenticate if this repository is private:

hf download wzebrowski/OpenSML-150M --local-dir ./OpenSML-150M
cd OpenSML-150M
python -m pip install -r requirements.txt
python inference.py --prompt "Say hello in one sentence." --max-new-tokens 32

The CLI verifies the weight hash and tokenizer manifest, then uses greedy FP32 reference MLX inference. Prompts use User: {prompt}\nAssistant:; --raw-completion skips this wrapper. The 2,048-token context includes the requested generation budget; overflow is rejected. Output includes text, generated token IDs, and the stop reason. Native loading and a short generation were verified locally; this is not a Transformers AutoModel or mlx-lm loader package. Other formats and cross-framework parity remain unverified.

See checkpoint provenance, loading verification, and bundle hashes.

Model Specification

Field Value
Parameters (model-card label) 150,439,188 (150.44M)
Layers / hidden width / FFN width 20 / 768 / 2,048
Attention GQA: 12 query heads / 4 KV heads; head dimension 64
Position encoding RoPE, base 10,000
Normalization / MLP RMSNorm, Q/K normalization, SwiGLU
Vocabulary / context 32,000 / 2,048
Embeddings Tied input and output
Linear biases / dropout Neither used
Pretraining precision BF16 compute; FP32 master weights and optimizer state
Benchmark precision FP32 scoring, explicit vanilla MLX attention

The independently fitted tokenizer is a 32,000-token byte-level BPE. Its recorded held-out audit had zero round-trip failures. Full configuration, parameter accounting, and tokenizer checks are in the technical report.

Training and Model Lineage

The selected pretrained base is Stage B step 73,243, after 7,800,086,528 lifetime pretraining tokens. Stage B continued Stage A with the adjusted mixture below. Percentages are configured token shares.

Source Initial token share Stage B token share Selection
FineWeb-Edu 55% 55% fineweb-edu-dedup
DCLM-Edu 25% 20% edu_int_score >= 3
FineWiki 10% 15% English
Cosmopedia v2 10% 10% Textbook/tutorial/blog/educational format filter

FineWeb-Edu and Cosmopedia v2 come from two subsets of the same SmolLM Corpus repository. Pretraining used no dedicated code or math dataset.

V1 pretrained base → Unified384 → Repair512 → OpenSML-150M (SFT step 768)
                      +384         +128       +256 updates

Unified384 used Smol-SmolTalk, UltraChat, SQuAD v2, ARC training questions, and locally authored follow-ups. Repair512 added Dolly and Tulu Persona instruction examples. The final stage used SmolTalk constraints, SQuAD v2, SciQ, and Repair512 replay. All model parameters were updated; the selected checkpoint is a direct continuation, without LoRA or parameter interpolation. Dataset exposure, source pins, optimizer schedules, and checkpoint hashes are in the technical report.

Pretraining Loss Curves

Recorded pretraining validation loss for Stage A and Stage B

Recorded held-out loss without smoothing. The star marks the selected base; later points are outside its training exposure. The late-training panel uses a magnified scale. This chart uses the same original validation set throughout.

Stage B detail and per-source curves

Stage B validation loss on the expanded held-out set

This panel uses a different, expanded validation set; its absolute losses should not be merged with the overview. The selected base has original-set loss 2.81855024 and expanded-set loss 2.76700753.

Original-set validation loss by data source

Per-source panels use separate vertical scales. Full validation accounting and historical measurements are in the technical report.

Evaluation

Zero-shot Likelihood Benchmarks

Existing full-split evaluations cover 15,428 multiple-choice examples. The pretrained base and selected SFT model are reported separately.

Benchmark Split Examples Pretrained base acc / acc_norm OpenSML-150M acc / acc_norm
ARC-Easy Test 2,376 54.67% / 48.53% 56.65% / 55.43%
ARC-Challenge Test 1,172 23.38% / 26.88% 26.02% / 29.52%
PIQA Validation 1,838 65.23% / 64.36% 64.53% / 64.09%
HellaSwag Validation 10,042 30.21% / 34.44% 30.55% / 33.94%

SFT improves both ARC metrics. PIQA declines slightly; HellaSwag raw accuracy rises slightly while normalized accuracy declines.

Likelihood prompts and scoring

The native evaluator uses zero-shot raw completion prompts, FP32 scoring, reference MLX attention, batch size 1, and no chat template or added BOS/EOS. acc ranks summed candidate log-likelihood; acc_norm divides by candidate Unicode character count, excluding the leading delimiter. No generation or LLM judge is involved. The evaluator follows pinned harness conventions but is not an installed lm-evaluation-harness run. Dataset pins and integrity receipts are in the technical report and results record.

Comparison with Small Language Models

External scores below are previously published measurements. Each cell shows acc / acc_norm; — means not reported in the selected source. Bold marks the highest available reported value.

Model ARC-Easy acc / acc_norm ARC-Challenge acc / acc_norm PIQA acc / acc_norm HellaSwag acc / acc_norm
OpenSML-150M (SFT) 56.65% / 55.43% 26.02% / 29.52% 64.53% / 64.09% 30.55% / 33.94%
GPT-2 (124M, base) — / 39.48% — / — — / 62.51% 28.92% / 31.14%
OPT-125M (base) 43.52% / 39.98% 18.94% / 22.78% 63.00% / 62.02% — / —
Pythia-160M (base) 43.52% / 39.65% 18.77% / 23.29% 62.73% / 61.64% — / —

OpenSML's largest reported advantage is normalized ARC-Easy: +15.45 percentage points over OPT-125M and +15.78 over Pythia-160M. These are cross-source comparisons, not a controlled ranking: prompts, normalization, numerical settings, and dataset revisions are not verified identical. OpenSML is SFT-trained; the references are base models. Source records are in the technical report.

Instruction Following: IFEval

Checkpoint Prompt strict Instruction strict Prompt loose Instruction loose 1,280-token cap hits
OpenSML-150M 15.16% (82/541) 25.30% (211/834) 15.71% (85/541) 25.78% (215/834) 33/541

These results apply to OpenSML-150M, not the pretrained base. All 541 prompts and 834 instructions were scored using programmatic strict/loose verifiers, without an LLM judge.

IFEval formatting, generation, and scoring

Generation is zero-shot and greedy, with plain User: {prompt}\nAssistant: formatting, no added system message, and a 1,280-new-token cap. The context is 2,048 tokens; document-end EOS is ID 1. No additional stop strings or repetition penalty are used. All responses are scored as produced: 508 stop at EOS and 33 reach the generation cap; none hit the context limit. The dataset's only split is named train, but these prompts are evaluation inputs, not supervised records.

Complete-answer Diagnostics

Saved development checks report a follow-up joint proxy of 22/32 and named-constraint passes on 5/32 public cases. Repetition was detected on 4/63 legacy turns and 18/73 public-development turns. These are mechanical development proxies; a systematic independent review of complete-answer correctness is still pending. Likelihood and IFEval do not establish reliable factual answers. No MT-Bench or coding-success score is reported for the selected model.

Intended Use and Limitations

Intended for small-language-model research and controlled experimentation. Responses may be incorrect, repetitive, biased, or unsafe. Production use, safety behavior, tool use, multilingual ability, and extended dialogue remain unvalidated. Benchmark reuse during selection limits evaluation independence; a comprehensive contamination audit has not been established.

Checkpoint Verification

Saved hashes and integrity receipts identify the selected weights and completed evaluations. They do not establish cross-framework inference parity. An unchanged native weight export and standalone inference CLI are included. The export has a local loading/generation check; a full benchmark rerun of this download bundle has not been performed. See the results and provenance record and technical report.

Licensing and Data Provenance

Author-controlled project contributions use Apache-2.0. Upstream materials retain their own terms. SFT includes SQuAD v2 (CC-BY-SA-4.0) and SciQ (CC-BY-NC-3.0); repository metadata does not grant blanket commercial rights to all training materials. Weight-release licensing review remains pending.

Downloads last month
905
Safetensors
Model size
0.2B params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train wzebrowski/OpenSML-150M