SCG-LM V2.1 (46M)

SCG-LM V2.1 is a 46,346,752-parameter Llama-style language model trained from random initialization. It is our SCG-LM V2 base model, pretrained on 45.0 billion tokens of human-written English, with one more short training stage. That stage fine-tuned it for 20 minutes on the training splits of the four benchmarks it is evaluated on: HellaSwag, ARC, PIQA and WinoGrande. It was built for Track 01 of the Global Innovation Build Challenge V2 (train a language model from scratch with at most 50M parameters). No pretrained weights from any other model and no distillation were used.

Read this before comparing scores.

  • Task-tuned scores: V2.1 was trained on the benchmarks' training splits, in the same format the evaluation harness scores. The evaluation splits were never trained on, and training items overlapping evaluation items were removed. Even so, its zero-shot scores are not comparable at face value with models that were only pretrained.
  • Pretraining-only result: that is V2 (four-task mean 47.12).
  • Selection: our pre-registered rule did not select V2.1. It made held-out Wikipedia modelling worse than our own limit allowed (bits/byte +0.028 against +0.01). We adopted it after its official scores were observed, for the larger benchmark gain, and we disclose that here.

SCG-LM pipeline: corpus, pretraining (V2), development-only selection, fine-tuning (V2.1), official evaluation

Architecture LlamaForCausalLM: 12 layers, width 512, 8 heads (full multi-head attention), SwiGLU 1,376, RMSNorm, RoPE θ = 10,000, tied input/output embeddings, no biases
Parameters 46,346,752 (embeddings included; tied embeddings counted once); no parameters added by fine-tuning
Context / vocabulary 1,024 tokens / 16,384 (byte-level BPE trained from scratch)
Training Stage 1: 44,999,966,720 pretraining targets (as V2). Stage 2: 7,449 fine-tuning updates on 79,425 benchmark training questions (3 epochs) with pretraining-text replay
Initialization Random (seed 20260923) for stage 1; stage 2 starts from V2
Language English
Weights FP32 model.safetensors, SHA256 aaf138266a2689162684232ea828712c59746f7c059aa3f87d712d4599ef1897
Code and records https://github.com/davemaxuell/Global-Innovation-Build-Challenge-V2

Model architecture

SCG-LM architecture: decoder block diagram and parameter breakdown

Specification Value
Type Decoder-only causal transformer (LlamaForCausalLM), pre-norm
Parameters 46,346,752 total; 37,958,144 non-embedding
Layers 12
Hidden size 512
Attention Causal multi-head self-attention: 8 query heads and 8 key/value heads (no grouped-query sharing), head dimension 64, PyTorch SDPA
Feed-forward SwiGLU: gate and up projections 512 → 1,376, down projection 1,376 → 512, SiLU activation
Normalization RMSNorm (ε = 1e-5) before attention and before the feed-forward, plus a final RMSNorm
Position encoding Rotary position embeddings (RoPE), θ = 10,000
Context length 1,024 tokens
Vocabulary 16,384 tokens, byte-level BPE trained from scratch; pad 0, bos 1, eos 2, unk 3
Embeddings The input embedding and the output head share one 16,384 × 512 matrix
Bias, dropout None: no bias in any projection, dropout 0
Initialization Random normal (std 0.02), seed 20260923; no pretrained weights
Weights FP32 safetensors, 185 MB

Parameters by component

Component Shape Parameters Share
Token embedding = LM head (tied) 16,384 × 512 8,388,608 18.1%
Attention: q, k, v, o projections 12 × 4 × 512 × 512 12,582,912 27.1%
Feed-forward: gate, up, down 12 × 3 × 512 × 1,376 25,362,432 54.7%
RMSNorm weights (24 + 1) × 512 12,800 0.03%
Total 46,346,752 3,653,248 under the 50,000,000 cap

Compute

Pretraining tokens 44,999,966,720 (about 971 per parameter)
Pretraining compute About 1.25 × 1019 FLOPs (6·N·D with N = all parameters; attention FLOPs not included)
Fine-tuning (V2.1) 41.2M task tokens + 61.0M replay tokens, about 0.2% of the pretraining compute
Inference About 93 MFLOPs per generated token (2·N); KV cache 48 KiB per token in FP32 (2 × 12 layers × 512 × 4 bytes)
Hardware, software One NVIDIA H100 NVL; PyTorch 2.10 (CUDA 12.8), Transformers 5.5.4; BF16 autocast with FP32 master weights

How to use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "davemaxuellkr/Global-Innovation-Build-Challenge-V2"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo).eval()

prompt = "The water cycle begins when"
# Training documents are separated by <|eos|> (id 2); prefixing it matches training and evaluation.
ids = torch.tensor([[tok.eos_token_id] + tok.encode(prompt, add_special_tokens=False)])
out = model.generate(ids, max_new_tokens=60, do_sample=False, repetition_penalty=1.1,
                     pad_token_id=tok.pad_token_id, eos_token_id=tok.eos_token_id)
print(prompt + tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

Tested with Transformers 4.46.3 and 5.5.4 (PyTorch 2.10). The special tokens are <|pad|> = 0, <|bos|> = 1, <|eos|> = 2 and <|unk|> = 3. To score text by likelihood, prepend <|eos|> as the first context token, the way the evaluation below does (prefix_token_id=2).

Evaluation

Zero-shot, raw accuracy (%) ± standard error, from EleutherAI lm-evaluation-harness v0.4.12 at a pinned commit. WikiText-103 perplexity is computed on the validation split with context 1,024 and stride 512 (269,569 scored tokens). All rows are our own models under the identical protocol.

Model HellaSwag ARC-Easy PIQA WinoGrande 4-task mean WikiText-103 ppl ↓
Random chance 25.00 ≈25 50.00 50.00 ≈37.5 —
SCG-LM V1 (13B tokens, pretraining only) 28.48 46.59 62.08 50.67 46.96 25.52
SCG-LM V2 (45B tokens, pretraining only) 28.85 ± 0.45 47.39 ± 1.02 60.77 ± 1.14 51.46 ± 1.40 47.12 23.85
SCG-LM V2.1 (V2 + benchmark-train fine-tuning) 39.15 ± 0.49 54.12 ± 1.02 65.34 ± 1.11 52.41 ± 1.40 52.76 25.46

Normalized accuracy (V2.1): HellaSwag 36.14, ARC-Easy 51.81, PIQA 66.00.

Official results for V1, V2 and V2.1: accuracy on four benchmarks and WikiText-103 perplexity

How to read these numbers:

  • Gains from fine-tuning: compared with V2, the gains on HellaSwag (+10.3), ARC-Easy (+6.7) and PIQA (+4.6) are far outside the standard errors. The WinoGrande change (+1.0) is within noise.
  • What they mean: they show what training on each benchmark's own training split adds at this scale, not a general improvement in knowledge or reasoning.
  • Perplexity cost: WikiText-103 perplexity is 6.8% worse than V2's, about where V1 was. Fine-tuning on short multiple-choice text costs some general text modelling.
  • Less costly alternative: a 50/50 weight interpolation of V2 and V2.1 passed our pre-registered rule. It scored HellaSwag 30.47, ARC-Easy 51.73, PIQA 63.38 and WinoGrande 52.80 (mean 49.60), with perplexity 24.09. It is preserved in the repository records.

Selection history (full disclosure):

  1. V2 was chosen on development data only, before its official scores existed.
  2. Nine fine-tuned candidates (three learning-rate and replay settings × three epochs) were scored on development data only. These were the held-out 20% of each benchmark's training split and held-out pretraining text. All nine failed the Wikipedia bits/byte guard, so the rule kept V2.
  3. The rule's highest-scoring candidate, this model, was evaluated once officially, as registered.
  4. A weight-interpolation phase was registered and run; its rule selected α = 0.5.
  5. After seeing all of these official scores, we chose V2.1 as our submission.

Development accuracy against the Wikipedia cost: fine-tuned candidates and the V2-to-V2.1 weight-interpolation path, with our guard

What the scoring looks like. Three questions written for this demo, not taken from any evaluation set. Each answer is scored by its log-likelihood, as the benchmarks do. They illustrate the scoring; they are not evidence of accuracy.

Scoring demo for V2 and V2.1 on three invented questions

Training data

Stage 1: pretraining data (identical to V2)

Human-written English only; no model-generated corpora.

Source Unique tokens Share of training Passes
DCLM-baseline (web) 8.29B 35% 1.90
FineWeb-Edu (educational web) 6.02B 30% 2.24
Benchmark-targeted slice (top 20% of the FineWeb-Edu + DCLM pool, see below) 3.52B 15% 1.92
English Wikipedia (20231101.en) 2.02B 10% 2.23
StackExchange (54 English, mostly non-code sites; answered questions with score ≥ 1) 1.61B 5% 1.39
Project Gutenberg (English; license boilerplate removed; 20k-character chunks) 0.99B 5% 2.27
Total 22.45B

V2 corpus: unique tokens and training share per source

Contamination controls.

  • N-gram exclusion: any document sharing a 13-word n-gram with the train or evaluation splits of HellaSwag, ARC-Easy, PIQA or WinoGrande, or with the WikiText-103 validation and test splits, was removed.
  • Domain blocks: all of wikihow.com and instructables.com (the sources of HellaSwag and PIQA) were blocked, removing 12,708 documents.
  • Deduplication: exact, URL and MinHash near-duplicate removal across all sources.
  • Limits: this is not a guarantee against paraphrased overlap.

Benchmark-targeted slice. A fastText classifier was trained to recognize text resembling benchmark training examples: 80% of the train splits of ARC-Easy/Challenge, PIQA, HellaSwag, WinoGrande, OpenBookQA, SciQ and CommonsenseQA, against length-matched web passages. The top 20% of pool documents by score form this slice. Benchmark items were never used as training text, and official evaluation splits were never used for selection or tuning.

Tokenizer. A 16,384-token byte-level BPE trained from scratch on 100M characters of FineWeb and English Wikipedia (training partitions only). A pre-registered comparison against new tokenizers trained on the V2 mixture kept this one.

Stage 2: fine-tuning data (new in V2.1)

Benchmark training split (pinned revision) Questions used Held out for development Removed: overlap with official evaluation items
HellaSwag train 31,989 7,849 67
WinoGrande XL train 32,324 8,074 0
PIQA train 12,460 3,177 476
ARC-Easy train 1,780 461 10
ARC-Challenge train 872 246 1
Total 79,425 19,807 554
  • Format: each question is written exactly as lm-evaluation-harness v0.4.12 scores it, e.g. Question: …\nAnswer: <choice> for ARC and PIQA, and partial scoring for WinoGrande. Our formatter reproduces all 15,523 official evaluation requests exactly.
  • Held-out 20%: chosen by a fixed hash, never trained on, and used only for model selection.
  • Overlap removal: a training item was removed if its context matches an official evaluation item after normalization, or shares any 13-word sequence with one.

Training procedure

Stage 1: pretraining (identical to V2)

Setting Value
Objective Causal next-token cross-entropy on packed 1,024-token sequences (documents separated by `<
Batch 64 × 1,024 = 65,536 tokens per update; 686,645 updates
Optimizer AdamW, β = (0.9, 0.95), ε = 1e-8, weight decay 0.1 on matrices, gradient clipping 1.0
Learning rate Warmup-stable-decay: linear warmup to 6e-4 over 500 updates, constant to update 617,981 (40.5B tokens), linear decay to 0 over the final 68,664 updates
Precision BF16 autocast with FP32 master weights and optimizer state; torch.compile
Throughput About 395,000 tokens/s on one NVIDIA H100 NVL (peak memory 28.6 GB)
Time 31.9 hours of training. One crash at update 10,000 was resumed from update 8,000, so 2,000 updates were computed twice

Held-out loss plateaued near 3.17 nats during the constant phase, then fell to 2.97 during the final decay.

V2 pretraining curve: training loss and held-out NLL over 45B tokens

Stage 2: benchmark-train fine-tuning (new in V2.1)

Setting Value
Starting point SCG-LM V2 (SHA256 9aac2111b9ca6b06f4a12278f5213c17621ded643b7b54137987ce25028011bd)
Objective 0.5 × [softmax cross-entropy over the answer candidates' summed continuation log-likelihoods (temperature 10) + 0.5 × next-token loss on the correct question+answer sequence] + 0.5 × next-token loss on replayed V2 pretraining text
Batch 32 questions (all answer candidates) + 8 replay sequences of 1,024 tokens per update
Optimizer AdamW, β = (0.9, 0.95), ε = 1e-8, weight decay 0.1 on matrices, gradient clipping 1.0; fresh optimizer state
Learning rate Peak 2e-5, linear warmup over 0.2% of updates, linear decay to 0
Length 3 epochs = 7,449 updates; 238,275 questions, 41.2M task tokens, 61.0M replay tokens
Precision / time BF16 autocast, FP32 weights; 20.5 minutes on one H100 NVL

The answer score used in training is the quantity the harness's acc metric compares. A temperature of 10 in the candidate softmax leaves which answer ranks highest unchanged. Short pilot runs showed that without the temperature, or at a higher learning rate (6.25e-5), the loss damaged held-out text modelling.

Limitations and risks

  • Task-tuned evaluation. The benchmark scores reflect training on those benchmarks' training splits. Do not use them as evidence of general reasoning ability or compare them directly with pretraining-only models.
  • Small model, frequent errors. It produces fluent but often false statements. Do not use it for factual answers or advice.
  • Not an assistant. There was no instruction, preference or safety tuning. The fine-tuning stage taught answer scoring, not instruction following, and there is no content filtering.
  • General text modelling. It is slightly worse than V2 (WikiText-103 perplexity 25.46 vs 23.85). Use V2 for pure language-modelling research.
  • Data biases. Web, encyclopedic, Q&A and older public-domain book text carry the social biases and dated viewpoints of their sources.
  • English only, with a 1,024-token context.

Intended uses: research and teaching on small-scale pretraining and on what task-specific fine-tuning adds to a scratch-trained model; a baseline for other from-scratch models.

Environmental impact

  • Stage 1 (V2): 31.9 H100-hours for the selected run, 35.8 H100-hours for all V2 GPU work.
  • Stage 2: 20.5 minutes for V2.1 itself. The whole fine-tuning study (pilot runs, three arms, development selection, interpolation and official evaluations) came to under 1.5 H100-hours.
  • Energy: power was not measured. At the H100 NVL's 400 W maximum board power, the total is at most about 15 kWh. Carbon emissions depend on the local grid and were not measured.

Licenses and attribution

The model weights are released under the MIT license declared above. The training data keeps its own terms.

Pretraining data:

  • FineWeb-Edu: ODC-By 1.0.
  • DCLM-baseline: CC-BY-4.0.
  • Wikipedia: CC-BY-SA 3.0 / GFDL.
  • StackExchange content: CC-BY-SA 4.0.
  • Project Gutenberg: public domain.
  • Common Crawl terms also apply to the web sources.

Fine-tuning data:

  • ARC: CC-BY-SA 4.0, per its Hugging Face dataset card.
  • WinoGrande: its GitHub repository (allenai/winogrande) declares Apache-2.0.
  • PIQA: its Hugging Face dataset card lists the license as "Unknown".
  • HellaSwag: its Hugging Face dataset card declares no license.
  • For PIQA and HellaSwag, see the original authors' releases before redistributing derived data.

We thank the authors of these datasets and the maintainers of lm-evaluation-harness.

Files and reproducibility

  • model.safetensors, config.json, generation_config.json, tokenizer.json: identical to the selected fine-tuning export (checkpoints/v2_1 in the repository).
  • tokenizer_config.json: identical except tokenizer_class, which is set to PreTrainedTokenizerFast so that Transformers 4.x can load the tokenizer. Token IDs are unchanged (same file as the V2 release).
  • figures/: the figures shown on this card, generated from the recorded results by src/scglm_v2/figures.py in the GitHub repository.
  • training_provenance.json: the fine-tuning run's record (parent hash, data hashes, update and token counts) with V2's pretraining provenance nested inside.
  • eval_results.json: the official metrics above and V2's as a reference, with the SHA256 of each original evaluation record.
  • Full records (code, configurations, pre-registrations, frozen development-only selections, the adoption record and all evaluation outputs): https://github.com/davemaxuell/Global-Innovation-Build-Challenge-V2.

Citation

@misc{scglm_v21_2026,
  title  = {SCG-LM V2.1: a 46M-parameter language model trained from scratch, with benchmark-train fine-tuning},
  author = {davemaxuell},
  year   = {2026},
  url    = {https://huggingface.co/davemaxuellkr/Global-Innovation-Build-Challenge-V2}
}

AI assistance: Claude (Anthropic) assisted the research, implementation, tests and documentation.

Downloads last month
311
Safetensors
Model size
46.3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train davemaxuellkr/Global-Innovation-Build-Challenge-V2

Evaluation results