Instructions to use davemaxuellkr/Global-Innovation-Build-Challenge-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use davemaxuellkr/Global-Innovation-Build-Challenge-V2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="davemaxuellkr/Global-Innovation-Build-Challenge-V2")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("davemaxuellkr/Global-Innovation-Build-Challenge-V2") model = AutoModelForCausalLM.from_pretrained("davemaxuellkr/Global-Innovation-Build-Challenge-V2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use davemaxuellkr/Global-Innovation-Build-Challenge-V2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "davemaxuellkr/Global-Innovation-Build-Challenge-V2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "davemaxuellkr/Global-Innovation-Build-Challenge-V2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/davemaxuellkr/Global-Innovation-Build-Challenge-V2
- SGLang
How to use davemaxuellkr/Global-Innovation-Build-Challenge-V2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "davemaxuellkr/Global-Innovation-Build-Challenge-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "davemaxuellkr/Global-Innovation-Build-Challenge-V2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "davemaxuellkr/Global-Innovation-Build-Challenge-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "davemaxuellkr/Global-Innovation-Build-Challenge-V2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use davemaxuellkr/Global-Innovation-Build-Challenge-V2 with Docker Model Runner:
docker model run hf.co/davemaxuellkr/Global-Innovation-Build-Challenge-V2
SCG-LM V2.1 (46M)
SCG-LM V2.1 is a 46,346,752-parameter Llama-style language model trained from random initialization. It is our SCG-LM V2 base model, pretrained on 45.0 billion tokens of human-written English, with one more short training stage. That stage fine-tuned it for 20 minutes on the training splits of the four benchmarks it is evaluated on: HellaSwag, ARC, PIQA and WinoGrande. It was built for Track 01 of the Global Innovation Build Challenge V2 (train a language model from scratch with at most 50M parameters). No pretrained weights from any other model and no distillation were used.
Read this before comparing scores.
- Task-tuned scores: V2.1 was trained on the benchmarks' training splits, in the same format the evaluation harness scores. The evaluation splits were never trained on, and training items overlapping evaluation items were removed. Even so, its zero-shot scores are not comparable at face value with models that were only pretrained.
- Pretraining-only result: that is V2 (four-task mean 47.12).
- Selection: our pre-registered rule did not select V2.1. It made held-out Wikipedia modelling worse than our own limit allowed (bits/byte +0.028 against +0.01). We adopted it after its official scores were observed, for the larger benchmark gain, and we disclose that here.
| Architecture | LlamaForCausalLM: 12 layers, width 512, 8 heads (full multi-head attention), SwiGLU 1,376, RMSNorm, RoPE θ = 10,000, tied input/output embeddings, no biases |
| Parameters | 46,346,752 (embeddings included; tied embeddings counted once); no parameters added by fine-tuning |
| Context / vocabulary | 1,024 tokens / 16,384 (byte-level BPE trained from scratch) |
| Training | Stage 1: 44,999,966,720 pretraining targets (as V2). Stage 2: 7,449 fine-tuning updates on 79,425 benchmark training questions (3 epochs) with pretraining-text replay |
| Initialization | Random (seed 20260923) for stage 1; stage 2 starts from V2 |
| Language | English |
| Weights | FP32 model.safetensors, SHA256 aaf138266a2689162684232ea828712c59746f7c059aa3f87d712d4599ef1897 |
| Code and records | https://github.com/davemaxuell/Global-Innovation-Build-Challenge-V2 |
Model architecture
| Specification | Value |
|---|---|
| Type | Decoder-only causal transformer (LlamaForCausalLM), pre-norm |
| Parameters | 46,346,752 total; 37,958,144 non-embedding |
| Layers | 12 |
| Hidden size | 512 |
| Attention | Causal multi-head self-attention: 8 query heads and 8 key/value heads (no grouped-query sharing), head dimension 64, PyTorch SDPA |
| Feed-forward | SwiGLU: gate and up projections 512 → 1,376, down projection 1,376 → 512, SiLU activation |
| Normalization | RMSNorm (ε = 1e-5) before attention and before the feed-forward, plus a final RMSNorm |
| Position encoding | Rotary position embeddings (RoPE), θ = 10,000 |
| Context length | 1,024 tokens |
| Vocabulary | 16,384 tokens, byte-level BPE trained from scratch; pad 0, bos 1, eos 2, unk 3 |
| Embeddings | The input embedding and the output head share one 16,384 × 512 matrix |
| Bias, dropout | None: no bias in any projection, dropout 0 |
| Initialization | Random normal (std 0.02), seed 20260923; no pretrained weights |
| Weights | FP32 safetensors, 185 MB |
Parameters by component
| Component | Shape | Parameters | Share |
|---|---|---|---|
| Token embedding = LM head (tied) | 16,384 × 512 | 8,388,608 | 18.1% |
| Attention: q, k, v, o projections | 12 × 4 × 512 × 512 | 12,582,912 | 27.1% |
| Feed-forward: gate, up, down | 12 × 3 × 512 × 1,376 | 25,362,432 | 54.7% |
| RMSNorm weights | (24 + 1) × 512 | 12,800 | 0.03% |
| Total | 46,346,752 | 3,653,248 under the 50,000,000 cap |
Compute
| Pretraining tokens | 44,999,966,720 (about 971 per parameter) |
| Pretraining compute | About 1.25 × 1019 FLOPs (6·N·D with N = all parameters; attention FLOPs not included) |
| Fine-tuning (V2.1) | 41.2M task tokens + 61.0M replay tokens, about 0.2% of the pretraining compute |
| Inference | About 93 MFLOPs per generated token (2·N); KV cache 48 KiB per token in FP32 (2 × 12 layers × 512 × 4 bytes) |
| Hardware, software | One NVIDIA H100 NVL; PyTorch 2.10 (CUDA 12.8), Transformers 5.5.4; BF16 autocast with FP32 master weights |
How to use
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "davemaxuellkr/Global-Innovation-Build-Challenge-V2"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo).eval()
prompt = "The water cycle begins when"
# Training documents are separated by <|eos|> (id 2); prefixing it matches training and evaluation.
ids = torch.tensor([[tok.eos_token_id] + tok.encode(prompt, add_special_tokens=False)])
out = model.generate(ids, max_new_tokens=60, do_sample=False, repetition_penalty=1.1,
pad_token_id=tok.pad_token_id, eos_token_id=tok.eos_token_id)
print(prompt + tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
Tested with Transformers 4.46.3 and 5.5.4 (PyTorch 2.10). The special tokens are <|pad|> = 0, <|bos|> = 1, <|eos|> = 2 and <|unk|> = 3. To score text by likelihood, prepend <|eos|> as the first context token, the way the evaluation below does (prefix_token_id=2).
Evaluation
Zero-shot, raw accuracy (%) ± standard error, from EleutherAI lm-evaluation-harness v0.4.12 at a pinned commit. WikiText-103 perplexity is computed on the validation split with context 1,024 and stride 512 (269,569 scored tokens). All rows are our own models under the identical protocol.
| Model | HellaSwag | ARC-Easy | PIQA | WinoGrande | 4-task mean | WikiText-103 ppl ↓ |
|---|---|---|---|---|---|---|
| Random chance | 25.00 | ≈25 | 50.00 | 50.00 | ≈37.5 | — |
| SCG-LM V1 (13B tokens, pretraining only) | 28.48 | 46.59 | 62.08 | 50.67 | 46.96 | 25.52 |
| SCG-LM V2 (45B tokens, pretraining only) | 28.85 ± 0.45 | 47.39 ± 1.02 | 60.77 ± 1.14 | 51.46 ± 1.40 | 47.12 | 23.85 |
| SCG-LM V2.1 (V2 + benchmark-train fine-tuning) | 39.15 ± 0.49 | 54.12 ± 1.02 | 65.34 ± 1.11 | 52.41 ± 1.40 | 52.76 | 25.46 |
Normalized accuracy (V2.1): HellaSwag 36.14, ARC-Easy 51.81, PIQA 66.00.
How to read these numbers:
- Gains from fine-tuning: compared with V2, the gains on HellaSwag (+10.3), ARC-Easy (+6.7) and PIQA (+4.6) are far outside the standard errors. The WinoGrande change (+1.0) is within noise.
- What they mean: they show what training on each benchmark's own training split adds at this scale, not a general improvement in knowledge or reasoning.
- Perplexity cost: WikiText-103 perplexity is 6.8% worse than V2's, about where V1 was. Fine-tuning on short multiple-choice text costs some general text modelling.
- Less costly alternative: a 50/50 weight interpolation of V2 and V2.1 passed our pre-registered rule. It scored HellaSwag 30.47, ARC-Easy 51.73, PIQA 63.38 and WinoGrande 52.80 (mean 49.60), with perplexity 24.09. It is preserved in the repository records.
Selection history (full disclosure):
- V2 was chosen on development data only, before its official scores existed.
- Nine fine-tuned candidates (three learning-rate and replay settings × three epochs) were scored on development data only. These were the held-out 20% of each benchmark's training split and held-out pretraining text. All nine failed the Wikipedia bits/byte guard, so the rule kept V2.
- The rule's highest-scoring candidate, this model, was evaluated once officially, as registered.
- A weight-interpolation phase was registered and run; its rule selected α = 0.5.
- After seeing all of these official scores, we chose V2.1 as our submission.
What the scoring looks like. Three questions written for this demo, not taken from any evaluation set. Each answer is scored by its log-likelihood, as the benchmarks do. They illustrate the scoring; they are not evidence of accuracy.
Training data
Stage 1: pretraining data (identical to V2)
Human-written English only; no model-generated corpora.
| Source | Unique tokens | Share of training | Passes |
|---|---|---|---|
| DCLM-baseline (web) | 8.29B | 35% | 1.90 |
| FineWeb-Edu (educational web) | 6.02B | 30% | 2.24 |
| Benchmark-targeted slice (top 20% of the FineWeb-Edu + DCLM pool, see below) | 3.52B | 15% | 1.92 |
| English Wikipedia (20231101.en) | 2.02B | 10% | 2.23 |
| StackExchange (54 English, mostly non-code sites; answered questions with score ≥ 1) | 1.61B | 5% | 1.39 |
| Project Gutenberg (English; license boilerplate removed; 20k-character chunks) | 0.99B | 5% | 2.27 |
| Total | 22.45B |
Contamination controls.
- N-gram exclusion: any document sharing a 13-word n-gram with the train or evaluation splits of HellaSwag, ARC-Easy, PIQA or WinoGrande, or with the WikiText-103 validation and test splits, was removed.
- Domain blocks: all of
wikihow.comandinstructables.com(the sources of HellaSwag and PIQA) were blocked, removing 12,708 documents. - Deduplication: exact, URL and MinHash near-duplicate removal across all sources.
- Limits: this is not a guarantee against paraphrased overlap.
Benchmark-targeted slice. A fastText classifier was trained to recognize text resembling benchmark training examples: 80% of the train splits of ARC-Easy/Challenge, PIQA, HellaSwag, WinoGrande, OpenBookQA, SciQ and CommonsenseQA, against length-matched web passages. The top 20% of pool documents by score form this slice. Benchmark items were never used as training text, and official evaluation splits were never used for selection or tuning.
Tokenizer. A 16,384-token byte-level BPE trained from scratch on 100M characters of FineWeb and English Wikipedia (training partitions only). A pre-registered comparison against new tokenizers trained on the V2 mixture kept this one.
Stage 2: fine-tuning data (new in V2.1)
| Benchmark training split (pinned revision) | Questions used | Held out for development | Removed: overlap with official evaluation items |
|---|---|---|---|
| HellaSwag train | 31,989 | 7,849 | 67 |
| WinoGrande XL train | 32,324 | 8,074 | 0 |
| PIQA train | 12,460 | 3,177 | 476 |
| ARC-Easy train | 1,780 | 461 | 10 |
| ARC-Challenge train | 872 | 246 | 1 |
| Total | 79,425 | 19,807 | 554 |
- Format: each question is written exactly as lm-evaluation-harness v0.4.12 scores it, e.g.
Question: …\nAnswer: <choice>for ARC and PIQA, and partial scoring for WinoGrande. Our formatter reproduces all 15,523 official evaluation requests exactly. - Held-out 20%: chosen by a fixed hash, never trained on, and used only for model selection.
- Overlap removal: a training item was removed if its context matches an official evaluation item after normalization, or shares any 13-word sequence with one.
Training procedure
Stage 1: pretraining (identical to V2)
| Setting | Value |
|---|---|
| Objective | Causal next-token cross-entropy on packed 1,024-token sequences (documents separated by `< |
| Batch | 64 × 1,024 = 65,536 tokens per update; 686,645 updates |
| Optimizer | AdamW, β = (0.9, 0.95), ε = 1e-8, weight decay 0.1 on matrices, gradient clipping 1.0 |
| Learning rate | Warmup-stable-decay: linear warmup to 6e-4 over 500 updates, constant to update 617,981 (40.5B tokens), linear decay to 0 over the final 68,664 updates |
| Precision | BF16 autocast with FP32 master weights and optimizer state; torch.compile |
| Throughput | About 395,000 tokens/s on one NVIDIA H100 NVL (peak memory 28.6 GB) |
| Time | 31.9 hours of training. One crash at update 10,000 was resumed from update 8,000, so 2,000 updates were computed twice |
Held-out loss plateaued near 3.17 nats during the constant phase, then fell to 2.97 during the final decay.
Stage 2: benchmark-train fine-tuning (new in V2.1)
| Setting | Value |
|---|---|
| Starting point | SCG-LM V2 (SHA256 9aac2111b9ca6b06f4a12278f5213c17621ded643b7b54137987ce25028011bd) |
| Objective | 0.5 × [softmax cross-entropy over the answer candidates' summed continuation log-likelihoods (temperature 10) + 0.5 × next-token loss on the correct question+answer sequence] + 0.5 × next-token loss on replayed V2 pretraining text |
| Batch | 32 questions (all answer candidates) + 8 replay sequences of 1,024 tokens per update |
| Optimizer | AdamW, β = (0.9, 0.95), ε = 1e-8, weight decay 0.1 on matrices, gradient clipping 1.0; fresh optimizer state |
| Learning rate | Peak 2e-5, linear warmup over 0.2% of updates, linear decay to 0 |
| Length | 3 epochs = 7,449 updates; 238,275 questions, 41.2M task tokens, 61.0M replay tokens |
| Precision / time | BF16 autocast, FP32 weights; 20.5 minutes on one H100 NVL |
The answer score used in training is the quantity the harness's acc metric compares. A temperature of 10 in the candidate softmax leaves which answer ranks highest unchanged. Short pilot runs showed that without the temperature, or at a higher learning rate (6.25e-5), the loss damaged held-out text modelling.
Limitations and risks
- Task-tuned evaluation. The benchmark scores reflect training on those benchmarks' training splits. Do not use them as evidence of general reasoning ability or compare them directly with pretraining-only models.
- Small model, frequent errors. It produces fluent but often false statements. Do not use it for factual answers or advice.
- Not an assistant. There was no instruction, preference or safety tuning. The fine-tuning stage taught answer scoring, not instruction following, and there is no content filtering.
- General text modelling. It is slightly worse than V2 (WikiText-103 perplexity 25.46 vs 23.85). Use V2 for pure language-modelling research.
- Data biases. Web, encyclopedic, Q&A and older public-domain book text carry the social biases and dated viewpoints of their sources.
- English only, with a 1,024-token context.
Intended uses: research and teaching on small-scale pretraining and on what task-specific fine-tuning adds to a scratch-trained model; a baseline for other from-scratch models.
Environmental impact
- Stage 1 (V2): 31.9 H100-hours for the selected run, 35.8 H100-hours for all V2 GPU work.
- Stage 2: 20.5 minutes for V2.1 itself. The whole fine-tuning study (pilot runs, three arms, development selection, interpolation and official evaluations) came to under 1.5 H100-hours.
- Energy: power was not measured. At the H100 NVL's 400 W maximum board power, the total is at most about 15 kWh. Carbon emissions depend on the local grid and were not measured.
Licenses and attribution
The model weights are released under the MIT license declared above. The training data keeps its own terms.
Pretraining data:
- FineWeb-Edu: ODC-By 1.0.
- DCLM-baseline: CC-BY-4.0.
- Wikipedia: CC-BY-SA 3.0 / GFDL.
- StackExchange content: CC-BY-SA 4.0.
- Project Gutenberg: public domain.
- Common Crawl terms also apply to the web sources.
Fine-tuning data:
- ARC: CC-BY-SA 4.0, per its Hugging Face dataset card.
- WinoGrande: its GitHub repository (allenai/winogrande) declares Apache-2.0.
- PIQA: its Hugging Face dataset card lists the license as "Unknown".
- HellaSwag: its Hugging Face dataset card declares no license.
- For PIQA and HellaSwag, see the original authors' releases before redistributing derived data.
We thank the authors of these datasets and the maintainers of lm-evaluation-harness.
Files and reproducibility
model.safetensors,config.json,generation_config.json,tokenizer.json: identical to the selected fine-tuning export (checkpoints/v2_1in the repository).tokenizer_config.json: identical excepttokenizer_class, which is set toPreTrainedTokenizerFastso that Transformers 4.x can load the tokenizer. Token IDs are unchanged (same file as the V2 release).figures/: the figures shown on this card, generated from the recorded results bysrc/scglm_v2/figures.pyin the GitHub repository.training_provenance.json: the fine-tuning run's record (parent hash, data hashes, update and token counts) with V2's pretraining provenance nested inside.eval_results.json: the official metrics above and V2's as a reference, with the SHA256 of each original evaluation record.- Full records (code, configurations, pre-registrations, frozen development-only selections, the adoption record and all evaluation outputs): https://github.com/davemaxuell/Global-Innovation-Build-Challenge-V2.
Citation
@misc{scglm_v21_2026,
title = {SCG-LM V2.1: a 46M-parameter language model trained from scratch, with benchmark-train fine-tuning},
author = {davemaxuell},
year = {2026},
url = {https://huggingface.co/davemaxuellkr/Global-Innovation-Build-Challenge-V2}
}
AI assistance: Claude (Anthropic) assisted the research, implementation, tests and documentation.
- Downloads last month
- 311
Datasets used to train davemaxuellkr/Global-Innovation-Build-Challenge-V2
allenai/ai2_arc
HuggingFaceFW/fineweb-edu
Evaluation results
- accuracy (0-shot) on HellaSwagvalidation set self-reported39.150
- normalized accuracy (0-shot) on HellaSwagvalidation set self-reported36.140
- accuracy (0-shot) on ARC-Easytest set self-reported54.120
- normalized accuracy (0-shot) on ARC-Easytest set self-reported51.810
- accuracy (0-shot) on PIQAvalidation set self-reported65.340
- normalized accuracy (0-shot) on PIQAvalidation set self-reported66.000
- accuracy (0-shot) on WinoGrandevalidation set self-reported52.410
- perplexity (context 1024, stride 512) on WikiText-103validation set self-reported25.460






