1gpu-llm Small EN/IT/CODE β€” Acquisition-Release Base (step_11900)

This repository is the canonical Small base release for the 1gpu-llm EN/IT/CODE 15B, tokenizer-48k, ctx2500 lineage: the benchmark-selected winner of the acquisition-release decay branch.

1gpu-llm is a family of language models trained from scratch on a single consumer GPU.

Reference training hardware: NVIDIA GeForce RTX 4060 Ti 16GB, single GPU.

Checkpoint identity:

  • family 1gpu-llm, tier small, role: decayed acquisition-release base
  • EN / IT / code, trained from scratch
  • GPT-2-style decoder, pre-layernorm, tied embeddings, learned absolute positions
  • vocab 48000, context 2500, dim 768, layers 12, heads 12
  • parameters 160,752,000 (~160.8M)
  • global step 11900, stored LR 2.389e-05
  • SHA256: d5ce343cb05b34c25e33502cd39fddf85de00bbd428d1e893127d645dbacc1ab

This is a base pretrained model, not instruction tuned.

Training Lineage

acquisition (WSD, peak LR 2e-4)
  -> best pre-collapse step_11000 (val_loss_mixed 3.3720)
  -> optimizer-preserving / scheduler-reset decay-only branch
     (1100 steps, 2e-4 -> 2e-5, inverse-proportional, ckpt every 100)
  -> dense checkpoints 11100...12100
  -> repo-native full benchmark (12 candidates)
  -> winner step_11900 (900 decay steps in, NOT the 12100 endpoint)

Parent acquisition checkpoint: step_11000.pt, SHA256 19c7bf81…fc4b, published separately as nazdef/1gpu-llm-small-en-it-code-acquisition-step11000 (revision 26e6479763e4612c7babbd571cd40b03d17b7527).

Selection Evidence

Full cohort (cuda/bf16/bs4/seed 1337, suite 20260919_pretrain_minimal_48k_en_it_code):

  • step_11900: mixed 4.9239, en 4.6377, it 3.9426, ppl 137.5
  • runner-up step_11700: mixed 4.9359 (+0.012)
  • vs source step_11000: mixed 5.1550 β†’ βˆ’4.5%

Shortlist CPU/FP32 confirmation (11700 / 11900 / 12100, same suite+seed):

  • step_11900: mixed 4.9193, en 4.6333, it 3.9477, ppl 136.9, loop 0.375, lc 0.950/0.775
  • step_11700: mixed 4.9375 (+0.018) β€” step_12100: mixed 4.9522 (+0.033)

11900 won mixed loss under both contracts and holds the best EN in both. Small honest trade-off: IT is ~0.03 better at 11700/12100. No decoding grid was needed.

Relation to the Non-Decayed Artifact

Both are intentionally retained:

  • step_11000 (acquisition repo): non-decayed, plastic, SFT-source candidate
  • step_11900 (this repo): decayed, cooled, ready-to-use base and future Continual parent

Approximate Exposure

Project accounting (approximate exposure, not unique coverage):

  • 240,000 sequence tokens / optimizer step
  • ~2.856B sequence tokens by step 11900

The public dataset is document-level; training used its deterministic 48k/ctx2500 packed derivative (manifest 1e8b33f3…, shuffle pcg64/1337). Frozen dataset revision: 333c4001551757aedb3454f30c8a6c08eaf23e12.

Files Included

  • step_11900.pt β€” original exact checkpoint (model + optimizer + scheduler + RNG + cursor 1142400)
  • step_11900.safetensors β€” checkpoint-native weights, bit-identical (max delta 0.0)
  • step_11900.safetensors.json β€” export manifest / provenance sidecar
  • exact tokenizer bundle (SHA256 8aef9589…24a08b)
  • source_decay_launch_config.yaml β€” decay config that produced this branch
  • release_selection.json β€” GPU/BF16 + CPU/FP32 selection numbers
  • release_manifest.json, SHA256SUMS

No standard-Transformers model.safetensors / config.json yet (see limitation).

Usage (repo-native)

import torch
from safetensors.torch import load_file

state = load_file("step_11900.safetensors")  # exact model tensors

Native architecture: repo GPT2DecoderLM (pre-LN, fused QKV); load with project code at github.com/nazarenodefrancesc/nanochat-llm-training.

Limitations

  • base pretrained model, not instruction tuned
  • native format is authoritative; standard Transformers conversion omitted because the GPT-2 mapping drops head.bias (max abs 1.327, mean abs 0.058)
  • benchmark is limited and does not prove factual reliability
  • EN/IT/code balance is not uniform (IT ~0.03 better at neighboring checkpoints)
  • downstream deployment requires separate evaluation

License / Data Provenance

No blanket license claim: the upstream licensing of the 15B EN/IT/CODE corpus mixture could not be established from the dataset release metadata (no license field on the dataset card). Same honest caveat as the acquisition repo. Users must verify upstream terms against the frozen revision above.

Summary

If you want the ready-to-use decayed Small base of this lineage β€” benchmark-selected under two contracts, cooled gradients, exact native weights β€” this is the one.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train nazdef/1gpu-llm-small-en-it-code-base