Width Vs Depth

We tested two different architecture configurations on 500M tokens of FineWeb to investigate whether depth or width is more effective for tiny language models (TLMs).

Model Architectures

Config A: Deep-Narrow (code name: depth_311)

  • Base Architecture: LlamaForCausalLM
  • Tokenizer: Harley-ml/Dillionv2-1.3M
  • Transformers Version: 5.13.1
  • Hidden Size: 128
  • Vocab Size: 2564
  • Number of Layers: 21
  • Number of Heads: 4
  • Number of KV Heads: 2
  • Intermediate Size: 313
  • Head Dim: 32
  • Max Position Embeddings: 256
  • RoPe Theta: 2125.0
  • Tie Word Embeddings: true
  • Hidden Activation: silu
  • MLP Bias: false
  • Initializer Range: 0.2
  • RMS Norm Eps: 1e-06
  • Pretraining Tp: 1
  • Use Cache: false
  • Hidden Size/Layers: 6.10
  • Total Parameters: 3,889,920

Config B: Shallow-Wide (code name: width_311)

  • Base Architecture: LlamaForCausalLM
  • Tokenizer: Harley-ml/Dillionv2-1.3M
  • Transformers Version: 5.13.1
  • Hidden Size: 192
  • Vocab Size: 2564
  • Number of Layers: 9
  • Number of Heads: 6
  • Number of KV Heads: 2
  • Intermediate Size: 469
  • Head Dim: 32
  • Max Position Embeddings: 256
  • RoPe Theta: 2125.0
  • Tie Word Embeddings: true
  • Hidden Activation: silu
  • MLP Bias: false
  • Initializer Range: 0.2
  • RMS Norm Eps: 1e-06
  • Pretraining Tp: 1
  • Use Cache: false
  • Hidden Size/Layers: 21.33
  • Total Parameters: 3,816,576

Training Setup

  • Epochs: 1
  • Max Steps: -1.0
  • Batch Size: 256
  • Sequence Length: 256
  • Gradient Accumulation: 2
  • Gradient Clipping: 1.0
  • Gradient Checkpointing: true
  • Learning Rate: 3e-3
  • Eval Split: 0.001
  • Weight Decay: 0.01
  • Optimizer: AdamW
  • AdamW Betas: (0.9, 0.95)
  • AdamW Eps: 1e-8
  • Scheduler: WSD
  • WSD Warmup Ratio: 0.015
  • WSD Stable Ratio: 0.78
  • WSD Decay Ratio: 0.20
  • WSD Minium LR Ratio: 0.0
  • WSD Number of Cycles: 0.5
  • DType: float16
  • Torch.Compile: false
  • DataLoader Workers: 4
  • Seed: 311

Results

Accuracy is normalized by length and shown as a percentage.

Config Final Val Loss ↓ Arc Easy ↑ Arc Challenge ↑ HellaSwag ↑ PiQA ↑ Swag ↑ Blimp ↑ Avg ↑
Config A 3.13697 29.17 21.67 27.01 54.03 32.65 67.66 38.70
Config B 3.14935 28.91 20.73 26.93 53.65 32.24 68.18 38.44

Config A scores higher than Config B on nearly every task, demonstrating that even at minuscule scales, greater depth can outperform greater width.

Model Checkpoints

The two models are stored separately in different folders in this repository. To load them, use:

from transformers import AutoModelForCausalLM, AutoTokenizer

config_a = AutoModelForCausalLM.from_pretrained(
    "fromziro/Width-Vs-Depth",
    subfolder="config_a",
)

# load config b instead:
# config_b = AutoModelForCausalLM.from_pretrained(
#    "fromziro/Width-Vs-Depth",
#    subfolder="config_b",
# )

tokenizer = AutoTokenizer.from_pretrained("fromziro/Width-Vs-Depth")

License

Apache 2.0.

Citation

@misc{width-vs-depth,
  title        = {Width-vs-Depth at Small Scales},
  organization = [FromZero],
  authors      = {Paul Courneya, Jonathon LY},
  year         = {2026},
  url          = {https://huggingface.co/fromziro/Width-Vs-Depth]
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train fromziro/Width-Vs-Depth