GPT-2 Nepali WikiMulti

GPT-2-small model pretrained from scratch on the Nepali portion of WikiMulti.

Dataset

  • Dataset: dehanalkautsar/WikiMulti
  • File: 20260801_ne_wiki.parquet
  • Training field: text

Architecture

Standard GPT-2-small architecture:

  • Layers: 12
  • Hidden size: 768
  • Attention heads: 12
  • Maximum context: 1024
  • Vocabulary size: 43846

All GPT-2 weights were initialized from scratch.

No pretrained GPT-2 model weights were used.

Tokenizer

A separate Byte-Level BPE tokenizer was trained from scratch using only the Nepali training split.

The original English GPT-2 vocabulary and BPE merge rules were not used.

Vocabulary size: 43846

Training

  • GPU: NVIDIA RTX A6000
  • Precision: FP16
  • TF32: True
  • Maximum epochs: 50
  • Batch size per GPU: 8
  • Gradient accumulation: 4
  • Sequence length: 1024
  • Learning rate: 5e-05
  • Scheduler: linear
  • Validation frequency: every 500 optimizer steps
  • Early stopping patience: 999
  • Early stopping threshold: 0.0
  • Random seed: 42

Training was stopped after validation loss failed to improve for 999 consecutive evaluation rounds, unless the maximum epoch ceiling was reached first.

Best Model

Best validation loss:

1.5498346090316772

Final best-model validation loss:

1.5498346090316772

Perplexity:

4.71069101241065

Downloads last month
964
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support