TinyStories 8.3M Foundation Storyteller

A high-performance 8.3M parameter modern decoder-only transformer trained from scratch on the complete 2.1 Million TinyStories dataset across a dedicated Tesla T4 GPU.

Model Specifications

  • Parameters: 7,155,360 (~8.33M)
  • Layers: 6
  • Embedding Dimension: 288
  • Attention Heads: 8
  • FFN Hidden Dimension: 768
  • Context Length: 256 tokens
  • Vocabulary: 4,096 tokens (Byte-Level BPE)
  • Architecture: RoPE, Float32-Precision StableRMSNorm, SwiGLU, SDPA (Flash Attention), Tied Embeddings

Status

  • Progress: Step 5,000 / 33,120
  • Validation Loss: 1.7634
  • Validation Perplexity: 5.83

Sample Output

Prompt: Once upon a time, there was a little girl named Mia who loved stars.
Generation: Once upon a time, there was a little girl named Mia who loved stars. The hand m snowman pow and theJo loo jourver the trees. wait, the bird shiningâ that together itins. but m a re, yellow arri! The arri said, I kick,ppildfore toerfter op kne comingec

's bird mri carefully and place, I pie, wow!itedTh't need toerfterver theauseideurn, the bird dry away as bra as he sp.

Howuffy, a shapes replro areOn

Author

Trained by Zyroxx66.

Downloads last month
87
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Zyroxx66/tinystories-8m-storyteller