TinyChat-5M

A 4.3M parameter transformer trained from scratch on the TinyChat dataset.

Architecture

Parameter Value
Parameters 4,327,872
Layers 6
Hidden dim 192
Attention heads 3
Head dim 64
FFN mult 4x (SwiGLU)
Context 512
Vocab 4,096 (BPE)
Norm RMSNorm
Positional RoPE
Embeddings Tied (input = output)

Training

  • Dataset: starhopp3r/TinyChat (1M rows, ~740M tokens)
  • Tokens seen: 1.64B (90% train / 10% val split, multi-epoch)
  • Optimizer: AdamW, LR 1e-3, cosine decay + warmup
  • Batch: 32 × 512 tokens
  • Steps: 100,000 (completed cleanly)
  • Hardware: RTX 5090 (32 GB)
  • Time: ~46 minutes

Results

Metric Value
Final val loss (step 100k) 1.7603
Final val perplexity 5.81
Best val loss (step 96k) 1.7531
Best val perplexity 5.77

Generation Samples

Prompt: Hello, how are you today? Output: I feel quite disgusted by the mess in our office today. Yes, it is necessary to have a better work space for everyone. The clutter makes everyone a bit unwell and can affect their mood. It really does, especially when we just want things to be clean again. A tidy environment helps us enjoy our work more clearly each day.

Prompt: Once upon a time Output: when the thread interrupts our plans and makes everything feel more tense. It is so annoying how simple things can turn into chaos, isn't it? Yes, it really is, like trying to find noise that won't stop bothering us. I wish we could just enjoy the sewing process without all these problems.

Prompt: The cat sat on the Output: bench, waiting for lunch to arrive. That is surprising and cute! I hope it unfolds before they try to play outside. Yes, it will be a delightful surprise, but we must stay calm as her. Do you remember that day when the weather was so nice?

Prompt: I think that Output: is a good idea because sometimes sadness can feel so heavy and hard to handle. Yes, it feels like emotions are holding us down like a thin needle poking at our hearts. Do you believe sharing feelings helps lighten the load we carry in some way?

The model produces coherent conversational English appropriate for its size. It does not answer factual questions correctly (expected at 4.3M params) but maintains consistent tone and topic.

Notes

  • v4 fixes v3's NaN divergence (LR reduced from 2e-3 to 1e-3, NaN detection added).
  • Training completed all 100k steps without divergence.
  • Model is for research/educational purposes. Not suitable for production use.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support