Text Generation
Safetensors
English
qwen3
from-scratch
small-language-model
conversational

Training Loss?

#1
by qikp - opened

Can you provide the Training Loss? I'm curious if the model is underfitting or not.

Its an 800k param model, its never not gonna underfit...

Can you provide the Training Loss? I'm curious if the model is underfitting or not.

I don't exactly remember, but it was something like ~2-3.5

AxionLab-official changed discussion status to closed

Sign up or log in to comment