Kite

🎉 You are looking at Kite 8.1, which is larger and uses ClimbMix!

Kite is a small, trained, 11 million parameter language model.

Training

It was trained on the first shard of Andrej Karpathy's shuffle of nvidia/Nemotron-ClimbMix, using 1 epoch, 12 batch size, 0.001 learning rate, and the pika 5 tokenizer.

Limitations

Due to its size, the model is not suitable for production workloads.

Downloads last month
377
Safetensors
Model size
11M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train qikp/kite-8.1-11m-base