Text Generation
Transformers
Safetensors
English
llama
text-generation-inference

Dumb-1.2-RC1

new SoTA dumb

This model was trained with a high-quality data set of 34.6M parameters.

  • training token:

What this model can do

  • text generation
  • Predict the following words
  • Explain simple questions in an interesting way

What this model can't do

  • calculate arithmetic
  • math
  • Fluent conversation
  • Don't lie as much as possible

benchmark

vs 1-20M param

evals our model grint 1.3(1M) michel-nano(5.9M)
WikiText bytePPL 2.838 3.06 3.2461
arc-easy 34.13% 29.0% 33.38%

vs 20M~ param

1: (XX%) is Calculated with our ai / better ai what percentage of the score is for what is better.

evals our model supra-50M-base
WikiText bytePPL 2.838(95%) 2.7
arc-easy 34.13%(76%) 45.2%
BLiMP 64.87%(96%) 67.4%

If you look at these two tables, BLiMP and PPL are catching up with the competition, but ARC-Easy is not very powerful. In addition, it is 14% higher in wikitext than michel-nano (22M param). However, this model is still a preview version, detailed benchmarks have not been carried out, and the official version may be more powerful. In addition, the low score of ARC-Easy indicates that you don't know much about scientific knowledge, and you can see that you should train using scientific data sets.

In addition, it is completely inferior to the more advanced Supra model (some items are comparable, but not completely catching up)

Downloads last month
951
Safetensors
Model size
34.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 56m/Dumb-1.2-RC1

Finetuned
(1)
this model
Quantizations
1 model

Datasets used to train 56m/Dumb-1.2-RC1

Spaces using 56m/Dumb-1.2-RC1 2