Gome
Custom experimental language model.
Training
- max_steps: 19518
- seq_len: 512
- global_batch_size: 32
- lr: 0.0001
- tokenizer: gpt2
- dataset: HuggingFaceFW/fineweb-edu
Notes
This is a custom architecture and requires custom model code to load for inference. use the run_gome.py to run the model after you download everyting this is my first model which has 24M params but it wasnt able to finish training because it started to take longer and longer but after some testing it genuinely generate text, not in the most coherent way but it works, and I really how you enjoy it, please and thank you.
I had to fix it a little so now the whole brain of Gome is in that one model.safetensors file
- Downloads last month
- 7