Gome

Custom experimental language model.

Training

  • max_steps: 19518
  • seq_len: 512
  • global_batch_size: 32
  • lr: 0.0001
  • tokenizer: gpt2
  • dataset: HuggingFaceFW/fineweb-edu

Notes

This is a custom architecture and requires custom model code to load for inference. use the run_gome.py to run the model after you download everyting this is my first model which has 24M params but it wasnt able to finish training because it started to take longer and longer but after some testing it genuinely generate text, not in the most coherent way but it works, and I really how you enjoy it, please and thank you.

I had to fix it a little so now the whole brain of Gome is in that one model.safetensors file

Downloads last month
7
Safetensors
Model size
24.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support