crumb
/

gpt2023

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

crumb commited on Apr 30, 2023

Commit

4b77f6b

•

1 Parent(s): 421585d

Update README.md

Files changed (1) hide show

README.md +2 -0

README.md CHANGED Viewed

@@ -8,6 +8,8 @@ language:
 This is the smallest GPT-2 model (124m) from OpenAi finetuned on approximately 2.23B tokens (almost the 2.48B needed to 'chinchilla-optimally' pretrain it!) consisting of 1.3B from common crawl sites from 2023, 540M from ArXiv, and 390M from GitHub.
 *(from GPT-2 model card)*
 ### Model description

 This is the smallest GPT-2 model (124m) from OpenAi finetuned on approximately 2.23B tokens (almost the 2.48B needed to 'chinchilla-optimally' pretrain it!) consisting of 1.3B from common crawl sites from 2023, 540M from ArXiv, and 390M from GitHub.
+The model was trained with a learning rate of 1e-4, with a warmup of 1024 steps, then decaying to 0. There were 4000 total steps during training at a batch size of 512 examples with a context length of 1024. The batch size and context length are the same as the pre-training of GPT2 itself.
 *(from GPT-2 model card)*
 ### Model description