The learning rate was not displayed as it should. (#3)

- The learning rate was not displayed as it should. (52b9f0d034e7c0c2627e41a35fc6b5d31b6cb8ca)

Co-authored-by: kapllan <kapllan@users.noreply.huggingface.co>

Files changed (1) hide show

README.md CHANGED Viewed

@@ -107,7 +107,7 @@ For further details see [Niklaus et al. 2023](https://arxiv.org/abs/2306.02069?u
 - batche size: 512 samples
 - Number of steps: 1M/500K for the base/large model
 - Warm-up steps for the first 5\% of the total training steps
-- Learning rate: (linearly increasing up to) $1e\!-\!4$
 - Word masking: increased 20/30\% masking rate for base/large models respectively
 ## Evaluation

 - batche size: 512 samples
 - Number of steps: 1M/500K for the base/large model
 - Warm-up steps for the first 5\% of the total training steps
+- Learning rate: (linearly increasing up to) 1e-4
 - Word masking: increased 20/30\% masking rate for base/large models respectively
 ## Evaluation