JingzeShi
/

Doge-60M

@@ -16,7 +16,6 @@ Doge is an ongoing research project where we aim to train a series of small lang
 In addition, Doge uses Dynamic Mask Attention as sequence transformation and can use Multi-Layer Perceptron or Cross Domain Mixture of Experts as state transformation. Dynamic Mask Attention allows the Transformer to use self-attention during training and state space during inference, and Cross Domain Mixture of Experts can directly inherit the weights of Multi-Layer Perceptron for further training. This model is trained by Jingze Shi, it only allows text input and text generation, for detailed algorithm and model architecture, please refer to [Wonderful Matrices](https://arxiv.org/abs/2412.11834), the ongoing research repository is [Wonderful Matrices](https://github.com/LoserCheems/WonderfulMatrices).
 ## Uses
 ```python
@@ -37,18 +36,24 @@ In addition, Doge uses Dynamic Mask Attention as sequence transformation and can
 > TODO: The larger model is under training and will be uploaded soon.
-|| Training Data | Epochs | Steps | Content Length | Tokens | LR | Batch Size | Precision |
 |---|---|---|---|---|---|---|---|---|
-| [Doge-20M](https://huggingface.co/LoserCheems/Doge-20M) | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 2 | 10k | 2048 | 5B | 8e-4 | 0.25M | bfloat16 |
-| [Doge-60M](https://huggingface.co/LoserCheems/Doge-60M) | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 2 | 20k | 2048 | 20B | 6e-4 | 0.5M | bfloat16 |
-**Training Environment**:
 - Image: nvcr.io/nvidia/pytorch:24.10-py3
 - Hardware: 1x NVIDIA RTX 4090
 - Software: Transformers
 ## Citation
 ```bibtex

 In addition, Doge uses Dynamic Mask Attention as sequence transformation and can use Multi-Layer Perceptron or Cross Domain Mixture of Experts as state transformation. Dynamic Mask Attention allows the Transformer to use self-attention during training and state space during inference, and Cross Domain Mixture of Experts can directly inherit the weights of Multi-Layer Perceptron for further training. This model is trained by Jingze Shi, it only allows text input and text generation, for detailed algorithm and model architecture, please refer to [Wonderful Matrices](https://arxiv.org/abs/2412.11834), the ongoing research repository is [Wonderful Matrices](https://github.com/LoserCheems/WonderfulMatrices).
 ## Uses
 ```python
 > TODO: The larger model is under training and will be uploaded soon.
+**Training**:
+| Model | Training Data | Epochs | Steps | Content Length | Tokens | LR | Batch Size | Precision |
 |---|---|---|---|---|---|---|---|---|
+| [Doge-20M](https://huggingface.co/JingzeShi/Doge-20M) | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 2 | 10k | 2048 | 5B | 8e-4 | 0.25M | bfloat16 |
+| [Doge-60M](https://huggingface.co/JingzeShi/Doge-60M) | [HuggingFaceTB/smollm-corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 2 | 20k | 2048 | 20B | 6e-4 | 0.5M | bfloat16 |
+**Evaluation**:
+| Model | TriviaQA | MMLU | ARC | PIQA | HellaSwag | OBQA | Winogrande |
+|---|---|---|---|---|---|---|---|
+| [Doge-20M](https://huggingface.co/JingzeShi/Doge-20M) | - | 26.01 | 36.15 | 56.26 | 26.60 | 26.60 | 50.12 |
+| [Doge-60M](https://huggingface.co/JingzeShi/Doge-60M) | - | 25.81 | 45.49 | 61.37 | 29.65 | 27.40 | 52.57 |
+**Environment**:
 - Image: nvcr.io/nvidia/pytorch:24.10-py3
 - Hardware: 1x NVIDIA RTX 4090
 - Software: Transformers
 ## Citation
 ```bibtex