File size: 412 Bytes
b45f55a 1b59448 b45f55a e1c87fb 1b59448 0132ff0 a01b372 |
1 2 3 4 5 6 7 8 9 10 11 |
---
license: apache-2.0
datasets:
- HuggingFaceTB/cosmopedia
---
An untrained precursor MoE created from Cosmo using mergekit.
Gate routing initialized using prompt hidden state method. Five are based on the visualized topic clusters of Cosmopedia data, three are task-oriented.
Degenerate layers were 0, 1, and 2. Expert gates for layers 0, 1, and 2 have been randomly initialized to with luck mitigate this.
|