File size: 412 Bytes
b45f55a
 
1b59448
 
b45f55a
e1c87fb
1b59448
0132ff0
 
a01b372
1
2
3
4
5
6
7
8
9
10
11
---
license: apache-2.0
datasets:
- HuggingFaceTB/cosmopedia
---
An untrained precursor MoE created from Cosmo using mergekit.

Gate routing initialized using prompt hidden state method. Five are based on the visualized topic clusters of Cosmopedia data, three are task-oriented.

Degenerate layers were 0, 1, and 2. Expert gates for layers 0, 1, and 2 have been randomly initialized to with luck mitigate this.