Update README.md
Browse files
README.md
CHANGED
@@ -65,4 +65,9 @@ A small 101M param (total) decoder model. This is the first version of the model
|
|
65 |
|
66 |
- 768 hidden size, 6 layers
|
67 |
- GQA (24 heads, 8 key-value), context length 1024
|
68 |
-
- train-from-scratch
|
|
|
|
|
|
|
|
|
|
|
|
65 |
|
66 |
- 768 hidden size, 6 layers
|
67 |
- GQA (24 heads, 8 key-value), context length 1024
|
68 |
+
- train-from-scratch
|
69 |
+
|
70 |
+
For the chat version of this model, please [see here](https://youtu.be/dQw4w9WgXcQ?si=3ePIqrY1dw94KMu4)
|
71 |
+
|
72 |
+
---
|
73 |
+
|