BEE-spoke-data
/

smol_llama-101M-GQA

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

pszemraj commited on Oct 26, 2023

Commit

f49a072

•

1 Parent(s): be538b7

Update README.md

Files changed (1) hide show

README.md +6 -1

README.md CHANGED Viewed

@@ -65,4 +65,9 @@ A small 101M param (total) decoder model. This is the first version of the model
 - 768 hidden size, 6 layers
 - GQA (24 heads, 8 key-value), context length 1024
-- train-from-scratch

 - 768 hidden size, 6 layers
 - GQA (24 heads, 8 key-value), context length 1024
+- train-from-scratch
+For the chat version of this model, please [see here](https://youtu.be/dQw4w9WgXcQ?si=3ePIqrY1dw94KMu4)
+---