LS-92M-A22M

task 63M / 16M act 92M / 22M act 220M / 25M act chance
arc_easy 31.2% 35.0% 35.8% 25.1%
hellaswag 27.3% 28.5% 31.8% 25.0%
piqa 56.5% 60.0% 60.8% 50.0%
lambada 13.2% 18.5% 19.2% 0%
mmlu 24.8% 23.8% 24.5% 25.0%

Miniature mixture of experts model with top-1 routing. Trained entirely on a 1660 super.

This model is highly experimental and may not perform up to standards.

4 layers were duplicated from the base model LS-63M-A16M and a high LR was applied to them. This checkpoint marks the second epoch of training complete, around another 2B tokens over 38 GPU hours. Total 4B tokens over 70 hours.

Total parameters: 92M Active parameters: 22M context length: 4096, sliding window 1024

Training end evals: NA

Chat:

you> hi
bot> Hello! How can I assist you?

      [12 tok, 187.5 tok/s, ctx 22/16384]

you> what is the capital of france?
bot> The capital of France is Louis XVI.
      [13 tok, 171.1 tok/s, ctx 53/16384]

Code:

you> write a python function that reverses a string
bot> Here's an example of a Python function that reverses a string:

```python
def reverse_string(s):
    return s[::-1]
```

To test this function, you can call it with a test case like this:

```python
print(reverse_string("Hello world"))
```

This will print the reversed string in the order of slicing.

Note: It is important to note that reverse() is a built-in function in Python, so if you have a string that you want to reverse, you should replace `"Hello"` with the string you want to reverse.

      [138 tok, 170.8 tok/s, ctx 155/16384]

(as seen, Code aspect is a major improvement over the base 63M A16M model.)

To try it yourself:

Download tinylm.exe and LS-92M-A22M-q8.bin (placed in \models), run command tinylm chat LS-92M-A22M-q8 2048

Note: the bundled tinylm.exe is likely outdated. For the latest version, check here for the source code of tinylm and windows prebuilts.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dsg2/LS-92M-A22M

Base model

Dsg2/LS-63M-A16M
Finetuned
(1)
this model
Quantizations
1 model

Collection including Dsg2/LS-92M-A22M