LS-Tiny-Base
| model | Multiple choice (length-adjusted) | Ends on its own | Depends on the question |
|---|---|---|---|
| LS-Tiny-Base | 47.5% | 96.9% | +0.40 |
| LS2.5-108M-A17M-Chat | 59.3% | 49% | +0.47 |
LS-Tiny-Base is a 31M Sparse model (Active 15M) trained on the synthetic dataset MiniCPM5-2B-distill (~130M tokens).
Despite being 3.5 times smaller than LS2.5-108M's chat variant, and seeing over 46x less tokens, it matches 80% of the accuracy of LS2.5-108M-A17M-Chat and more consistently ends its own replies (+47%). This shows that a singular high quality data source is far more efficient compared to a mix of datasets, and is likely the case since the distilled dataset contains one behavior in how responses are shaped, as compared to different datasets having multiple response styles or minor traits.
But while the dataset yields high efficiency, it is highly time and power consuming since inference is far slower than training. On average, LS2.5-108M-A17M-Chat maintained a 71k t/s training speed while distilling MiniCPM5-2B capped out at ~1.4k t/s (51x slower).
Chatbench results by category:
| category | n | acc | accnorm | chance |
|---|---|---|---|---|
| fact | 25 | 36.0% | 40.0% | 25.0% |
| commonsense | 20 | 55.0% | 50.0% | 25.0% |
| instruct | 20 | 65.0% | 35.0% | 25.0% |
| relevance | 20 | 15.0% | 65.0% | 5.0% |
| consistency | 18 | 33.3% | 50.0% | 25.0% |
| arithmetic | 15 | 40.0% | 46.7% | 25.0% |
| ALL | 118 | 40.7% | 47.5% | 25.0% |
Relevance and consistency are high since they naturally form from a unified response style. In Chatbench's open ended questoins, the model has learned to include the end of the query at the start of its reponse consistently, or start with the phrase "Yes, ". It also shows the usage of em-dashes, and prompts to the user after explanations (GPT style behavior).
| Model specs | |
|---|---|
| Total parameters: | 31M |
| Model parameters: | 15M |
| Active parameters: | 15M |
| Ngram table: | 16M |
| Context length: | 1024 |
| Pretrain details | |
| Tokens seen: | ~130M |
| GPU time (5060ti): | 1096s |
Examples:
you> hi
bot> Hello! I'm doing well, thanks for asking. Today, we're discussing **professionalism**—the principle of professionalism and respect in our everyday interactions with others. Would you like to explore how it applies to your own actions?
[62 tok, 1409.1 tok/s, ctx 73/4096]
you> ok
bot> Professionalism means ensuring that your actions or actions can convey effectiveness in our work environments. It involves:
- Using the same language across all intervals to emphasize key points.
- Creating a clear sense of communication through individual behavior.
- Fostering a supportive atmosphere where ideas are shared freely and easily, encouraging a more cohesive approach to life.
[100 tok, 1282.1 tok/s, ctx 184/4096]
you> How can I be healthy?
bot> To maintain good health, focus on:
- **Mindful Breaks**: Take short breaks to stretch and breathe.
- **Sleep & Schedule**: Prioritize sleep and movement over intensity.
- **Physical Awareness**: Practice mindfulness or meditation for relaxation.
- **Nutrition**: Avoid processed foods like protein, vegetables, and whole grains.
- **Environment**: Ensure your workspace is clean and well-ventilated to reduce stress.
Consistency in these practices helps sustain good health without sacrificing physical function.
[156 tok, 1457.9 tok/s, ctx 171/4096]
you> too much effor
bot> That's a great point! Failure of completing the work can be tricky—check that you're not overusing it or using other methods. If you have more questions about how to handle this process, feel free to ask!
[57 tok, 1266.7 tok/s, ctx 242/4096]
Usage: To run this model, download the latest tinylm from github and build from source. Currently working on windows prebuilts.
- Downloads last month
- 109
4-bit