Instructions to use nowordsxiaomu/DeepSeek-Flash-Mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nowordsxiaomu/DeepSeek-Flash-Mini with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nowordsxiaomu/DeepSeek-Flash-Mini", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DeepSeek-Flash-Mini
A from-scratch, lightweight Mixture-of-Experts language model that reproduces the core recipe of DeepSeek-V3 / V4-style efficient transformers at toy scale. It is fully trainable and runnable on a laptop (Apple M-series MPS supported).
β οΈ This is a 15M-parameter research toy trained on ~5.3M tokens of TinyStories. It tells small coherent stories. It is not a general assistant and will score near random on knowledge/QA benchmarks (see MMLU note below). The value here is the architecture, not the raw score.
Architecture
| Component | What it does |
|---|---|
| MLA (Multi-head Latent Attention) | Low-rank KV compression β only a kv_lora_rank-dim latent is cached, so KV memory is ~6.4Γ smaller than a same-size MHA. Two mathematically-equivalent paths: naive (explicit K/V restore, for training) and absorb (compute attention in latent space, for long-context decode). |
| DeepSeekMoE | Fine-grained routed experts + shared experts. Every token also passes a shared expert; the first n_dense_layers stay dense FFN. |
| Aux-loss-free load balancing | Each expert carries a gradient-free bias b_i; routing uses argtopk(s_i + b_i) while aggregation still uses the true s_i. Biases self-adjust every step toward balanced load β no auxiliary loss term needed. |
| MTP (Multi-Token Prediction) | A lightweight head predicts t_{i+2} sharing the embedding/output. At inference it acts as a draft model for self-speculative decoding (distribution-identical to autoregressive decoding), giving ~1.37Γ speedup. |
Training
- Config:
nanoβ dim 256, 6 layers, 8 routed experts (2 active) + 1 shared,kv_lora_rank64.- Total params 14.92M, activated 6.90M (46%).
- Data: TinyStories (22.5M chars β 5.3M tokens), custom BPE vocab 8192.
- Recipe: 3500 steps, cosine schedule + warmup, bf16, AdamW, gradient clip, MPS.
- Time: ~32 minutes on Apple M-series (MPS).
Evaluation β where it actually ranks
The meaningful benchmark for this model is TinyStories validation perplexity (lower = better), compared with other small models trained on the same corpus:
| Model | Params | Train tokens | Val PPL (TinyStories) |
|---|---|---|---|
| karpathy/stories15M | 15.2M | ~1B | 2.92 |
| child-12m | 12.3M | ~1B | 3.45 |
| pluto-15M | 15M | ~1B | 3.64 |
| TinyStories-28M | 28M | ~1B | ~3.0β3.5 |
| DeepSeek-Flash-Mini (nano) | 14.9M | 5.3M | 7.67 |
| microgpt | ~? | 327M | 9.49 |
Reading the table: our 7.67 sits in the middle. The gap to the top models is dominated by training-token count (5.3M vs ~1B), not architecture β MLA/MoE/MTP here are an engineering demonstration. (Cross-tokenizer perplexities are not strictly comparable; we use a custom BPE, so a fully fair comparison would report bits-per-byte. The table is for rough orientation.)
Efficiency metrics (measured)
| Metric | Value |
|---|---|
| KV cache vs same-size MHA | 6.4Γ smaller (960 B vs 6144 B per token) |
| MTP speculative decoding | 1.37Γ speedup (92% draft acceptance) |
| MPS decode throughput | ~40 tok/s (naive) |
MMLU (subset, measured)
Loglikelihood multiple-choice (mean-NLL argmin), CPU, 10 subjects / 1,432 questions:
- Overall accuracy: 23.9% (random-chance baseline 25%)
- Per-subject range: 17% (computer_security) β 31% (high_school_mathematics, abstract_algebra)
Interpretation: essentially at the random floor. Expected β the model was trained on 5.3M
tokens of TinyStories and holds no world knowledge. The small deviations (e.g. math 31%) are
statistical noise, not competence. Reported for honesty, not as a ranking claim. See
mmlu_result.json for the per-subject breakdown. The honest "ranking" for this model is the
TinyStories perplexity table above.
How to use
pip install torch safetensors
# (this repo already bundles config.py / model/ / dataio/ / generate.py)
from load_and_generate import load_model
model, cfg = load_model(".") # needs config.json + model.safetensors in repo dir
# or CLI
python load_and_generate.py --prompt "Once upon a time" --max-new-tokens 80
python load_and_generate.py --prompt "Once upon a time" --spec # MTP speculative decoding
Loads the safetensors weights into the bundled model code and generates TinyStories-style text.
Files
model.safetensorsβ nano weights (68 MB)config.jsonβ architecture hyperparameters (ModelConfigschema)tokenizer.jsonβ custom BPE tokenizer (vocab 8192)config.py,model/,dataio/,generate.pyβ self-contained inference codeload_and_generate.pyβ convenience loader + CLI
Limitations
- Tiny context (512 tokens), English TinyStories only, no instruction-tuning.
- Not a chat/QA model; do not expect factual answers.
- The three presets (
nano/small/base) are defined inconfig.py; onlynanois trained and shipped here.
License
MIT β Β© 2026 nowordsxiaomu.
- Downloads last month
- -
Evaluation results
- validation perplexity on TinyStoriesself-reported7.670