Squeal-Studio/drakon_28m-base
drakon_28m-base is a compact ~28M parameter language model pretrained from scratch on Russian-language text. It uses a hybrid architecture combining Depthwise Causal Conv, MinGRU, and Full Causal Attention with RoPE — every 3rd block is an attention block.
This is a base model (pretraining only). Instruction-tuned (SFT) version is in progress.
Research and Educational Model. Designed for experimentation and research. Given its size, performance on complex tasks will be limited.
Model Description
- Architecture: Drakon (Depthwise Conv + MinGRU + Causal Attention with RoPE)
- Parameters: ~28.5M
- Tokenizer: Custom BPE, vocab_size = 24,001
- Context length: 1024 tokens
Architecture Details
| Parameter | Value |
|---|---|
| hidden_size | 448 |
| num_layers | 12 |
| num_attention_heads | 8 |
| attention_every_n_layers | 3 |
| num_attention_layers | 4 |
| num_mingru_layers | 8 |
| conv_kernel_size | 4 |
| ffn_ratio | 2.0 |
| dropout | 0.1 |
| position_embedding | RoPE |
| normalization | RMSNorm (pre-norm) |
| block_pattern | [ConvMinGRU, ConvMinGRU, Attention] x 4 |
Training Details
- Dataset: ~5.5M documents, 10gb, ~1.5B unique tokens, ~3B tokens total (2 epochs)
- Preprocessing: Unicode normalization (NFKC), Cyrillic-ratio filtering, degenerate-repetition filtering, digit/punctuation-ratio filtering, Wikipedia, Fineweb2, cultura paragraph-level chunking, short-utterance grouping, exact deduplication (SHA-256), and approximate deduplication (MinHash/LSH)
- Sequence length: 1024 tokens
- Epochs: 2
- Batch size: 16 per device, gradient accumulation steps 8 (effective batch size 128)
- Learning rate: 4e-4 (cosine schedule, warmup 1500 steps)
- Weight decay: 0.01
- Max grad norm: 1.0
- Steps: 22,924
- Precision: fp16
- Seed: 42
Sources:
| Source | Documents | Characters |
|---|---|---|
| cultura_ru_edu | 649,700 | 1,194,216,253 |
| fineweb2_ru | 649,493 | 1,107,124,811 |
| wikipedia | 642,803 | 1,078,959,165 |
| taiga_proza | 549,206 | 967,503,171 |
| habr | 499,169 | 830,954,448 |
| ru_news | 329,573 | 552,286,936 |
| opensubtitles | 199,970 | 104,917,728 |
Evaluation
| Step | Eval Loss | Perplexity |
|---|---|---|
| 5,000 | 3.4508 | 31.5 |
| 10,000 | 3.2175 | 25.0 |
| 15,000 | 3.1293 | 22.9 |
| 22,924 | 3.0893 | 22.0 |
Training Curve
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Squeal-Studio/drakon_28m-base"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval().cuda()
inputs = tok("Однажды в лесу", return_tensors="pt").to("cuda")
inputs.pop("token_type_ids", None)
out = model.generate(
**inputs,
max_new_tokens=150,
do_sample=True,
temperature=0.8,
top_k=50,
repetition_penalty=1.1,
)
print(tok.decode(out[0], skip_special_tokens=True))
Generation Examples
Prompt: Однажды в лесу
Однажды в лесу случилось нечто, что называется: "Мы, друзья!", — сказал он.
В тот день все жители города были бездарно голодны. Они жили в деревне и не
знали где сейчас спрятаться от этих мыслей. 1995 год настал. На берегу реки
было тихо и тепло. Мы сидели на ветке возле дома для людей, где их не было.
Prompt: Президент заявил
Президент заявил, что у него есть все основания полагать, что на его долю
претендуют не только "опытные" россияне, но и люди с повышенной долей страха.
"Но не стоит забывать о том, что этот постулат – это в лучшем случае серьезный
вызов для тех, кто хочет получить помощь", — сказал он после публикации заявления.
Prompt: Наука доказала
Наука доказала, что наука об этом говорит, хотя и не имеет никакого отношения
к проблеме современной науки. Мы считаем, что вопрос по проблеме "многомерности"
был предметом обсуждения в ходе рабочей группы ученых, которая определила важность
того, как люди могут жить, и при этом сформулировала эту проблему.
Prompt: Рецепт борща очень простой
Рецепт борща очень простой – 2 ст. ложки растительного масла, порезанного в одном
масле. Смотрите рецепт салата и нарежьте его на кислые кусочки. Настаивать нужно
несколько часов. Добавить готовый блинчик на стол, в котором будут находиться овощи.
Scope & Limitations
- The architecture is still raw and prototype-level — expect rough edges
- Designed for architectural experiments and educational purposes
- Context length of 1024 tokens limits long-range coherence
- Not intended for production use
License
Apache 2.0
- Downloads last month
- 66
