BananaMind-IDK-3M
A 2,715,984-parameter small language model trained from scratch on 8B tokens of FineWeb-Edu, requested by @Banaxi-Tech and trained by @Compactbot.
Architecture
| Parameter | Value |
|---|---|
| Params (learnable) | 2,715,984 |
| Hidden dim (d) | 144 |
| Layers | 6 |
| Query heads | 2 |
| KV heads | 1 (GQA) |
| Head dim | 80 |
| FFN | SwiGLU, 3ร expansion |
| Position encoding | RoPE |
| Normalization | RMSNorm |
| Vocab size | 8,192 (BPE) |
| Tied embeddings | Yes |
Training
- Data: 8B tokens of FineWeb-Edu (cycled from a 3.6 GB local sample)
- Optimizer: AdamW, lr 5e-3, cosine decay with 200-step warmup
- Batch: 4 ร grad-accum 8 = effective 32, seq 2048
- Steps: 122,070 (reached the 8B-token target)
- Hardware: RTX 5090 (shared, batch reduced to fit ~2 GB free VRAM)
Evaluation (zero-shot, loglikelihood)
| Benchmark | Score |
|---|---|
| PIQA | 53.7% |
| ARC-Easy | 30.77% |
| ARC-Challenge | 16.47% |
| HellaSwag | 26.01% |
| ArithMark-3.0 | โ (not scored) |
These are honest numbers for a 2.7M-param model. For reference, chance is 50%/25%/25%/25% respectively.
Known limitations
This model produces degenerate greedy generation โ outputs collapse into repetition loops after the first sentence. It learned token-level statistics (grammatical first sentences) but not enough structure to sustain coherent multi-sentence generation. This is expected at 2.7M params even with 8B tokens of training data.
The model is published as a research artifact / data point, not as a usable generation model.
Files
final.ptโ full checkpoint (state_dict + optimizer + step), load withtorch.load('final.pt', map_location='cpu')['model']config.jsonโ architecture configtokenizer.jsonโ BPE tokenizer (8192 vocab, 5922 merges)tokenizer_config.jsonโ tokenizer settings
Usage
import torch
ckpt = torch.load('final.pt', map_location='cpu')
state_dict = ckpt['model'] # 128 tensors, 2,715,984 learnable params
The model uses a custom architecture (GQA attention + SwiGLU FFN + RoPE + RMSNorm). A loading script is not included; the state_dict keys follow the pattern:
model.embed_tokens.weight[8192, 144]model.layers.{i}.{attn|mlp|norm}.*for i in 0..5model.norm.weight[144]- (lm_head is tied to embed_tokens)
- Downloads last month
- 50