İvme-Conversate-P-v1

İvme (Turkish: acceleration) is a series of stupidly small language models built to punch above their weight. P-v1 is the smallest one yet: a 3.23M parameter decoder-only base model, about 13% the size of İvme-Conversate-v3-Base, trained entirely by logit distillation from v3.

The question behind it was simple: if you take v3's tokenizer and architecture, shrink the model to almost nothing, and train it to match v3's full output distribution on every token, how much of v3 survives? The short answer: most of its grammar, a good chunk of its language modelling, and very little of its (already thin) world knowledge.


Model Details

Parameter Value
Architecture Decoder-only transformer, dense, same code as v3
Parameters 3,228,928
Layers 6
Hidden dim 128
FFN dim SwiGLU (341)
Attention heads 4, full attention (no GQA), head_dim 32
Attention variant QK-Norm + XSA (Exclusive Self Attention, arXiv:2603.09078)
Context length 1024 tokens
Vocab size 16,000 (v3's digit-atomic BPE, unchanged)
Positional encoding RoPE (θ=10,000)
Normalization RMSNorm (pre-norm)
Embeddings Tied input/output
Biases None

Nothing about the architecture is new. P-v1 is v3's exact model code (RoPE, QK-Norm, XSA, SwiGLU, RMSNorm) with the shape cut from 16 × 320 down to 6 × 128. Keeping the code identical was deliberate: the only thing that changes between teacher and student is size, so whatever the student loses is down to capacity, not a different design.

At this size, the embedding table dominates. With a 16k vocabulary and tied embeddings, 2.05M of the 3.23M parameters (63%) are the token lookup table, leaving roughly 1.18M for the actual transformer. P-v1 is, quite literally, mostly vocabulary.


Distillation

Because P-v1 shares v3's tokenizer exactly, distillation is direct: no vocabulary alignment, no projection, just token-by-token matching of the full 16,000-way next-token distribution.

loss = 0.9 · KL(p_teacher ‖ p_student)  +  0.1 · CE(next token)
  • Forward KL (teacher ‖ student), temperature 1.0. Forward KL is mass-covering: the student is pushed to put probability everywhere the teacher does, not just on the teacher's top choice.
  • Online teacher. v3 runs in bf16 inside the training loop. At 24.8M parameters it's cheap enough that caching logits would have cost far more disk than it saved compute.
  • A small hard-label term (10% CE on the real next token) keeps the student anchored to the data, not only to the teacher.

The useful thing about a hard target is that it says "the next token was X." The useful thing about the teacher's distribution is that it also says "Y and Z were plausible, and these thousands of tokens weren't." For a model with ~1.2M non-embedding parameters, that's a much easier signal to learn from.


Benchmarks

All results zero-shot with lm-evaluation-harness. v3 numbers are from its model card, measured the same way.

Benchmark P-v1 (3.2M) v3-Base (24.8M) Share of v3's above-chance score
WikiText-2 (byte perplexity) ↓ 2.6116 2.1362 —
BLiMP (macro-average) ↑ 70.33% 78.49% 71%
ARC-Easy (acc) ↑ 35.35% 43.60% 56%
ARC-Easy (acc_norm) ↑ 32.03% 39.02% 50%
ARC-Challenge (acc_norm) ↑ 21.16% 24.06% below chance
HellaSwag (acc_norm) ↑ 26.23% 28.82% 32%
PIQA (acc_norm) ↑ 54.57% 57.78% 59%

The last column measures how much of v3's lead over random guessing P-v1 keeps: (student − chance) / (teacher − chance).

BLiMP is the standout. P-v1 keeps about 71% of v3's above-chance grammaticality score at 13% of the size. Grammar is local, rule-like structure, exactly what a very small model can still represent, and exactly what the teacher's soft distribution encodes well (it spreads probability across grammatical alternatives, not just the single observed token).

Language modelling holds up reasonably. WikiText-2 byte perplexity goes from 2.136 to 2.612, about 1.10 → 1.39 bits per byte. A real cost, but a modest one for an 8× smaller model.

Knowledge and commonsense benchmarks mostly don't transfer, and that's expected. v3 itself sits barely above chance on HellaSwag (28.8% vs 25%) and ARC-Challenge (24.1%, effectively chance). A student can't inherit a capability the teacher barely has. ARC-Challenge acc_norm landing slightly below chance (21.2%) is common for very small models on that metric and isn't a sign of a broken model. ARC-Easy and PIQA keep a little over half of v3's margin.


Sample output

Temperature 0.8, top_k 40, seed 42. This is a base model, not instruction-tuned: it continues text rather than answering like a chat assistant.

Prompt: "To bake a great cake you should always"

To bake a great cake you should always be careful to make it very good for the family and to make them look more attractive. While you should eat a little more and drink a lot of water for the next several days, it is always advisable to drink a little more. You should also be vigilant. If you should be a good drinker, the best way to make a good match is to pick at least one meal to make a little more.

Expect more repetition and topic drift than v3 at the same settings. That's the price of 3M parameters.


Training

Data

~2.0B tokens from HuggingFaceFW/fineweb-edu (sample-10BT), streamed, tokenized with v3's tokenizer and packed into 1024-token sequences. That's roughly 600 tokens per parameter, far past compute-optimal on purpose, same philosophy as the rest of the İvme series: the model is tiny, so overtraining it is cheap.

The student only ever sees fineweb-edu text, but every token's target comes from v3, which was trained on the much broader v3 mix (fineweb-edu, DCLM, FineWiki, FineMath, Cosmopedia, SimpleStories). Some of that breadth reaches P-v1 through the teacher's predictions.

Hyperparameters

Setting Value
Loss 0.9 · forward KL (T=1.0) + 0.1 · CE
Optimizer AdamW (β = 0.9, 0.95)
Learning rate 3e-3
LR schedule Warmup-Stable-Decay (WSD)
Warmup steps 1,000
Decay fraction 10% of training
Weight decay 0.1 (none on embeddings and norms)
Gradient clipping 1.0
Batch 128 × 1024 tokens (131,072 tokens / step)
Steps 15,258
Precision bfloat16 autocast, fp32 loss
Compilation torch.compile (student and teacher)

Hardware

Trained on a single NVIDIA RTX PRO 6000 Blackwell (96GB) at ~374k tokens/s, about 1.5 hours end to end, teacher included.

Reproducing

Everything used to train and evaluate this model is in the repo:

File What it is
training/distill.py The full training script (data streaming, online teacher, KD loss, WSD schedule)
training/launch_distill.sh Exact command for this model
training/launch_control.sh The CE-only control: same model, same data, no teacher
eval.sh Reproduces the benchmark table
inference.py Minimal generation example
eval/lm_eval_results.json Raw lm-evaluation-harness output for the numbers above
cd training && ./launch_distill.sh

Inference

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "IvmeLabs/Ivme-Conversate-P-v1", trust_remote_code=True, dtype=torch.float32,
)
tokenizer = AutoTokenizer.from_pretrained("IvmeLabs/Ivme-Conversate-P-v1", trust_remote_code=True)
model.eval()

inputs = tokenizer("Once upon a time, there was a", return_tensors="pt")
out = model.generate(
    **inputs, max_new_tokens=200, do_sample=True,
    temperature=0.8, top_k=40, pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(out[0], skip_special_tokens=True))

trust_remote_code=True is required (custom architecture: RoPE + QK-Norm + XSA + SwiGLU + RMSNorm dense decoder, shared with v3).

At 3.23M parameters, P-v1 runs comfortably on CPU.


Limitations

  • Base model only, not instruction-tuned, will not follow instructions or answer questions
  • English only
  • 1024 token context window
  • Inherits v3's weaknesses and adds its own: weaker long-distance syntax, more repetition, faster topic drift
  • World knowledge and commonsense reasoning (ARC, HellaSwag, PIQA) are at or near chance; don't rely on it for facts
  • Arithmetic: like v3, free-form generated arithmetic is unreliable. Don't use it for computation
  • Most of the parameter budget is the embedding table, so there is very little room for anything beyond local language structure

What's Next

The CE-only control run (training/launch_control.sh), to measure exactly how much of P-v1's performance comes from distillation versus simply training a model this size on the same tokens. After that: shorter token budgets, where distillation's advantage is usually largest, and a factorized embedding to shrink the 16k-vocab lookup table and give the transformer body a bigger share of the budget.

You can check our other models on our organization card!


Citation

@misc{ivme-conversate-p-v1,
  author       = {IvmeLabs},
  title        = {İvme-Conversate-P-v1},
  year         = {2026},
  publisher    = {Hugging Face},
  url          = {https://huggingface.co/IvmeLabs/Ivme-Conversate-P-v1}
}

Built by IvmeLabs. Small models, deliberate choices.

Downloads last month
-
Safetensors
Model size
3.23M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IvmeLabs/Ivme-Conversate-P-v1

Finetuned
(1)
this model

Dataset used to train IvmeLabs/Ivme-Conversate-P-v1

Paper for IvmeLabs/Ivme-Conversate-P-v1