Klyra-64M-Instruct

Klyra-64M-Instruct is the first instruction-tuned checkpoint in the Klyra-64M family.

It starts from Klyra-64M-Midtrain and was fine-tuned on filtered HuggingFaceTB/smol-smoltalk data using assistant-only loss.

SFT

  • Train examples: 226,049
  • Eval examples: 992
  • Epochs: 2
  • Context length: 1,024
  • Precision: BF16

Final validation:

  • Loss: 0.9828
  • Perplexity: 2.672

Benchmark

ARC-Easy PIQA OpenBookQA HellaSwag Social-IQA Mean
32.74 57.94 28.80 28.33 35.31 36.63

Intended Use

Use this checkpoint as the general instruction-tuned Klyra baseline or as a parent for downstream post-training experiments.

About Klyra

Klyra-64M is a compact language-model research project initiated and developed by a student of Informatics Engineering at Politeknik Negeri Jakarta (State Polytechnic of Jakarta).

The model uses a MiniMind-compatible decoder-only architecture and was trained through a staged pipeline from random initialization.

Main project: Jahirrrr/Klyra-64M

Architecture

Component Value
Unique runtime parameters ~63.9M
Transformer layers 8
Hidden size 768
Attention heads 8
KV heads 4
Vocabulary ~6.4K
Context length 1,024
MoE No

Parameter note: Klyra uses tied input/output embeddings at runtime. Some exported safetensors files may contain both model.embed_tokens.weight and lm_head.weight as separate serialized tensors. The intended runtime model has 63,912,192 unique parameters after weight tying.

Downloads last month
12
Safetensors
Model size
68.8M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Jahirrrr/Klyra-64M-Instruct