Berna Prototype 162M

Decoder-only Transformer trained from scratch on 98M tokens.

Model Details

Property Value
Parameters 162.42M
Architecture Decoder-only
Layers 12
Hidden size 768
Intermediate size 3072
Attention heads 12
Vocab size 32000
Context length 2048
Position encoding RoPE
Normalization RMSNorm
Activation SwiGLU
Precision BF16

Training Details

Property Value
Hardware RTX 5090
Time 40 minutes
Tokens 98,295,808
Throughput 40,815 tok/s
Optimizer AdamW
Scheduler Cosine warmup
Peak LR 6e-4
Batch size 16 x 512
Steps 12000
Final loss 5.58
Perplexity 265

Architecture Features

  • RMSNorm pre-normalization
  • RoPE position embeddings
  • SwiGLU activation
  • Separate LM head
  • Causal attention mask
  • No bias terms

Usage

pip install torch transformers
from transformers import AutoConfig, AutoModelForCausalLM
from transformers import PreTrainedTokenizerFast
from modeling_berna import BernaConfig, BernaForCausalLM

AutoConfig.register("berna", BernaConfig)
AutoModelForCausalLM.register(BernaConfig, BernaForCausalLM)

model = AutoModelForCausalLM.from_pretrained(".", local_files_only=True)
tokenizer = PreTrainedTokenizerFast.from_pretrained(".", local_files_only=True)
model = model.to("cuda")

Generate

prompt = "Hello"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(output[0]))

Status

Prototype. Trained on 98M tokens only. Text quality is limited at this stage.

Quality Tokens Progress
Basic words 300M 33%
Phrases 1B 10%
Sentences 3B 3%
Coherent 40B 0.25%

Roadmap

  • Environment setup
  • BPE tokenizer (32k)
  • Llama-style architecture
  • Atomic checkpointing
  • Kill -9 recovery
  • HF-compatible export
  • Train on 98M tokens
  • Scale to 1.5B

License

This model is released under the Berna Custom License. See LICENSE for full terms.

Key restrictions:

  • Name "Berna" must be preserved (no renaming).
  • Modification of weights is exclusive to the creator.
  • Commercial use requires separate written license.

For commercial licensing: contact @muhamedkamil on HuggingFace.

Downloads last month
362
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support