Cagliari-114M

Base language model, 114M params. Qwen3-compatible architecture, GPT-2 BPE tokenizer with 4 additional special tokens. Pretrained from scratch on 8.16B tokens.

No instruction tuning. No chat behavior. Completion only.

Architecture

Layers 24
Hidden size 576
Intermediate size 1536
Attention GQA, 9 Q / 3 KV heads, head_dim 64
Norm RMSNorm + QK-Norm
Embeddings tied
RoPE theta 10000.0
Context 1024
Vocab 50304 (50261 used)
Tokenizer GPT-2 BPE
Specials <|im_start|>, <|im_end|>, <think>, </think>, <|endoftext|>
EOS / BOS / PAD 50256

Corpus

Source Tokens License
Gutenberg 4.5B Public Domain
Wikipedia 2.37B CC BY-SA 4.0
OpenThoughts 0.82B Apache 2.0
OpenHermes 0.39B Apache 2.0
StackExchange 0.08B CC BY-SA 4.0
Total 8.16B

Training

Training Time 4 hours
Mix change (S2b) OpenThoughts 12% -> 2%, StackExchange 4% -> 2%
Tokens seen 5.24B (64.3% of corpus)
Optimizer AdamW, b1 0.9, b2 0.95, wd 0.01, clip 1.0
Batch 128 x 1024
Hardware TPU v5e-8

Benchmarks

Evaluated using the lm-evaluation-harness (v0.4.13) suite, zero-shot (num_fewshot=0) on full splits.

Benchmark Metric Score
BLiMP Accuracy 80.22%
ARC-Easy Accuracy 38.30%
Normalized Accuracy 35.40%
WikiText-2 Byte Perplexity 1.96
Word Perplexity 36.95
Bits Per Byte 0.97

Limitations

  • 114M params. Limited factual recall and reasoning.
  • Base model. No instruction following, no finetuning, no chat behavior.
  • Context 1024 tokens.
  • Outputs reflect pretraining data biases.
  • Not for production or safety-critical use.

License

Apache 2.0.

Citation

@misc{cagliari114m,
  title  = {Cagliari-114M},
  author = {Meridian-MRM},
  howpublished = {\url{https://huggingface.co/Meridian-MRM/cagliari-114m}}
}

Credits

  • Project Gutenberg (Public Domain)
  • Wikipedia (CC BY-SA 4.0)
  • OpenThoughts (Apache 2.0)
  • OpenHermes (Apache 2.0)
  • StackExchange (CC BY-SA 4.0)
  • GPT-2 BPE (OpenAI)
  • JAX / Flax / Optax, tiktoken, llama.cpp, Hugging Face
Downloads last month
834
Safetensors
Model size
0.1B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Meridian-MRM/cagliari-114m 1

Evaluation results