Zedev-134M Logo

Zedev-134M

Parameters Context

Base causal language model with 134.47M parameters.


Architecture

Field Value
Parameters 134.47M
Layers 26
Hidden size 640
Attention GQA (10 Query / 2 KV heads, head_dim 64)
Intermediate size 1536 SwiGLU
Norm RMSNorm (eps 1e-5) + per-head QK-Norm
RoPE theta 10000.0
Context length 1024
Vocab size 50304 (50261 used, padded to /128)
Tokenizer GPT-2 BPE
Embeddings Tied (lm_head == embed_tokens)
HF Layout Qwen3ForCausalLM

Training & Corpus

Parameter Value
Training Time 17 hours, 41 minutes
Trained on Context Length 1024
Tokens seen 17.652B~

Pretraining Data

Source Tokens License
Wikipedia (20231101.en) 4.585B CC BY-SA 4.0
Gutenberg 4.500B Public Domain
DCLM-baseline s1 4.384B CC-BY-4.0
DCLM-baseline s2 4.261B CC-BY-4.0
OpenHermes 0.385B Apache 2.0
Total 18.115B

Benchmarks

Evaluated with lm-evaluation-harness (0-shot, float16, batch size 16).

Benchmark Metric Cagliari Zedev Improvement
ARC-Challenge acc_norm 22.78% 25.51% 🟢 +2.73%
acc 19.03% 23.12% 🟢 +4.09%
ARC-Easy acc 38.30% 45.62% 🟢 +7.32%
acc_norm 35.40% 41.67% 🟢 +6.27%
BoolQ acc 41.44% 45.84% 🟢 +4.40%
HellaSwag acc_norm 27.41% 31.17% 🟢 +3.76%
acc 27.04% 28.72% 🟢 +1.68%
OpenBookQA acc_norm 28.20% 29.00% 🟢 +0.80%
acc 14.80% 17.40% 🟢 +2.60%
PIQA acc 58.27% 61.43% 🟢 +3.16%
acc_norm 55.98% 60.94% 🟢 +4.96%
WinoGrande acc 50.36% 49.64% 🔴 -0.72%
BLiMP acc 80.22% 80.32% 🟢 +0.10%
WikiText-2 word_perplexity (↓) 36.95 29.67 🟢 -7.28
byte_perplexity (↓) 1.96 1.89 🟢 -0.07
bits_per_byte (↓) 0.97 0.9146 🟢 -0.0554

Special Tokens

Token ID
<|endoftext|> 50256 (EOS/BOS/PAD)
<|im_start|> 50257
<|im_end|> 50258
<think> 50259
</think> 50260

Usage

Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("Meridian-MRM/Zedev-134m")
model = AutoModelForCausalLM.from_pretrained(
    "Meridian-MRM/Zedev-134m", torch_dtype=torch.float16
).cuda().eval()

prompt = "The old lighthouse keeper walked to the edge of the cliff and"
ids = tok(prompt, return_tensors="pt").input_ids.cuda()

with torch.no_grad():
    out = model.generate(
        ids, max_new_tokens=128, do_sample=True,
        temperature=0.7, top_k=40, top_p=0.95,
        repetition_penalty=1.15, pad_token_id=50256,
    )
print(tok.decode(out[0], skip_special_tokens=True))

llama.cpp f16.gguf

./llama-cli -m zedev-134m-f16.gguf \
    -p "The old lighthouse keeper walked to the edge of the cliff and" \
    -n 128 --temp 0.7 --top-k 40 --top-p 0.95 --repeat-penalty 1.15

Limitations

  • English only.

License

Apache 2.0.

Credits

  • JAX / Flax / Optax (Google)
  • Wikipedia (CC BY-SA 4.0)
  • Project Gutenberg (Public Domain)
  • DCLM Dataset (CC-BY-4.0)
  • OpenHermes (Apache 2.0)
  • GPT-2 BPE (OpenAI)
  • Hugging Face, llama.cpp
  • Steve Carell

Citation

@misc{zedev134m,
  title  = {Zedev-134M},
  author = {Meridian-MRM},
  howpublished = {\url{https://huggingface.co/Meridian-MRM/zedev-134m}}
}
Downloads last month
295
Safetensors
Model size
0.1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Meridian-MRM/Zedev-134m

Quantizations
1 model

Datasets used to train Meridian-MRM/Zedev-134m

Space using Meridian-MRM/Zedev-134m 1

Evaluation results

  • Accuracy on BLiMP
    lm-eval
    80.320
  • Accuracy on PIQA
    lm-eval
    61.430
  • Normalized Accuracy on PIQA
    lm-eval
    60.940
  • Accuracy on WinoGrande
    lm-eval
    49.640
  • Accuracy on BoolQ
    lm-eval
    45.840
  • Accuracy on AI2 Reasoning Challenge
    lm-eval
    45.620
  • Normalized Accuracy on AI2 Reasoning Challenge
    lm-eval
    41.670
  • Accuracy on OpenBookQA
    lm-eval
    17.400