My new 128M code completion model, named Vega 1.5 both because it has a new tokenizer... and I don't like how "VegaLM1" looks anyway.

This model was trained on ~5.5B tokens of stack-v3-train (python subset), and only understands python. Attempted FIM but didn't work because I'm a dummy. Bananamind Base Bench 1.1 code completion Elo is 1434, which is a meaningful jump over 42M-CodeCompletion (and is second only to SmolLM as of now!)

Model specs

  • Layers: 16
  • Hidden size: 768
  • Attention heads: 12
  • KV heads: 6
  • Context length: Trained on 2,048 tokens
  • Intermediate size: 2048
  • Vocabulary: 32,003
  • Parameters: 128,410,368
Downloads last month
31
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train CNWPlayer/Vega-1.5-128M-CodeCompletion