BerkeliumGPT2-Coder-3b (MLX & Transformers, bf16)

This is the official full-precision bfloat16 (bf16) model fine-tuned from Qwen/Qwen2.5-3B on the BerkeliumCoding dataset across 176 permissive repositories with a 4,096 token context window.

Native SafeTensors format ready to run directly in both Apple MLX and Hugging Face Transformers / vLLM.

Model Details

  • Architecture: Qwen 2.5 (3 Billion Parameters)
  • Precision: bfloat16 (bf16 full precision, non-quantized)
  • Context Length: 4,096 tokens
  • Training Acceleration: Apple Silicon Unified Memory (Metal) via MLX
  • Format: Native SafeTensors

Run with MLX (Apple Silicon)

1. Installation

pip install mlx-lm

2. Run from CLI

mlx_lm.generate \
  --model Berkelium-ai/BerkeliumGPT2-Coder-3b \
  --prompt "def merge_intervals(intervals: list[list[int]]) -> list[list[int]]:" \
  --max-tokens 256 \
  --temp 0.2

3. Python API

from mlx_lm import load, generate

model, tokenizer = load("Berkelium-ai/BerkeliumGPT2-Coder-3b")
response = generate(
    model,
    tokenizer,
    prompt="def merge_intervals(intervals: list[list[int]]) -> list[list[int]]:",
    max_tokens=256,
    verbose=True
)
print(response)

Run with Hugging Face Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Berkelium-ai/BerkeliumGPT2-Coder-3b"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "def merge_intervals(intervals: list[list[int]]) -> list[list[int]]:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.2)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Berkelium-ai/BerkeliumGPT2-Coder-3b-MLX

Base model

Qwen/Qwen2.5-3B
Finetuned
(554)
this model