Vesper-Coder-1.5B (v1, Gen-2 Checkpoint)

Vesper-Coder-1.5B (v1) is an experimental 1.5B-parameter dense code generation checkpoint initialized from Qwen/Qwen2.5-Coder-1.5B-Instruct and trained via two rounds of on-policy Direct Preference Optimization (DPO) on a 464-task MBPP training pool.

Important Evaluation Update (Normalized Harness Audit): Under an EvalPlus-normalized evaluation harness (max_new_tokens=384 + AST post-processing sanitizer) with strictly leak-free prompt-example gating, the Vesper-Coder-1.5B (v1, Gen-2) weights do not improve over Base Qwen2.5-Coder-1.5B-Instruct (all paired two-sided McNemar $p > 0.05$; slightly trailing base on raw counts).

Earlier gains reported under our initial harness (max_new_tokens=256, no AST sanitizer: 54.86% vs. 47.47% on MBPP; 63.41% vs. 58.54% on HumanEval) were an artifact of the 256-token generation cap truncating verbose Base Qwen2.5-Coder-1.5B-Instruct outputs. Raising max_new_tokens to 384 and applying standard AST sanitization increased Base Qwen2.5-Coder-1.5B-Instruct 1-shot accuracy by +10.51 pp on MBPP (47.47% -> 57.98%) and +6.70 pp on HumanEval (58.54% -> 65.24%), eliminating the v1 weight advantage. We retain this checkpoint publicly as a reproducible artifact and negative-result audit for small-model on-policy DPO.


1. Checkpoint Provenance & How DPO Preference Pairs Were Selected

  • Exported Checkpoint: Gen-2 (second generation of iterative on-policy LoRA DPO, $r=16, \alpha=32$, merged into dense bfloat16 weights via merge_and_unload()). In the initial v1 experiments, Gen-1 (55 steps) and Gen-2 (115 steps) were evaluated directly on the test sets without a separate validation split; the uploaded weights in this repository are Gen-2.
  • Training Pool (N = 464 Tasks): Constructed from the full google-research-datasets/mbpp dataset (974 tasks) by removing all 257 task IDs belonging to the official MBPP-Sanitized test split (Train ∩ MBPP-Sanitized Test = ∅). Note that 147 of the 378 tasks in EvalPlus MBPP+ overlap with this 464-task pool; therefore, MBPP+ is reported both on all 378 tasks and on the strictly disjoint N = 231 subset.
  • How DPO Winners (y_win) and Losers (y_lose) Were Selected:
    1. For each task in the 464-task training pool, the model first generated a 1-shot greedy (T = 0.0, max_new_tokens = 256) solution.
    2. Solutions were executed in a sandbox against the training task's ground-truth test_list assertions.
    3. Whenever the T = 0.0 attempt failed (y_lose), up to 2 retry turns were sampled (T = 0.40–0.60, max_new_tokens = 256) conditioned on the sandbox execution error and a truncated function-header prefix. Any retry that passed all test_list assertions was recorded as a verified winner (y_win), forming an on-policy preference pair (prompt, y_win > y_lose).
    4. Why v1 DPO Did Not Transfer Under the 384-Token Normalized Harness: Because rollouts during v1 harvesting were capped at max_new_tokens = 256, many y_lose trajectories failed simply due to token truncation at 256 tokens, and y_win selected shorter responses that fit inside 256 tokens. Once evaluation was normalized to max_new_tokens = 384 with AST sanitization, Base Qwen2.5-Coder-1.5B-Instruct no longer suffered from 256-token truncation.

2. Normalized-Harness Benchmark Results (max_new_tokens=384 + AST Sanitizer)

All evaluations below use identical generation caps (max_new_tokens=384), identical AST sanitization, and strictly leak-free test-time retry gating (retry loops execute only the public example already shown in the prompt: test_list[0] on MBPP and docstring >>> examples on HumanEval).

A. 1-Shot (T = 0.0) & 3-Turn Retry Accuracy Across 6 Suites

Model & Evaluation Regime MBPP Full (N=257) MBPP Held-Out (test_list[1:], N=257) MBPP+ All (35x, N=378) MBPP+ Disjoint (35x, N=231) HumanEval (N=164) HumanEval+ (80x, N=164)
Qwen2.5-Coder-1.5B-Instruct (1-Shot, Base) 149 / 257 (57.98%) 155 / 257 (60.31%) 182 / 378 (48.15%) 104 / 231 (45.02%) 107 / 164 (65.24%) 93 / 164 (56.71%)
Vesper-Coder-1.5B (1-Shot, Gen-2 Weights) 143 / 257 (55.64%) 149 / 257 (57.98%) 170 / 378 (44.97%) 103 / 231 (44.59%) 106 / 164 (64.63%) 91 / 164 (55.49%)
Qwen2.5-Coder-1.5B-Instruct (Naive 3-Turn Self-Debug) 152 / 257 (59.14%) 159 / 257 (61.87%) 186 / 378 (49.21%) 107 / 231 (46.32%) 108 / 164 (65.85%) 94 / 164 (57.32%)
Qwen2.5-Coder-1.5B-Instruct (3-Turn Partial Prune) 167 / 257 (64.98%) 173 / 257 (67.32%) 203 / 378 (53.70%) 120 / 231 (51.95%) 106 / 164 (64.63%) 93 / 164 (56.71%)
Vesper-Coder-1.5B (3-Turn Partial Prune, Gen-2) 163 / 257 (63.42%) 168 / 257 (65.37%) 189 / 378 (50.00%) 114 / 231 (49.35%) 109 / 164 (66.46%) 93 / 164 (56.71%)
Qwen2.5-Coder-1.5B-Instruct (3-Turn Full Reset / Best-of-3) 181 / 257 (70.43%) 185 / 257 (71.98%) 213 / 378 (56.35%) 126 / 231 (54.55%) 111 / 164 (67.68%) 97 / 164 (59.15%)

B. Paired Two-Sided Exact McNemar Tests: Vesper-Coder-1.5B (Gen-2) vs. Base Qwen2.5-Coder-1.5B-Instruct

Here b counts tasks where Vesper passes and Base fails; c counts tasks where Base passes and Vesper fails:

Comparison & Suite Base Pass Vesper Pass Net Diff b (Vesper only) c (Base only) Both Pass Neither Pass Exact Two-Sided McNemar $p$
1-Shot: MBPP Full (N=257) 149 143 -6 17 23 126 91 0.4296 (no improvement)
1-Shot: MBPP Held-Out test_list[1:] (N=257) 155 149 -6 18 24 131 84 0.4408 (no improvement)
1-Shot: EvalPlus MBPP+ All (N=378) 182 170 -12 21 33 149 175 0.1337 (no improvement)
1-Shot: OpenAI HumanEval (N=164) 107 106 -1 9 10 97 48 1.0000 (no improvement)
1-Shot: EvalPlus HumanEval+ (N=164) 93 91 -2 10 12 81 61 0.8318 (no improvement)
3-Turn Partial Prune: MBPP Full (N=257) 167 163 -4 10 14 153 80 0.5413 (no improvement)
3-Turn Partial Prune: MBPP Held-Out (N=257) 173 168 -5 11 16 157 73 0.4421 (no improvement)
3-Turn Partial Prune: EvalPlus MBPP+ All (N=378) 203 189 -14 19 33 170 156 0.0704 (no improvement)
3-Turn Partial Prune: OpenAI HumanEval (N=164) 106 109 +3 11 8 98 47 0.6476 (no improvement)
3-Turn Partial Prune: EvalPlus HumanEval+ (N=164) 93 93 0 10 10 83 61 1.0000 (no improvement)

3. Usage (transformers)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "axieyangb/Vesper-Coder-1.5B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

prompt = "Write a Python function `longest_increasing_subsequence(nums: list[int]) -> int` that returns the length of the longest strictly increasing subsequence."
messages = [
    {"role": "system", "content": "You are an expert Python programmer. Output clean, well-structured Python code."},
    {"role": "user", "content": prompt},
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)

with torch.no_grad():
    generated_ids = model.generate(**inputs, max_new_tokens=384, do_sample=False)

response = tokenizer.decode(generated_ids[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)

License

Released under the Apache-2.0 License.

Downloads last month
556
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for axieyangb/Vesper-Coder-1.5B

Finetuned
(214)
this model
Quantizations
1 model