math-test-maxx

Merged 16-bit weights: unsloth/Qwen2.5-3B-Instruct + ReasonMaxxer offline-search LoRA (v2 recipe).

Source: ansz42/ReasonMaxxer (offline_search/, pack configs/test_pack_qwen25_3b.yaml).

Same-protocol greedy 0-shot

vLLM, boxed chat prompt, MathVerifier. Not the official 8-shot Qwen numbers.

Model GSM8K (n=1319) MATH-500 (n=500)
Base unsloth/Qwen2.5-3B-Instruct 84.6% 61.6%
This merge 85.0% 62.8%
vs base +0.4 pp +1.2 pp

An earlier train recipe (lr 2e-4, batch 1×4, clip 1.0) regressed to 82.9% / 53.8%. This upload is the milder retry.

Train recipe (preferred)

Knob Value
LoRA r=16, α=32, QKVO
learning_rate 2e-5
batch × grad accum 2 × 4 (effective 8)
max_grad_norm 0.1
data 300 MATH-500 items (seed 42), 12 offline search samples, entropy-weighted signed loss
steps 774 micro-steps / 194 Adam updates (one pass over informative rows)

In-loop 300-item eval (temp 0.6, n=4): pass@1 61.3%, pass@4 74.0%.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ba2han/math-test-maxx

Base model

Qwen/Qwen2.5-3B
Adapter
(477)
this model