PCSS-Qwen3-4B

GitHub Repository | Blog Post | Reproduction Notebook

PCSS-Qwen3-4B is an experimental checkpoint demonstrating that fine-tuning small base models on pure logic deduction traces can elicit substantial out-of-domain reasoning generalization. Starting from unsloth/Qwen3-4B-Base (Unsloth's mirror of Qwen/Qwen3-4B-Base), this model was fine-tuned on 100 5x5 zebra puzzles from the tamewild/zebra_100 dataset, which contains zero mathematical data.

For fine-tuning on these long-form logic traces, the model was trained using PCSS (Per-Example Calibrated Sigmoid Scaler), an adaptive loss scaler derived from KTO. The training process leveraged MiSS (Matrix Shard Sharing) for parameter-efficient fine-tuning (PEFT), completing in approximately 6.5 minutes on a single NVIDIA H200 NVL GPU. (Note: Given the minimal 100-example training regime, independent runs exhibit higher variance than 500-sample configurations; reproduction runs generally yield MATH-500 scores in the 80%–86% range due to optimization and inference non-determinism).

Experimental Checkpoint Notice: PCSS-Qwen3-4B is an exploratory research artifact developed solely to study out-of-domain reasoning generalization. Unlike PCSS-Qwen3.5-9B, this model has a significantly narrower capability profile and lacks general conversational robustness. It is not intended for chat, multi-turn dialogue, or general assistant tasks.


Benchmark Results

Our fine-tuned checkpoint was evaluated via vLLM (0.19.1) using greedy decoding with 16,384 max completion tokens. All reported scores are Pass@1 estimates. To account for runtime batching non-determinism in vLLM, scores for MATH-500, 4x4 Zebra, and 5x5 Zebra were averaged over 3 evaluation runs, and AIME 2025 was averaged over 6 runs. Full MATH (5,000 problems) was evaluated with a single greedy run.

Baseline scores for the untrained base model and the official post-trained model in non-thinking mode are self-reported from the official Qwen 3 Technical Report.

Benchmark Qwen 3 4B Base (PCSS + MiSS) Qwen 3 4B Post-Trained (Self-Reported Non-Thinking Mode) Qwen 3 4B Base (Self-Reported Baseline)
MATH-500 84.60% 84.80% —
Full MATH (5,000) 85.26% (+31.16% delta) — 54.10% (4-shot)
AIME 2025 21.67% 19.10% —
4x4 Zebra 31.67% — —
5x5 Zebra 4.33% — —

Training Configuration & Hyperparameters

For full training details and mathematical derivations, please refer to the Reproduction Notebook and Blog Post.

  • Base Model: unsloth/Qwen3-4B-Base
  • Dataset: tamewild/zebra_100 (100 5x5 zebra puzzles)
  • PEFT Method: MiSS (Matrix Shard Sharing), rank r=512
  • Loss Function: PCSS (beta=0.65, peak_scale=5.0)
  • Optimizer: 8-bit AdamW (LR: 5e-6, beta2: 0.99994, Weight Decay: 0.01)
  • LR Scheduler: Constant learning rate with 1-epoch linear warmup
  • Training Duration: 5 epochs at batch size 1 (500 steps total)

Usage & Inference

Serving with vLLM

The checkpoint bundles the pre-configured chat_template.jinja. To serve the model via vLLM:

vllm serve tamewild/PCSS-Qwen3-4B \
  --max-model-len 32000 \
  --generation-config vllm \
  --host 127.0.0.1 \
  --port 18000 \
  --gpu-memory-utilization 0.90

Prompt Formats

Evaluations were conducted using the following zero-shot Chain-of-Thought templates:

1. Mathematical Benchmarks (MATH-500, Full MATH, AIME 2025)

{problem}

Please reason step by step, and put your final answer within \boxed{}

2. Logic Grid Benchmarks (Zebra Puzzles)

{problem}

Provide the solution grid as the final answer.

Please reason step by step, and put your final answer within \boxed{}
Downloads last month
19
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tamewild/PCSS-Qwen3-4B

Finetuned
(232)
this model

Dataset used to train tamewild/PCSS-Qwen3-4B

Collections including tamewild/PCSS-Qwen3-4B

Article mentioning tamewild/PCSS-Qwen3-4B