PCSS-Qwen3.5-9B

GitHub Repository | Blog Post | Reproduction Notebook

PCSS-Qwen3.5-9B is an experimental checkpoint demonstrating that fine-tuning small base models on pure logic deduction traces can elicit substantial out-of-domain reasoning generalization. Starting from Qwen/Qwen3.5-9B-Base, this model was fine-tuned on just 500 5x5 zebra puzzles from the tamewild/instruct5 dataset, which contains zero mathematical data.

For fine-tuning on these long-form logic traces, the model was trained using PCSS (Per-Example Calibrated Sigmoid Scaler), an adaptive loss scaler derived from KTO. The training process leveraged MiSS (Matrix Shard Sharing) for parameter-efficient fine-tuning (PEFT) alongside a preliminary PCSS extension for Multi-Token Prediction (MTP) with k=1, completing in approximately 40 minutes on a single NVIDIA H200 NVL GPU. (Note: The accompanying reproduction notebook was run on an H100 SXM, producing comparable results but with minor variance due to training non-determinism).


Chat Template & Evaluation Mode

  • Control Token Pre-training: As noted on the Qwen/Qwen3.5-9B-Base model page, the base checkpoint was released with chat template control tokens (such as <|im_start|> and <|im_end|>) explicitly pre-trained "to allow efficient LoRA-style PEFT with the official chat template," mitigating the need to fine-tune embeddings.
  • Non-Thinking Mode: The official Qwen 3.5 chat template includes a built-in toggle (enable_thinking) that controls whether the model generates a chain of thought within <think> tags prior to its final response. In our experiments, we only train and evaluate in non-thinking mode (enable_thinking=False). This ensures the untrained base, our fine-tuned checkpoint, and the official post-trained model are evaluated under strictly comparable conditions without generation within these tags.

Benchmark Results

Models were served via vLLM (0.19.1) using the following configuration:

  • Max Completion Tokens: 16,384
  • Sampling Parameters: temperature=1.0, top_p=0.95, top_k=20
  • Presence Penalty: 0.0 for the untrained base and our fine-tuned model (1.5 was applied for the official post-trained checkpoint as recommended by Qwen).

To account for sampling variance, most benchmarks were evaluated using multiple independent samples. All reported scores are Pass@1 estimates. Pass@3 estimates are only reported for benchmarks with sufficient sample sizes (n >= 6).

Benchmark Qwen 3.5 9B Base (PCSS + MiSS) Qwen 3.5 9B Official Post-Trained Qwen 3.5 9B Base (Untrained)
MMLU Redux 90.45% 90.15% 89.32%
GPQA Diamond 71.04% 79.63% 61.11%
MATH-500 96.60% 97.40% 93.53%
AIME 2025 60.67% (Pass@3: 75.58%) 60.56% (Pass@3: 77.50%) 49.44% (Pass@3: 61.83%)
AIME 2026 64.33% (Pass@3: 75.50%) 69.44% (Pass@3: 84.67%) 50.00% (Pass@3: 65.50%)
ArXivMath 05/26 22.92% (Pass@3: 34.88%) 24.17% (Pass@3: 31.87%) 5.42% (Pass@3: 12.00%)
5x5 Zebra 80.50% (Pass@3: 97.05%) 48.00% 17.00%
6x6 Zebra 26.00% 5.00% —

Training Configuration & Hyperparameters

For full training details and mathematical derivations, please refer to the Reproduction Notebook and Blog Post.

  • Base Model: Qwen/Qwen3.5-9B-Base
  • Dataset: tamewild/instruct5 (500 5x5 zebra puzzles)
  • PEFT Method: MiSS (Matrix Shard Sharing), rank r=512
  • Loss Function: PCSS (beta=0.65, peak_scale=5.0)
  • Optimizer: 8-bit AdamW (LR: 5e-6, beta2: 0.99994, Weight Decay: 0.01)
  • LR Scheduler: Constant learning rate with 1-epoch linear warmup
  • Training: 5 epochs at batch size 1

Usage & Inference

Serving with vLLM

The checkpoint bundles the pre-configured chat_template.jinja. To serve the model in non-thinking mode with native MTP speculative decoding:

# Note: Configured to 2 tokens for better throughput during our benchmarking, despite k=1 fine-tuning.
vllm serve tamewild/PCSS-Qwen3.5-9B \
  --max-model-len 32000 \
  --language-model-only \
  --generation-config vllm \
  --host 127.0.0.1 \
  --port 18000 \
  --gpu-memory-utilization 0.90 \
  --default-chat-template-kwargs '{"enable_thinking": false}' \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 2}'

Prompt Formats

All evaluations were performed without a system prompt. The following zero-shot Chain-of-Thought templates were used (note the specific divergence for zebra puzzles, where the official post-trained model was evaluated using a prompt that requested a markdown table):

1. Mathematical Benchmarks (MATH-500, Full MATH, AIME 2025, AIME 2026, ArXivMath)

{problem}

Please reason step by step, and put your final answer within \boxed{}

2. Multiple-Choice Benchmarks (GPQA Diamond, MMLU Redux)

{problem}

Please reason step by step, and put your final letter choice within \boxed{\text{}}

3. Logic Grid Benchmarks (Zebra Puzzles)

For the Untrained Base and PCSS Fine-Tuned Models:

{problem}

Provide the solution grid as the final answer.

Please reason step by step, and put your final answer within \boxed{}

For the Official Post-Trained Model:

{problem}

Provide the solution table as the final answer.

Because the official post-trained model struggled to consistently output parsable LaTeX grids, its prompt was adjusted to request a markdown table. While this adjusted prompt omitted the explicit Chain-of-Thought instruction, the model natively outputs extensive step-by-step reasoning traces on these puzzles regardless. The table below contextualizes the generation lengths (in tokens) alongside benchmark accuracy across the 5x5 and 6x6 zebra puzzles.

Model Task Accuracy Median Tokens Average Tokens
Qwen 3.5 9B Official Post-Trained 5x5 Zebra 48.00% 11,580.0 11,367.3
Qwen 3.5 9B Base (PCSS Fine-Tuned) 5x5 Zebra 80.50% 9,209.5 10,126.1
Qwen 3.5 9B Official Post-Trained 6x6 Zebra 5.00% 13,773.0 13,772.0
Qwen 3.5 9B Base (PCSS Fine-Tuned) 6x6 Zebra 26.00% 15,543.5 14,654.2
Downloads last month
463
Safetensors
Model size
10B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tamewild/PCSS-Qwen3.5-9B

Finetuned
(613)
this model

Dataset used to train tamewild/PCSS-Qwen3.5-9B

Collection including tamewild/PCSS-Qwen3.5-9B

Article mentioning tamewild/PCSS-Qwen3.5-9B