Instructions to use tamewild/PCSS-Qwen3-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tamewild/PCSS-Qwen3-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tamewild/PCSS-Qwen3-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tamewild/PCSS-Qwen3-4B") model = AutoModelForCausalLM.from_pretrained("tamewild/PCSS-Qwen3-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tamewild/PCSS-Qwen3-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tamewild/PCSS-Qwen3-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tamewild/PCSS-Qwen3-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tamewild/PCSS-Qwen3-4B
- SGLang
How to use tamewild/PCSS-Qwen3-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tamewild/PCSS-Qwen3-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tamewild/PCSS-Qwen3-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tamewild/PCSS-Qwen3-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tamewild/PCSS-Qwen3-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tamewild/PCSS-Qwen3-4B with Docker Model Runner:
docker model run hf.co/tamewild/PCSS-Qwen3-4B
PCSS-Qwen3-4B
GitHub Repository | Blog Post | Reproduction Notebook
PCSS-Qwen3-4B is an experimental checkpoint demonstrating that fine-tuning small base models on pure logic deduction traces can elicit substantial out-of-domain reasoning generalization. Starting from unsloth/Qwen3-4B-Base (Unsloth's mirror of Qwen/Qwen3-4B-Base), this model was fine-tuned on 100 5x5 zebra puzzles from the tamewild/zebra_100 dataset, which contains zero mathematical data.
For fine-tuning on these long-form logic traces, the model was trained using PCSS (Per-Example Calibrated Sigmoid Scaler), an adaptive loss scaler derived from KTO. The training process leveraged MiSS (Matrix Shard Sharing) for parameter-efficient fine-tuning (PEFT), completing in approximately 6.5 minutes on a single NVIDIA H200 NVL GPU. (Note: Given the minimal 100-example training regime, independent runs exhibit higher variance than 500-sample configurations; reproduction runs generally yield MATH-500 scores in the 80%–86% range due to optimization and inference non-determinism).
Experimental Checkpoint Notice:
PCSS-Qwen3-4Bis an exploratory research artifact developed solely to study out-of-domain reasoning generalization. UnlikePCSS-Qwen3.5-9B, this model has a significantly narrower capability profile and lacks general conversational robustness. It is not intended for chat, multi-turn dialogue, or general assistant tasks.
Benchmark Results
Our fine-tuned checkpoint was evaluated via vLLM (0.19.1) using greedy decoding with 16,384 max completion tokens. All reported scores are Pass@1 estimates. To account for runtime batching non-determinism in vLLM, scores for MATH-500, 4x4 Zebra, and 5x5 Zebra were averaged over 3 evaluation runs, and AIME 2025 was averaged over 6 runs. Full MATH (5,000 problems) was evaluated with a single greedy run.
Baseline scores for the untrained base model and the official post-trained model in non-thinking mode are self-reported from the official Qwen 3 Technical Report.
| Benchmark | Qwen 3 4B Base (PCSS + MiSS) | Qwen 3 4B Post-Trained (Self-Reported Non-Thinking Mode) | Qwen 3 4B Base (Self-Reported Baseline) |
|---|---|---|---|
| MATH-500 | 84.60% | 84.80% | — |
| Full MATH (5,000) | 85.26% (+31.16% delta) | — | 54.10% (4-shot) |
| AIME 2025 | 21.67% | 19.10% | — |
| 4x4 Zebra | 31.67% | — | — |
| 5x5 Zebra | 4.33% | — | — |
Training Configuration & Hyperparameters
For full training details and mathematical derivations, please refer to the Reproduction Notebook and Blog Post.
- Base Model:
unsloth/Qwen3-4B-Base - Dataset:
tamewild/zebra_100(100 5x5 zebra puzzles) - PEFT Method: MiSS (Matrix Shard Sharing), rank r=512
- Loss Function: PCSS (beta=0.65, peak_scale=5.0)
- Optimizer: 8-bit AdamW (LR: 5e-6, beta2: 0.99994, Weight Decay: 0.01)
- LR Scheduler: Constant learning rate with 1-epoch linear warmup
- Training Duration: 5 epochs at batch size 1 (500 steps total)
Usage & Inference
Serving with vLLM
The checkpoint bundles the pre-configured chat_template.jinja. To serve the model via vLLM:
vllm serve tamewild/PCSS-Qwen3-4B \
--max-model-len 32000 \
--generation-config vllm \
--host 127.0.0.1 \
--port 18000 \
--gpu-memory-utilization 0.90
Prompt Formats
Evaluations were conducted using the following zero-shot Chain-of-Thought templates:
1. Mathematical Benchmarks (MATH-500, Full MATH, AIME 2025)
{problem}
Please reason step by step, and put your final answer within \boxed{}
2. Logic Grid Benchmarks (Zebra Puzzles)
{problem}
Provide the solution grid as the final answer.
Please reason step by step, and put your final answer within \boxed{}
- Downloads last month
- 19