Instructions to use 5ivatej/LessThink-Qwen3-4B-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 5ivatej/LessThink-Qwen3-4B-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="5ivatej/LessThink-Qwen3-4B-v1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("5ivatej/LessThink-Qwen3-4B-v1") model = AutoModelForCausalLM.from_pretrained("5ivatej/LessThink-Qwen3-4B-v1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use 5ivatej/LessThink-Qwen3-4B-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "5ivatej/LessThink-Qwen3-4B-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "5ivatej/LessThink-Qwen3-4B-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/5ivatej/LessThink-Qwen3-4B-v1
- SGLang
How to use 5ivatej/LessThink-Qwen3-4B-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "5ivatej/LessThink-Qwen3-4B-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "5ivatej/LessThink-Qwen3-4B-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "5ivatej/LessThink-Qwen3-4B-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "5ivatej/LessThink-Qwen3-4B-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use 5ivatej/LessThink-Qwen3-4B-v1 with Docker Model Runner:
docker model run hf.co/5ivatej/LessThink-Qwen3-4B-v1
LessThink-Qwen3-4B-v1
Qwen3-4B (thinking mode) post-trained with GRPO to think less: a reward that keeps
correctness and penalizes thinking tokens only (everything before </think>), measured
relative to the base model's own thinking length on each prompt.
Results (32K token cap, full benchmark sets, multi-seed)
| Benchmark | Qwen3-4B | LessThink | Δ acc (pp) | Δ thinking tokens (matched) |
|---|---|---|---|---|
| GSM8K | 95.0 | 94.8 | -0.2 | -63% |
| MMLU-Pro | 72.0 | 70.5 | -1.6 | -57% |
| MATH-500 | 96.8 | 96.2 | -0.7 | -52% |
| GPQA-Diamond | 54.9 | 53.0 | -1.9 | -50% |
| AIME 2024 | 72.7 | 70.2 | -2.5 | -34% |
| HMMT Feb 2025 | 46.2 | 39.6 | -6.7 | -34% |
| AIME 2025 | 63.5 | 55.2 | -8.3 | -34% |
Matched Δ = geometric mean over questions of the per-question thinking-length ratio (seeds averaged first). Truncation on competition math drops from 8-9% to 1.5-2.5%; loop rate falls to ~0.
Accuracy under a token budget
A response counts as correct at budget B only if it finished within B total tokens.
| Budget | AIME24 base / ours | AIME25 base / ours | MATH-500 base / ours |
|---|---|---|---|
| 4K | 1.9 / 15.8 | 0.6 / 15.6 | 56.4 / 79.8 |
| 8K | 23.5 / 43.5 | 17.1 / 32.5 | 83.4 / 91.6 |
| 16K | 57.9 / 65.4 | 48.8 / 50.8 | 94.7 / 95.1 |
| 32K | 72.7 / 70.2 | 63.5 / 55.2 | 96.8 / 96.2 |
Full tables and budget curves for all benchmarks are in eval/.
Limitations
- On competition math with an unlimited budget, accuracy drops 2.5-8 pp.
- Final answers are 5-20% shorter than the base model's, although only thinking tokens were penalized.
- No coding, agentic or instruction-following evaluation yet.
Training
- GRPO (TRL 0.21), LoRA r=64 / alpha=128 on all linear layers, merged into the base weights. Released checkpoint: step 50 of 400.
- Reward: correct -> 1 + aclip(1 - L/L_ref, -1, 1); wrong -> -aclip(L/L_ref - 1, 0, 1), where L = thinking tokens, L_ref = the base model's mean thinking length on correct samples for that prompt, a = 0.3 (linear warmup over 50 steps).
- Data: 5,376 prompts (GSM8K train, DeepScaleR subset, ARC-Challenge, SciQ, CommonsenseQA), 13-gram decontaminated against every evaluation set.
- lr 1e-5, KL beta 0.02, 8 generations per prompt, 64 completions per step, max completion 8,192 tokens, 1x H100.
Evaluation setup
vLLM 0.10.0, temperature 1.0, top_p 0.95, top_k 20, min_p 0, 32,768-token cap, truncated responses scored wrong. Seeds: AIME/HMMT 16, GPQA 8, MATH-500 4, GSM8K 2, MMLU-Pro 1.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
m = "5ivatej/LessThink-Qwen3-4B-v1"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForCausalLM.from_pretrained(m, torch_dtype="auto", device_map="auto")
msgs = [{"role": "user", "content": "What is 17*23?"}]
x = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=True, return_tensors="pt").to(model.device)
out = model.generate(x, max_new_tokens=4096, temperature=1.0, top_p=0.95, top_k=20, do_sample=True)
print(tok.decode(out[0][x.shape[1]:], skip_special_tokens=True))
Recommended sampling: temperature 1.0, top_p 0.95, top_k 20, min_p 0.
- Downloads last month
- 221