Instructions to use Shaik1903/ThinkLess-2B-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Shaik1903/ThinkLess-2B-SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Shaik1903/ThinkLess-2B-SFT") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Shaik1903/ThinkLess-2B-SFT") model = AutoModelForMultimodalLM.from_pretrained("Shaik1903/ThinkLess-2B-SFT", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Shaik1903/ThinkLess-2B-SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Shaik1903/ThinkLess-2B-SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Shaik1903/ThinkLess-2B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Shaik1903/ThinkLess-2B-SFT
- SGLang
How to use Shaik1903/ThinkLess-2B-SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Shaik1903/ThinkLess-2B-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Shaik1903/ThinkLess-2B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Shaik1903/ThinkLess-2B-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Shaik1903/ThinkLess-2B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Shaik1903/ThinkLess-2B-SFT with Docker Model Runner:
docker model run hf.co/Shaik1903/ThinkLess-2B-SFT
ThinkLess-2B-SFT
The SFT stage of ThinkLess-2B: Qwen3.5-2B fine-tuned on its own shortest correct solutions (with a same-family 9B teacher filling in the hardest problems). It is the most accurate model in the ThinkLess family and cuts reasoning length by 43–72% vs the base model.
Use this model when you want the highest accuracy, including competition-level math (on HMMT Feb 2025 it matches the base model, 19.2 vs 18.8). Use ThinkLess-2B when you want the shortest reasoning at near-identical accuracy on everyday math and science.
Results (81,920-token budget, thinking on)
| Benchmark | Qwen3.5-2B (base) | ThinkLess-2B-SFT | Change (paired, 95% CI) | Mean tokens: base → SFT |
|---|---|---|---|---|
| GSM8K | 86.4 | 91.2 | +4.9 [+3.2, +6.7] | 18,351 → 5,078 (−72%) |
| MATH-500 | 83.5 | 89.6 | +6.1 [+3.6, +8.7] | 28,824 → 16,351 (−43%) |
| GPQA-Diamond | 44.2 | 54.8 | +10.6 [+5.6, +15.9] | 51,826 → 29,590 (−43%) |
Answers cut off at the limit: GSM8K 8.1% → 0.5%, MATH-500 14.4% → 3.0%, GPQA 31.1% → 7.1%.
Under a hard thinking budget (thinking stopped at B tokens, then the model must answer), ThinkLess-2B-SFT is the best model at tight limits: GSM8K 75.1 / 82.1 / 88.1 / 89.0 and MATH-500 46.4 / 50.8 / 60.2 / 72.7 at 2k / 4k / 8k / 16k tokens (base: 64.9 / 68.0 / 71.9 / 78.0 and 40.6 / 40.5 / 47.8 / 58.1).
Training
- Data: 8,890 problems from GSM8K train and MATH train (levels 3–5), 13-gram decontaminated against the evaluation
sets. For each problem, the shortest correct and finished solution among: 4 base-model samples at an 8k cap,
4 more at 16k for problems still unsolved, and 2 from Qwen3.5-9B for the rest (59.8% / 10.8% / 29.4% of the data).
GSM8K was capped at the number of MATH examples, keeping its shortest solutions. Released in
ThinkLess-data (config
sft). - Recipe: full fine-tune, 2 epochs (140 steps), lr 1e-5 cosine, effective batch 128, max length 17,408 tokens, fp32 master weights with bf16 autocast, 8×H100 (41 min). Loss 0.437 → 0.381.
Full details, evaluation protocol and limitations: ThinkLess-2B.
How to use
vllm serve Shaik1903/ThinkLess-2B-SFT --speculative-config '{"method":"mtp","num_speculative_tokens":2}'
Use Qwen3.5's thinking-mode sampling (temperature 1.0, top-p 0.95, top-k 20, presence penalty 1.5).
- Downloads last month
- 169

