Qwen3Loop-0.6B-SFT-Deep-Supervision (Hotfix 0.2: Suffix Layer Calibration & Unrolled Release)

Qwen3Loop 0.6B is an efficient recursive reasoning model available in two deployment formats:

  1. Native Looped Architecture (28 physical layers executed dynamically 56/42 times via custom engine patch).
  2. Standard Unrolled Architecture (42 physical layers running natively across ALL standard inference tools without any patches).

⚠️ Important Warning & Architecture Selection

🔄 Native Looped vs. Unrolled Architecture Trade-Off:

  • Native Looped Binaries (modelo_qwen3loop_sft_*.gguf): Maximum VRAM and disk compactness (604 MB in Q8_0, 2.8 GB VRAM). Requires our patched llama.cpp build or Python Qwen3LoopForCausalLM (trust_remote_code=True).
  • Unrolled Binaries (unrolled_modelo_qwen3loop_sft_*.gguf): 100% Universal Compatibility. This version trades off the weight-sharing VRAM benefit (827 MB in Q8_0, ~3.6 GB VRAM), but in exchange runs natively out-of-the-box in standard vanilla llama.cpp, Ollama, LM Studio, vLLM, and Hugging Face Transformers WITHOUT requiring any C++ patches or custom Python scripts!

🎮 Playground Readiness Notice (Hotfix 0.2)

Following the Hotfix 0.2 update (Suffix Layer Transduction Calibration), both the Native and Unrolled variants have completely eliminated previous formatting collapse, repetition loops, and token dissipation. The model is fully calibrated, highly responsive, and officially USABLE FOR PLAYGROUND and interactive testing!


🎯 Recommended Sampling & Inference Parameters

To achieve optimal generation quality (zero loops, crisp reasoning tags <think>, sharp answers), use the following empirically validated parameters:

Parameter Recommended Value Description
Temperature 0.60 (or 0.0 for code/math) Balanced creativity and precision
Top-K 40 Filters extreme long-tail tokens
Top-P 0.95 Nucleus sampling threshold
Repeat Penalty 1.12 Prevents loop degradation in small architectures
Repeat Last N 128 Repetition penalty lookback buffer
Context Window 32,768 tokens Native context length

🔬 Layerwise Probing Evolution (Hotfix 0.2)

Through Layerwise Early-Exit Probing across the 56 virtual execution steps:

Execution Stage / Layer Before Calibration (Old) Hotfix 0.2 (Step 75) Status
Prefix Exit (L06) 29.10% 5.20% Factual Anchoring
Loop Pass 1 Exit (L20_p1) 35.70% 100.00% Loop Convergence
Loop Pass 2 Exit (L20_p2) 67.20% 99.80% 🟢 Peak Recursive Reasoning
Suffix Entry (L21) 44.10% 99.61% 🟢 Smooth Thought Projection
Suffix Mid (L24) 27.10% 96.88% 🟢 Zero Noise
Suffix Final Exit (L27) 🛑 21.90% 🛑 90.62% 🏆 Dissipation Eliminated (+68.7%)

📊 97-Question Full Capability Benchmark

  • Global Accuracy: 79.38% (77/97) (Net +6.18% gain over DARE-TIES)
  • Writing & Text Generation: 100.00% (8/8) 🟢
  • Summarization: 100.00% (7/7) 🟢
  • Creativity: 100.00% (7/7) 🟢
  • Robustness: 100.00% (8/8) 🟢
  • Instruction Following: 90.00% (9/10) 🟢
  • Deductive Reasoning: 90.00% (9/10) 🟢
  • General Knowledge: 80.00% (8/10) 🟢
  • Mathematics & Algebra: 80.00% (8/10) 🟢
  • Coding & Syntax: 70.00% (7/10) 🟢
  • Inference Speed: 127.9 - 152.7 tokens/s (NVIDIA RTX 3060 CUDA)

📦 Model Files & Download Options

1. Universal Unrolled Models (No Patches Needed — Recommended for Ollama / LM Studio)

2. Compact Native Looped Models (Requires Patched llama.cpp Engine)

Downloads last month
1,688
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support