Q1D 4B Blk0 V2

Research artifact from Q1 Descent โ€” reconstructing per-block 1-bit "intelligence-density" recovery for open-weight Qwen3 models. Shelved 2026-06-19; published for the record. See the project debrief for context.

Method (Q1_0_g128, 1.125 bpw). Selected transformer block(s) quantized to 1-bit (sign + per-128-group FP16 scale) and trained back toward the full-precision teacher with a straight-through estimator (FP32 master, forward = exact Q1_0 so signs and scales move; KL to the FP16 teacher on C4). All other blocks stay F16. This is not a fully 1-bit model โ€” it isolates how far a small number of 1-bit blocks can be healed.

This model. Block 0 only at 1-bit (V2, 1200-step retrain); blocks 1-35 F16.

State / eval. GSM8K flex 0.84 / strict 0.81 (n=100; F16 teacher 0.76/0.56). Capability retained at 1 block. ~7/100 malformed on long generations.

Use. llama.cpp / LM Studio. Ships the closed-think (no-think) chat template; greedy (temp 0) recommended.

Caveats. Early research artifact, single seed, small-n evals (n=100, SEโ‰ˆ0.04 โ€” don't over-read sub-0.08 gaps). Known occasional malformed-token outputs in some contexts (a byte-level-tokenizer effect; see debrief). Not affiliated with PrismML or the Qwen team.

Downloads last month
5
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

1-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tg-techie-agents/Q1D-4B-Blk0-V2

Finetuned
Qwen/Qwen3-4B
Quantized
(311)
this model