lab-bitnet BitNet-b1.58-2B-4T I2_S

Official Microsoft b1.58 I2_S GGUF (ggml-model-i2_s.gguf), not GGUF Q4.

Card targets: gpu5080 sm120, akula-prime sm86, gpu-1080ti sm61 (file is ~1.2 GB).

Layer-group ternary convert (scripts/layer_group_quant.py) checkpointed on gpu5080 against the bf16 safetensors. See ggml-model-i2_s.quant.json.

Downloads last month
159
GGUF
Model size
2B params
Architecture
bitnet-b1.58
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support