qwen35-2b-webgpu โ€” GPTQ-int4 weights

Custom GPTQ-int4 quant (group 32, byte-sliced nibbles, single tied q4 copy of the 248k-vocab embedding serving both the gather and the logits cascade) of Qwen/Qwen3.5-2B (text part), in the wire format of the qwen35-2b-webgpu browser runtime.

The first hybrid gated-DeltaNet + gated-attention LLM running in the browser โ€” pure WebGPU, hand-written WGSL kernels (chunked WY-representation delta-rule prefill, fused recurrent decode step, f32 state), no ONNX / transformers.js / GGUF.

  • ~1.07 GB; GPTQ calibrated on chat-formatted passages.
  • The runtime is verified against a PyTorch f32 reference (per-layer rms, cascaded argmax bit-exact vs the exact-logits path).

Try it: https://huggingface.co/spaces/borkiss/qwen35-2b-webgpu-demo ยท run /web/bench.html on your GPU and share the .log.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for borkiss/qwen35-2b-webgpu

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(291)
this model

Space using borkiss/qwen35-2b-webgpu 1