Nanbeige 4.2 3B Q4_K_M for browser WebGPU

This is a six-part, browser-friendly GGUF of Nanbeige/Nanbeige4.2-3B. Each shard is below 500 MiB so it can be downloaded and cached by wllama.

The Q4_K_M tensor payload comes from Tdamre/Nanbeige4.2-3B-GGUF. The tensors were not requantized. The container metadata was normalized from the Nanbeige fork's older convention to the convention merged into upstream llama.cpp:

  • nanbeige.block_count: 22 physical layers
  • nanbeige.num_loops: 2
  • nanbeige.skip_loop_final_norm: false

Validated with llama.cpp merge commit b77d646751d01c0962bc203b6809e9d94f7d50b7.

Load the first shard; llama.cpp/wllama discovers the other five from their standard split names.

The canonical public assets are the nanbeige4.2-3b-q4km-v1 GitHub release.

SHA-256

8fdb05799b34cfd3d3b11afaf22f0cf17bbb26a04558ea887115dd1569d93d3c  Nanbeige4.2-3B-Q4_K_M-00001-of-00006.gguf
806cdd41859ce1f4956efcd46d1e171accd8c96496b3168d1a82418cac0f3a9b  Nanbeige4.2-3B-Q4_K_M-00002-of-00006.gguf
19e28679b716217bddbf91684cfb23c03f43777065a2ef5cd63519f8e1db9551  Nanbeige4.2-3B-Q4_K_M-00003-of-00006.gguf
4e2f164fedb13384a4ba654f9987fa1c2f9c1b782340c3746caa0c42ef187aa1  Nanbeige4.2-3B-Q4_K_M-00004-of-00006.gguf
146237aade6e3b32eaefc791c10ca7e4b6baaf4768147c856b59f732b2d6370f  Nanbeige4.2-3B-Q4_K_M-00005-of-00006.gguf
d13a2d60af1eb0e61092bd69fcdd48db6a8ff0374f1d96f9b7f7ba1185b37aeb  Nanbeige4.2-3B-Q4_K_M-00006-of-00006.gguf
Downloads last month
20
GGUF
Model size
4B params
Architecture
nanbeige
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Michionlion/Nanbeige4.2-3B-GGUF-WebGPU

Quantized
(37)
this model