Add AProjQ4 imatrix GGUF

#22
by 0pp0 - opened

Direct Q8_0 to Q4_K requantization of the 215 dense attention projections using the 220k routed-and-dense DS4 imatrix. The resulting GGUF passed the pre-0731 official continuation vectors, including long_code_audit, and Metal SSD-streaming validation.

Added the 0731 AProjQ4 GGUF in commit 0d193661:

DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ4-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf

  • Size: 84,420,584,288 bytes
  • Hub SHA-256: 413cf0a68ca8d084e89f3f810eef5046b5308174d441a80017a0ff388933c767
  • Source: 0731 AProjQ8 GGUF
  • Requantization: 215 dense attention projection tensors, Q8_0 -> Q4_K
  • Imatrix: imatrix/DeepSeek-V4-Flash-chat-v2-routed-and-dense-ds4-220k.dat
  • Official-continuation score: 100 cases / 2,313 target tokens, avg_nll=0.398263336

The upload was committed directly to refs/pr/22; Xet deduplicated the 84.4 GB file down to about 2.87 GB of new uploaded data.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment