Qwen3.8 Flash Next Uncensored โ€” quantization in progress

INCOMPLETE CHECKPOINT โ€” the GGUF sets cannot yet be loaded as complete models.

This repository is being populated from the pinned source checkpoint orcarouter/Qwen3.8-Flash-Next-Uncensored@8336e613ea508b13c2159bd0f68965d97a606b95. Files already present are real quantized source weights, checked against the per-tensor recipes and validated with the corresponding native GGUF reader. Each variant requires all 28 shards before inference. See delivery-status.json for the uploaded, hash-verified files. A completed model has not been delivered yet.

Folder Recipe Runtime
AD-3.84bpw-IQ4_XS-M64 AtomicChat actual 3.84 bpw tensor mapping and its imatrix llama.cpp with qwen4exp support
ROCmFP2-STRIX_LEAN-v2 pugant v2 actual 3.68 bpw tensor mapping and Unsloth imatrix ROCmFPX with the strix-nebulosa qwen4exp patches

The second variant is repacked into 28 shards to fit the bounded-storage workflow. Its per-tensor types and shapes follow the reference two-shard model. The tokenizer, chat template, architecture parameters, and integer PLE hash constants come from the uncensored source, rather than the reference cards. The text conversion omits MTP, matching both reference text-model sets. The shared BF16 vision projector is delivered as mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf (334 tensors, converted from all 333 original vision tensors by the pinned upstream converter). Whole-model inference, perplexity, and AMD device validation remain pending. No reference-model benchmark is claimed for this model.

References:

Weights remain subject to the source and base model licenses; see their model repositories. This checkpoint does not grant additional rights.

Source and reproduction notes

The root LICENSE is copied unchanged from the pinned uncensored source: Qwen Community License 1.0. The source model-card license tag differs from that file; this repository carries the actual license file with the weights.

provenance.json records source, recipes, quantizer commits, and calibration hashes. The recipes were recovered from the reference GGUF tensor headers. AtomicChat's displayed example command is not the selected 3.84 bpw tensor map. The ROCm variant follows the September 3 v2 files, including their selective Q3 expert tensors, and uses the separate Unsloth importance matrix.

Both variants preserve the uncensored source tokenizer and chat template. The reference importance matrices are reused; a fresh importance matrix was not generated from the uncensored model. Quality and AMD inference behavior still need to be evaluated after all shards are delivered.

Downloads last month
54
GGUF
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for gup98/test

Quantized
(24)
this model