Qwen3.8 Flash Next Uncensored โ quantization in progress
INCOMPLETE CHECKPOINT โ the GGUF sets cannot yet be loaded as complete models.
This repository is being populated from the pinned source checkpoint
orcarouter/Qwen3.8-Flash-Next-Uncensored@8336e613ea508b13c2159bd0f68965d97a606b95.
Files already present are real quantized source weights, checked against the
per-tensor recipes and validated with the corresponding native GGUF reader.
Each variant requires all 28 shards before inference. See delivery-status.json
for the uploaded, hash-verified files. A completed model has not been delivered yet.
| Folder | Recipe | Runtime |
|---|---|---|
AD-3.84bpw-IQ4_XS-M64 |
AtomicChat actual 3.84 bpw tensor mapping and its imatrix | llama.cpp with qwen4exp support |
ROCmFP2-STRIX_LEAN-v2 |
pugant v2 actual 3.68 bpw tensor mapping and Unsloth imatrix | ROCmFPX with the strix-nebulosa qwen4exp patches |
The second variant is repacked into 28 shards to fit the bounded-storage
workflow. Its per-tensor types and shapes follow the reference two-shard model.
The tokenizer, chat template, architecture parameters, and integer PLE hash
constants come from the uncensored source, rather than the reference cards.
The text conversion omits MTP, matching both reference text-model sets.
The shared BF16 vision projector is delivered as
mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf (334 tensors, converted
from all 333 original vision tensors by the pinned upstream converter).
Whole-model inference, perplexity, and AMD device validation remain pending. No reference-model benchmark is claimed for this model.
References:
- https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
- https://huggingface.co/pugant/Qwen3.8-Flash-Next-ROCMFP4_STRIX_LEAN-GGUF
- https://github.com/pugant/strix-nebulosa
Weights remain subject to the source and base model licenses; see their model repositories. This checkpoint does not grant additional rights.
Source and reproduction notes
The root LICENSE is copied unchanged from the pinned uncensored source:
Qwen Community License 1.0. The source model-card license tag differs from
that file; this repository carries the actual license file with the weights.
provenance.json records source, recipes, quantizer commits, and calibration
hashes. The recipes were recovered from the reference GGUF tensor headers.
AtomicChat's displayed example command is not the selected 3.84 bpw tensor map.
The ROCm variant follows the September 3 v2 files, including their selective
Q3 expert tensors, and uses the separate Unsloth importance matrix.
Both variants preserve the uncensored source tokenizer and chat template. The reference importance matrices are reused; a fresh importance matrix was not generated from the uncensored model. Quality and AMD inference behavior still need to be evaluated after all shards are delivered.
- Downloads last month
- 54
4-bit
Model tree for gup98/test
Base model
Qwen/Qwen3.8-Flash-Next