Qwen3.8-Flash-Next-GGUF

GGUF quantizations of Qwen/Qwen3.8-Flash-Next will be published here after the official weights and license are released and compatible conversion support is available.

🔔 Like and follow this repository for the GGUF release.
🐦 Follow @procrastiness on X for quantization updates.

Release status

The upstream model is currently listed by Qwen as an upcoming release scheduled for August 26, 2026. This repository does not contain model weights yet. It is being prepared for a real quantized release—not a renamed or unrelated checkpoint.

Qwen describes Qwen3.8-Flash-Next as:

  • A preview of the next-generation Qwen4 architecture
  • A multimodal Mixture-of-Experts (MoE) model
  • An upcoming open release under the official model ID Qwen/Qwen3.8-Flash-Next

Final architecture details, parameter counts, context length, license, chat template, multimodal projector requirements, and runtime compatibility will be copied from the official release—not inferred from rumors.

Planned GGUF files

The exact set will depend on the released architecture and practical file sizes. Intended variants include:

  • Q4_K_M
  • Q5_K_M
  • Q6_K
  • Q8_0

If the model requires a separate multimodal projector, compatible mmproj files will also be provided when the conversion toolchain supports them.

Release checklist

  • Verify the official source revision and license
  • Convert directly from Qwen/Qwen3.8-Flash-Next
  • Record the converter and llama.cpp revisions
  • Validate the tokenizer and chat template
  • Test text and multimodal inference where supported
  • Publish checksums, file sizes, and memory guidance
  • Credit Qwen and link the upstream model card

Expected usage

Usage commands will be added after compatibility is verified against an actual llama.cpp release. Commands will not be published before they can be tested with the released architecture.

Sources

Disclaimer

This is an independent community quantization project and is not affiliated with or endorsed by Qwen, Alibaba, or Hugging Face. The upstream model's license will govern redistribution and use of derived GGUF files.

Changelog

  • 2026-08-26: Repository prepared ahead of the official upstream release; no weights uploaded yet.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0xKitkat/Qwen3.8-Flash-Next-GGUF

Finetuned
(24)
this model