Qwen3.8-Flash-Next β€” Splash Package v3 (Q4_0 weights, Q8 output)

Splash-native serving package of Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF, built for the Splash Metal engine on Apple Silicon.

Contents

  • target/ β€” 30 target-layer MDFN0031 bins + embedding.bin + head.bin (Q4_0 experts, Q8 router/output)
  • draft/ β€” 5 MTP draft-layer bins + model.bin (MDFD0004)
  • tokenizer/, vision/ β€” tokenizer + vision tower
  • manifest.json β€” schema v5, splash-packed-q4-qwen4exp

Usage

splash serve-native <dir>/target <dir>/draft --tokenizer <dir>/tokenizer

Notes

  • Derived from upstream Qwen3.8-Flash-Next; quantized Q4_0 with Q8 output/attention weights.
  • Rebuildable from the GGUF shards in the v3-GGUF repo.
  • License: Qwen Community License 1.0.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for nitinpanj/Qwen3.8-Flash-Next-Splash-v3

Finetuned
(67)
this model