What is this?
Qwen4のアーキテクチャを先取り!Qwen3.8-Flash-NextをGGUFフォーマットに変換したものです。
imatrix dataset
日本語能力を重視し、日本語が多量に含まれるTFMC/imatrix-dataset-for-japanese-llmデータセットを使用しました。
なお、計算リソースの関係上imatrixの算出にはQ6_K量子化モデルを使用しました。
Quants
各クオンツ・推論努力とそのベンチマークスコア(API版Gemma4 31B採点によるElyza_tasks 100)をまとめておきます。
| クオンツ | スコア | コメント |
|---|---|---|
| Q5_K_M(No Think) | 4.53 | |
| Q4_K_M(No Think) | 4.52 | |
| IQ4_XS(No Think) | 4.55 | |
| MXFP4(No Think) | 4.43 | |
| reasoning_strength: low | 4.5 | |
| reasoning_strength: medium | 4.515 | |
| reasoning_strength: xhigh | 4.54 |
Note
-mm mmproj-Qwen3.8-Flash-Next-BF16.ggufでビジョンエンコーダーをロードし、Vision対応モデルとして使用することができます。
推論努力(Reasoning effort)はlow / medium / xhighから選択可能で、デフォルトはxhighです。
llama.cppのserverの場合は起動時に以下の引数を追加することで変更可能です。
--chat-template-kwargs '{\"reasoning_effort\":\"xhigh\"}'
License
qwen-community-1.0
Developer
Alibaba Cloud
- Downloads last month
- -
Hardware compatibility
Log In to add your hardware
We're not able to determine the quantization variants.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support