Instructions to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: llama cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: llama cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: ./llama-cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Use Docker
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
- LM Studio
- Jan
- vLLM
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
- Ollama
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with Ollama:
ollama run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
- Unsloth Studio
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF to start chatting
- Pi
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with Docker Model Runner:
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
- Lemonade
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next ROCmFP4-FAST GGUF
ROCmFP4 quantization of Qwen/Qwen3.8-Flash-Next, sized for the 96 GiB VRAM carve-out of a Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S). 83.65 GiB across 5 shards, 4.06 bpw.
This is an experimental build, made for speed on one machine rather than for quality. ROCmFPx is an experimental quant family, it is carried in a fork rather than upstream, and this file is quantized without an importance matrix. It measures 4.6785 perplexity against 4.0068 for the unquantized model, which is a wider gap than a good 4-bit quant should have. If you want quality, use a mainstream quant; if you want ROCmFP4 kernels on RDNA3.5, this is what it is for.
Setup
qwen4exp and the ROCmFPx quant types are not in upstream llama.cpp yet, so build this branch:
git clone https://github.com/LaurentZuijdwijk/llama.cpp
cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
Run
Point at the first shard; the rest follow automatically.
./build/bin/llama-cli \
-m Qwen3.8-Flash-Next-ROCmFP4-FAST-00001-of-00005.gguf \
-ngl 99 -c 32768
Runs fully on the GPU: ~85 GiB VRAM, negligible host RAM. KV is roughly 24 KiB per token.
The n-gram table ships as 16 per-head tensors of ~1.3 GiB rather than one 20.9 GiB
tensor. A single tensor that size is past maxStorageBufferRange (4 GiB on most Vulkan
devices), so joined it can only ever live on the host - and on a machine whose VRAM
carve-out leaves less than 21 GiB for the host, that means swap.
Quantization
| tensors | type | bpw |
|---|---|---|
| MoE experts, attention, GDN, hyper-connections | Q4_0_ROCMFP4_FAST |
4.25 |
| n-gram PLE table (51.2 B params) | Q3_0_ROCMFPX |
3.50 |
token_embd, output |
Q6_K |
6.56 |
Perplexity
wikitext-2 raw, 145 chunks at -c 2048.
| build | PPL |
|---|---|
| unquantized reference (as reported in PR 27742) | 4.0068 +/- 0.02271 |
| this file | 4.6785 +/- 0.02780 |
Quantized without an importance matrix. An imatrix build may follow and should close much of that gap at the same size.
Converting other qwen4exp GGUFs
Files built the upstream way carry the table joined. This fork reads them, but the table stays host-side. To split it per head:
python gguf-py/gguf/scripts/gguf_split_ple_heads.py in-00001-of-000NN.gguf out.gguf
Head bounds come from the file's own KV, and the quantized bytes are copied through untouched - no dequantize, no requantize, no quality change. Works on any quant and on split inputs. Verified on unsloth's UD-IQ4_XS, which goes from OOMing a 30 GB host to 88.6 GiB fully resident on the GPU.
MTP draft head
mtp/ holds the model's own multi-token-prediction head, 2.27 GiB, exported from the same
checkpoint. Qwen trains it jointly with the target, so it drafts better than a separate
small model would.
./build/bin/llama-server \
-m Qwen3.8-Flash-Next-ROCmFP4-FAST-00001-of-00005.gguf \
-md mtp/Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST.gguf \
-ngl 99 --n-gpu-layers-draft 99 \
--spec-type draft-mtp --spec-draft-n-max 3 -c 32768
Measured on a Radeon 8060S, 250 tokens at temp 0, each config warmed up first:
| draft | t/s | acceptance |
|---|---|---|
| none | 28.1 | -- |
| n-max 2 | 31.8 | 0.695 |
| n-max 3 | 32.4 | 0.612 |
It is quantized to match the target rather than above it. A Q8_0 draft measured worse on both throughput and acceptance and cost 1.5 GiB more: acceptance is the draft agreeing with the target, and two models quantized the same way are wrong in the same places.
Adds ~2.3 GiB to the ~85 GiB the target uses.
Credits
qwen4exp support is the work of Daniel Han
(@danielhanchen), from
ggml-org/llama.cpp#27742 - an
unmerged draft. If it lands upstream, prefer upstream.
Quant formats hand-ported from ciru-ai/ROCmFPX. Base model by the Qwen team.
Not included
Vision tower.
License
Qwen Community License 1.0, included as LICENSE.
- Downloads last month
- -
We're not able to determine the quantization variants.
Model tree for agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF
Base model
Qwen/Qwen3.8-Flash-Next