Instructions to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF # Run inference directly in the terminal: llama cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF # Run inference directly in the terminal: llama cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF # Run inference directly in the terminal: ./llama-cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Use Docker
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
- LM Studio
- Jan
- vLLM
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
- Ollama
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with Ollama:
ollama run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
- Unsloth Desktop
- Pi
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with Docker Model Runner:
docker model run hf.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
- Lemonade
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next ROCmFP4-FAST imatrix GGUF
The only quant of this model that fits fully in VRAM on a 128 GB unified-memory device without giving up quality: 87.06 GiB, 4.23 bpw, within 2.5% perplexity of the unquantized model (see below) — and it beats every alternative under 110 GiB we could find. Sized for the 96 GiB VRAM carve-out of a Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S). The whole model runs resident on GPU: nothing falls back to host RAM or CPU compute, not the experts, not the 51.2 B-parameter n-gram table.
Two layouts, same weights, same size (87.06 GiB either way — splitting the n-gram table per head is a byte-for-byte restructuring, not a re-quantize):
- root — n-gram table split per head, fully VRAM-resident.
joined/— n-gram table as one tensor, portable; needs--ngram-on-diskor host RAM for that tensor, since a single tensor that size is past what most Vulkan devices accept as one buffer.
Setup
Qwen3.8-Flash-Next itself is upstream (ggml-org/llama.cpp#27742). This fork is still needed for the ROCmFPx quant types and the per-head PLE layout.
git clone https://github.com/LaurentZuijdwijk/llama.cpp
cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
Run
# per-head table, fully on the GPU
./build/bin/llama-cli \
-m Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix.gguf \
-ngl 99 -c 32768
# joined table, off the GPU and off host RAM
./build/bin/llama-cli \
-m joined/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-joined.gguf \
-ngl 99 --ngram-on-disk --ngram-cache 8192 -c 32768
--ngram-cache defaults to 256 MiB; raise it for long generations or throughput drops
off over the course of a conversation.
The two layouts are interchangeable:
gguf_split_ple_heads.py
converts one to the other by copying quantized bytes verbatim — no dequantize, no
requantize, no quality change.
Perplexity
wikitext-2 raw, 145 chunks at -c 2048.
| build | PPL | vs. reference |
|---|---|---|
| unquantized reference (as reported in PR 27742) | 4.0068 +/- 0.02271 | - |
| this file | 4.1062 +/- 0.02329 | +2.48% |
Against AesSedai's quants, compared the fair way (each build's PPL against its own measured reference, since their test methodology differs from ours):
| build | size | PPL ratio vs. own reference |
|---|---|---|
| AesSedai IQ3_S | 107.38 GiB | +6.10% |
| this file | 87.06 GiB | +2.48% |
| AesSedai IQ4_XS | 117.13 GiB | +3.12% |
| AesSedai Q4_K_M | 135.38 GiB | +0.61% |
Beats their IQ4_XS and IQ3_S on quality at a smaller size.
MTP draft head
Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST-GGUF
holds the model's own multi-token-prediction head, exported from the same checkpoint.
Currently mismatched with this file's expert precision (draft is flat FP4, this file's
gate/up are Q4_K) — needs requantizing before it's worth enabling here.
Credits
qwen4exp support is the work of Daniel Han
(@danielhanchen), from
ggml-org/llama.cpp#27742, merged
upstream. This fork is only still needed for what's listed under Setup above.
Quant formats hand-ported from ciru-ai/ROCmFPX. Calibration corpora from bartowski and Thireus, credited above. Base model by the Qwen team.
Not included
Vision tower.
License
Qwen Community License 1.0, included as LICENSE.
- Downloads last month
- 421
We're not able to determine the quantization variants.
Model tree for agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Base model
Qwen/Qwen3.8-Flash-Next