Instructions to use AtomicChat/Qwen3.8-Flash-Next-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Use Docker
docker model run hf.co/AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AtomicChat/Qwen3.8-Flash-Next-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Qwen3.8-Flash-Next-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
- Ollama
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with Ollama:
ollama run hf.co/AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with Docker Model Runner:
docker model run hf.co/AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
- Lemonade
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AtomicChat/Qwen3.8-Flash-Next-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AtomicChat/Qwen3.8-Flash-Next-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
MTP support for AtomicChat Qwen3.8-Flash-Next GGUF
Hi, I'm using your Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64 GGUF with llama.cpp.
My setup is:
- RTX 5090 Laptop GPU β 24 GB VRAM
- 64 GB system RAM
- llama.cpp
b10713/ commit557614e02
The main model loads and runs correctly at around 11 tokens/sec. I'm trying to enable the model's 4B MTP layer for speculative decoding.
I tried using a separate Qwen3.8-Flash-Next-MTP-Q8_0.gguf, but llama.cpp fails when loading the MTP model with:
tensor 'blk.0.hc_attn_norm.weight' not found
The main AtomicChat GGUF itself loads and works normally.
Does the AtomicChat quant require a particular MTP GGUF, or is there a specific MTP file, conversion, llama.cpp branch, or configuration that you recommend for these AtomicChat quants?
Also, with my 24 GB VRAM + 64 GB RAM setup, would you expect the 4B MTP layer to provide a meaningful generation-speed improvement?
Thanks!
The error tensor 'blk.0.hc_attn_norm.weight' not found looks like the draft-load path rather than these quants.
Two things that may unblock you β note I have read the PRs but not tested MTP myself:
--spec-type draft-mtp for qwen4exp is not in main yet. It is #27836 (open, follow-up to #27742 which added the architecture). Its companion #28097, opened today, exists specifically for this failure: it adds draft-head-only GGUF support and fixes a draft-load regression, because #27836's loader requires hc_head_norm/hc_head_down/hc_head_up and the PLE block.
A correction to #6, where someone pointed at PR 27739 β that one and #27793 were discarded; #27742 is the merged one.
On the underlying request: these quants ship no draft head, so even with #27836 there is nothing to pair them with. @AtomicChat, would you consider publishing an MTP/NextN draft head alongside the M64 variants? #27836 includes converter support via --mtp.
Need MTP support . thanks!
https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/blob/main/MTP/README.md
that works, here is my bash file:
#!/bin/bash
# Resolve model path relative to this script's own directory, so it works from anywhere
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
MODEL="$SCRIPT_DIR/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64-00001-of-00033.gguf"
MMPROJ="$SCRIPT_DIR/mmproj-Qwen3.8-Flash-Next-BF16.gguf"
MPTFILE="$SCRIPT_DIR/mtp-Qwen3.8-Flash-Next-BF16.gguf"
/home/atb/Projects/llama.cpp.mtp/llama.cpp/build/bin/llama-server \
-m "$MODEL" \
--mmproj "$MMPROJ" \
-ngl 99 \
--n-cpu-moe 999 \
-fa on \
--parallel 1 \
--port 8080 \
--host 0.0.0.0 \
--ctx-size $((32*8*1024)) \
--no-warmup \
--temp 0.2 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--presence-penalty 0.0 \
--repeat-penalty 1.0 \
--seed 1902 \
--log-colors on \
--prio 2 \
--jinja \
--webui-mcp-proxy \
--spec-type draft-mtp \
--model-draft "$MPTFILE" \
--spec-draft-n-min 1 \
--spec-draft-n-max 2
Compared with the original version, how many tokens does MTP improve by?