How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="RESMP-DEV/Qwen3-Next-80B-A3B-Instruct-NVFP4")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("RESMP-DEV/Qwen3-Next-80B-A3B-Instruct-NVFP4")
model = AutoModelForCausalLM.from_pretrained("RESMP-DEV/Qwen3-Next-80B-A3B-Instruct-NVFP4")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

Qwen3-Next-80B-A3B-Instruct-NVFP4

Quantized version of Qwen/Qwen3-Next-80B-A3B-Instruct using LLM Compressor and the NVFP4 (E2M1 + E4M3) format.

This time it actually works! We think

This should be the start of a new series of hopefully optimal NVFP4 quantizations as capable cards continue to grow out in the wild.


Model Summary

Property Value
Base model Qwen/Qwen3-Next-80B-A3B-Instruct
Quantization NVFP4 (FP4 microscaling, block = 16, scale = E4M3)
Method Post-Training Quantization with LLM Compressor
Toolchain LLM Compressor
Hardware target NVIDIA Blackwell (Untested on RTX cards) / GB200 Tensor Cores
Precision Weights & activations = FP4 • Scales = FP8 (E4M3)
Maintainer RESMP.DEV

Description

This model is a drop-in replacement for Qwen/Qwen3-Next-80B-A3B-Instruct that runs in NVFP4 precision Accuracy remains within ≈ 1 % of the FP8 baseline on standard reasoning and coding benchmarks.

Downloads last month
16,340
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RESMP-DEV/Qwen3-Next-80B-A3B-Instruct-NVFP4

Quantized
(72)
this model