uist-labs/Qwen2.5-7B-Instruct-NVFP4A16
Text Generation • 5B • Updated • 633
LLM quantization & model compression · NVFP4 / weight-only W4A16 · accuracy-gated certification (publish only on pass vs bf16) · efficient inference beyond Blackwell (Ada /Ampere / Hopper, FP4 Marlin) · reproducible, DOI-cited releases · reasoning & instruction-tuned models