a2genesis/Qwen3.8-27B-NVFP4
Image-Text-to-Text • 18B • Updated • 362
Post-training quantization (NVFP4, FP8, NVIDIA ModelOpt) · self-hosted LLM inference with vLLM and SGLang · AI-native application development. Independent software studio in Bochum, Germany.
A2Genesis is an independent software studio in Bochum, Germany. We build custom applications and run language models on our own hardware.
Here we publish quantized checkpoints of open-weight models. The quantization recipe, the calibration data and the serving configuration are documented in every model card, so a result can be reproduced rather than taken on trust.
Qwen/Qwen3.8-27B, produced with NVIDIA TensorRT Model Optimizer 0.45. About 21 GB instead of 55 GB, so it fits on a single GPU. Apache 2.0.Not affiliated with or endorsed by NVIDIA or the Qwen team. NVIDIA and TensorRT are trademarks of NVIDIA Corporation.