Qwen3.8-27B AEON ULTIMATE โ€” ROCmFPX iMatrix GGUF

ROCmFPX iMatrix quantizations of the full BF16 checkpoint published by Aeon / AEON-7.

Attribution

The source weights are Aeon's Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16, pinned at Hub revision 8f76e82ed7ef4de7735f5d4148fce7b643b00fae. Aeon deserves attribution for the BF16 model and its model work. This repository contains derived GGUF quantizations produced by vmlinux with the ROCmFPX toolchain; it is not a new training run or a claim of ownership of the source model.

The source model declares Apache 2.0 licensing. Review the source model card and applicable terms before redistribution or deployment.

Quantizations

File Preset Size SHA-256
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf Q4_0_ROCMFP4 17,735,469,440 bytes 34c04aaec2399fab8179a921d0ea1e9d3b93da213201ab0ec2930ee1420307b0
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP6-iMatrix.gguf Q6_0_ROCMFPX 22,528,383,360 bytes d096fdd0ece150cefa793a8fdae55076b63d61d9e135571de6c73b016a641ba1
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP8-iMatrix.gguf Q8_0_ROCMFPX 28,193,397,120 bytes 10eff5653c51e23a8a816761feea69b235c9003e80821759ecef41814843a761

All three use the same model-specific importance matrix: Qwen3.8-27B-AEON-ULTIMATE-iMatrix.imatrix.gguf. The matrix was generated from 339 chunks of 512 tokens using the shared calibration corpus, and each quantizer consumed 496 entries.

Runtime

These are experimental ROCmFPX tensor types and require a compatible ROCmFPX-enabled llama.cpp build. Stock upstream llama.cpp will not load them.

Example ROCm0 invocation:

hf download vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF \
  Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
  --local-dir ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX

./llama-completion \
  -m ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
  -dev ROCm0 -ngl all -c 8192 -n 256 -p "Hello."

The GGUFs retain the native one-layer MTP head. Thinking is enabled by the embedded Qwen template by default; pass the appropriate chat-template kwargs when an application needs thinking disabled.

Validation and provenance

All three files loaded and generated a short completion on ROCm0 with all layers offloaded. Detailed public build information is in BUILD_RESULTS.md, with exact hashes in SHA256SUMS and source/toolchain details in PROVENANCE.md.

Downloads last month
1,348
GGUF
Model size
3.39M params
Architecture
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(29)
this model