MOSS-Music-8B-Instruct โ€” GGUF Q4_K_M

A 4-bit K-quant of the language model of MOSS-Music-8B-Instruct (OpenMOSS, Apache-2.0), for the native GGML runner ace-caption from HOT-Step and the YuE2 / MiniMax Music 3 studios, where it describes songs by ear for LoRA datasets.

File Size What
moss-lm-q4_k_m.gguf 5.2 GB language model, Q4_K_M

The audio tower is unchanged: take moss-aud-f16.gguf from scragnog/MOSS-Music-8B-Instruct-GGUF, which also has the f16 and q8_0 language models this quant was made from.

How it was made

moss-lm-f16.gguf from scragnog, requantised with HOT-Step's quantize (tools/quantize.cpp, the llama-quantize policy): attn_v and ffn_down one step up, the embedding and the output head at Q6_K, norms in F32. 254 of 399 tensors quantised; 16.4 GB โ†’ 5.2 GB.

Use

Same as the q8_0 file: point ace-caption at this language model and the f16 audio tower. It needs about 3 GB less video memory than q8_0.

Quality against q8_0 has not been measured yet; results will be added here.

Downloads last month
27
GGUF
Model size
8B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nerualdreming/MOSS-Music-8B-Instruct-GGUF

Quantized
(2)
this model