ACE-Step Captioner GGUF

This repository contains a llama.cpp-compatible GGUF conversion of ACE-Step/acestep-captioner.

The initial upload is intended to include:

  • acestep-captioner-Q4_K_M.gguf
  • acestep-captioner-Q6_K.gguf
  • acestep-captioner-Q8_0.gguf
  • acestep-captioner-mmproj-bf16.gguf

Files

  • acestep-captioner-Q4_K_M.gguf: Quantized text model for inference with llama.cpp
  • acestep-captioner-Q6_K.gguf: Higher-quality 6-bit quantized text model
  • acestep-captioner-Q8_0.gguf: Higher-quality 8-bit quantized text model
  • acestep-captioner-mmproj-bf16.gguf: Multimodal projector required for audio input

llama.cpp

This model requires a recent llama.cpp build with Qwen2.5-Omni audio support.

Tested with this fork / branch which fixed audio inference bugs: https://github.com/tpsjr7/llama.cpp/tree/ted/fix-qwen-audio-cleanup-merge

Example:

llama-cli -m ./acestep-captioner-Q4_K_M.gguf \
  --mmproj ./acestep-captioner-mmproj-bf16.gguf \
  --audio ./song.mp3 \
  -p "*Task* Describe this audio in detail" \
  -n 512 --temp 0 --single-turn --simple-io -ngl 999 --ctx-size 8192

Swap in acestep-captioner-Q6_K.gguf or acestep-captioner-Q8_0.gguf if you want a less aggressive quantization.

Downloads last month
240
GGUF
Model size
8B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tpsjr7/acestep-captioner-GGUF

Quantized
(3)
this model