FinTune GGUF (4-bit quantized)

4-bit (Q4_K_M) GGUF quantization of wrtdevcod/fintune-qwen2.5-1.5b-lora, a LoRA fine-tune of Qwen2.5-1.5B-Instruct for financial sentiment classification.

  • Original size (f16): 2.9 GB
  • Quantized size (Q4_K_M): 941 MB
  • Benchmark: 88.9% accuracy on 486-example held-out test set (base model: 50.6%)

CPU-friendly, works with llama.cpp / llama-cpp-python.

Downloads last month
58
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for wrtdevcod/fintune-qwen2.5-1.5b-gguf

Quantized
(254)
this model