Aether 2.5 Pro - GGUF

Pre-quantized GGUF binaries for Aether 2.5 Pro.

Aether 2.5 Pro is a fine-tuned version of Qwen2.5-3B-Instruct (SFT + LoRA). It offers improved reasoning, stronger instruction following, and better multilingual performance in German and English.

πŸ–₯️ Easiest way to run the model:
Download MonoAIStudio – our local chat application.
It comes pre-installed with Aether 2.5, Aether 2.5 Pro and Aether 2.5 Coder.
πŸ‘‰ Download MonoAIStudio.zip

πŸ”— Looking for the LoRA Adapter?
πŸ‘‰ Maxilicious20/Aether-2.5-Pro


πŸ“¦ Available Quantizations

Filename Quantization Quality Size Recommendation
aether-2.5-pro-f16.gguf F16 Maximum ~6.0 GB Best quality
aether-2.5-pro-q8_0.gguf Q8_0 Very High ~3.2 GB Excellent quality
aether-2.5-pro-q4_k_m.gguf Q4_K_M Balanced ~1.9 GB Recommended

πŸš€ How to Run

1. MonoAIStudio (Recommended)

  1. Download MonoAIStudio.zip
  2. Extract it
  3. Run MonoAIStudio.exe
  4. All Aether 2.5 models are already included

2. LM Studio

  1. Open LM Studio
  2. Search for Maxilicious20/Aether-2.5-Pro-GGUF
  3. Download aether-2.5-pro-q4_k_m.gguf
  4. Load the model

3. llama.cpp

./llama-cli -m aether-2.5-pro-q4_k_m.gguf -p "Hello!" -n 256
Downloads last month
73
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including Maxilicious20/Aether-2.5-Pro-GGUF