Dasha v6 (dasha-llama32-3b-instruct-v6)

GGUF quantizations of Dasha's own chat model — a QLoRA fine-tune of Llama-3.2-3B-Instruct that gives her her personality and conversational style. Used as the default local "brain" runtime by DariaOS, loaded through llama.cpp (core/runtimes/llm_runtime.py).

Files

File Quantization Size Use case
dasha-v6-Q5_K_M.gguf Q5_K_M ~2.2 GB Default — good quality/size tradeoff, what DariaOS installs by default
dasha-v6-f16.gguf F16 ~6 GB Full precision, for machines with the VRAM/RAM to spare

Usage

Loaded automatically by DariaOS's installer (install.py → "Поставить llama.cpp сейчас?"), or directly with llama.cpp:

llama-server -m dasha-v6-Q5_K_M.gguf

License

Inherits Meta's Llama 3.2 Community License from the base model.

Downloads last month
-
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dariumi/dasha-v6-gguf

Quantized
(510)
this model