Sarathi GGUF release

Upstream: Qwen/Qwen2.5-0.5B-Instruct at immutable revision 7ae557604adf67be50417f59c2c2f167def9a775. Conversion and quantization use llama.cpp revision 99f3dc32296f825fec94f202da1e9fede1e78cf9. Converted and quantized for Sarathi using llama.cpp.

Quantization File Bytes SHA-256 Minimum RAM (MiB) Recommended RAM (MiB)
Q5_K_M model-q5_k_m.gguf 420086080 0ba088dcd4d4ee6a6e74848de340b559157da286780487c3c2c510e6c6197730 2816 3840
Q8_0 model-q8_0.gguf 531068224 119b718b660a2778fd31dcb737daf867fcdbf0c027b92d4f4b3fbe6772d74621 2816 3840

Limitations

English-only deterministic MVP qualification. Desktop benchmarks are informational; verify physical-device compatibility separately.

Downloads last month
-
GGUF
Model size
0.5B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support