GSQ-RCO-GGUF Collection Non-uniform GGUF quantizations via GSQ + RCO: per-tensor mixed precision in standard GGUF form • 2 items • Updated 4 days ago • 30
Cerebras REAP Collection Sparse MoE models compressed using REAP (Router-weighted Expert Activation Pruning) method • 30 items • Updated Feb 25 • 152
view article Article GGML and llama.cpp join HF to ensure the long-term progress of Local AI +4 ggerganov, ngxson, allozaur, lysandre, victor, julien-c • Feb 20 • 510
Unsloth Dynamic 2.0 Quants Collection New 2.0 version of our Dynamic GGUF + Quants. Dynamic 2.0 achieves superior accuracy & SOTA quantization performance. • 121 items • Updated 21 days ago • 831