Memory-fit quantization · role-aware GGUF builds · running frontier LLMs on small hardware · MoE expert placement