You can now Run GLM-5.3-Flash Locally! ✨

#4
by danielhanchen - opened

Hey guys, GLM-5.3-Flash can now be run locally in Unsloth Desktop! ✨

Run 3-bit on 128GB RAM or 1-bit on 100GB. The bigger ones are still uploading.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.

Unsloth GitHub: https://github.com/unslothai/unsloth
Guide: https://unsloth.ai/docs/models/glm-5.3-flash

glm-5.3-flash unsloth desktop
danielhanchen pinned discussion

Please Unsloth make TQ1_0 I need fit on 96G

Does the MTP layer work? I tried loading them but it results in an error message saying nextm isn't implemented

Please Unsloth make TQ1_0 I need fit on 96G

Yes agreed please and thx

What type of speed are people getting with 128GB Macs?

GLM-5.3-Flash_Unsloth
Unsloth fork llama.cpp, 08-2026 pull
Default thinking runs ridiculously long, but 'low' seems okay so far
-mmap, -fit on
-b 4096 -ub 1024
full 1m ctx seems to be runnable

Major bottlneck 2ch DDR5, 4800mt/s.
At 108k context: PP 60-70, TG 6-7, rtx3090 gpus @100w /gpu, cpu @6threads 22%, DRAM use 118GB.
At 240k context: PP 8-11, TG 6

Increasing to 1M context drops it to a slog (why? until the ctx is actually filled, it should be fast... there must be a patch...)

Performance for this 176GB system is in same league as Minimax-M2.7, MiMo-2.5 and DeepSeek4-flash. All are so smart that it's hard to differentiate between them. GLM5.3 definitely up there - at these sizes the main productivity differences come from my degree of resonance and shared assumptions, the communication, the alignment with language.

Thank you for sharing your results.

Sign up or log in to comment