z-lab/Qwen3.8-27B-DFlash2
Text Generation • 2B • Updated • 21.1k • 162
Ignore my previous comment as I was testing the configuration on llama.cpp, which isn't as optimized as MLX based frameworks. However after trying this on mlx-vlm, I'm actually curious how you got this running long enough for your sanity testing. I tried to test it myself, and mlx-vlm is currently suffering from intense RAM spikes (4x the usage) which then very quickly causes OOMs. I opened an issue and I am working with collaborators who confirmed this is an issue.
Can you run this on a real benchmark? I tested this on my M4 Max with a real coding benchmark and I'm not seeing anything close to your 55 tok/s, I'm getting around 20.