This is the fastest Gemma 4 26b a4b model I've ever used.

#1
by DorkMckork1 - opened

It gets up to 58 tokens per second on my hardware with full context, GPU offload at max, CPU threads set to 8 (on AMD AM4 platform, Ryzen 5700x3d, Radeon RX 7700, Radeon RX 7600, 64 gigabytes of ram) It runs super fast and so far it stays solid in intelligence and reasoning at IQ4_XS quants. Excellent model! I'm getting this speed under Vulkan Llama.cpp. Awesome release. I'm using the mradermacher weighted quants (IQ4_XS) It runs great agentically, I'm using it in Odysseus and it does tool calls great. Thanks for this release, its really smart and really fast.

Sign up or log in to comment