[Question] update in template

#53
by msimunek - opened

Hi,
I would like to ask if you measure also improvement on gemma-4-26B-A4B model. In tweet you mentioned only 31B and 4B.
https://x.com/googlegemma/status/2077449152062247219

Thank you very much.

I saw 0% improvement running the new template. None of the model files on here show a recent update, so I'm assuming either:

  1. Google still hasn't uploaded them, or
  2. The new template is supposed to carry all the improvements, but fell far from the mark.

0 improvement here as well. I have tested the latest chat template with 12B and 26B and they still make the mistake of thinking about using the tool but then they will just skip right to the answer, or say they want to use the tool now but end the generation. The model needs to be retrained to fix those issues and I hope they will release a Gemma 4.1 in the very near future with better performance to be more competitive with Qwen. Perferably then all models will use the unified architecture of 12B and include audio as well.

Gemma 4 is something very special, it just needs refinement.

Google org

Hi,
I would like to ask if you measure also improvement on gemma-4-26B-A4B model. In tweet you mentioned only 31B and 4B.
https://x.com/googlegemma/status/2077449152062247219

Thank you very much.

yes, the performance related to flash attention should be for all Gemma 4 models (on the specific GPUs)

@plz12345 we updated the chat template, not the model files.

@Dampfinchen I'm trying to debug your prompt (from the other thread). The new chat template fixes many interactions and I'm trying to understand what really happens with the prompt you shared (thanks for that)!

Sign up or log in to comment