Gemma 4 12B GGUF fails to load in LM Studio / llama-server on Windows

#6
by nafishasan60 - opened

Screenshot 2026-06-04 011959

I’m trying to run this model in LM Studio on Windows:

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF

But it fails to load. LM Studio shows this error:

Failed to load the model

Engine protocol runtime llama-server for [model-id] exited before becoming healthy. exitCode=3221225620, signal=null

I’m using LM Studio, which appears to launch llama-server internally. The model crashes before it becomes healthy.

What I’ve checked so far:

I am trying to load the actual .gguf model file, not imatrix_unsloth.gguf_file
Restarted LM Studio
Tried lowering context length
Tried re-downloading the model file

My setup:

App: LM Studio
OS: Windows
Model: Gemma 4 12B Instruct GGUF

Questions:

Is Gemma 4 12B GGUF currently compatible with LM Studio?
Does LM Studio need a newer llama.cpp / llama-server build for this model?
Which quantization is recommended for LM Studio?
Is there any specific setting needed, such as context length, GPU offload, or flash attention?

Any help would be appreciated. Thanks!

Update -

When I try to load the model with mmproj-F32.gguf in the model folder I am getting above error.

when I remove the mmproj-F32.gguf from the model folder model loader perfectly.

its cuz new encoder arch iiuc, wait like a day cuz ts doesn't even have any downloads yet

never mind I am wrong, it doesn't even have an encoder it just uses like a projection for embedding

I'm having similar problems with llama.cpp:

srv load_model: [mtmd] failed to get memory usage of mmproj

failed to load model load_hparams: unknown projector type: gemma4uv

mtmd_init_from_file: error: Failed to load CLIP model

load_model: failed to load multimodal model

What I don't understand is why I've run even larger models on my PC without this happening.

I'm having similar problems with llama.cpp:

srv load_model: [mtmd] failed to get memory usage of mmproj

failed to load model load_hparams: unknown projector type: gemma4uv

mtmd_init_from_file: error: Failed to load CLIP model

load_model: failed to load multimodal model

What I don't understand is why I've run even larger models on my PC without this happening.

Update:

With a new version of llama.cpp released today, it was fixed.

how to solve the problem?

update llama.cpp

Sign up or log in to comment