vllm crash: missing module for Glm5NextTextLinearAttention

#26
by luonist - opened

I installed vllm nightly:

uv venv
source .venv/bin/activate
uv pip install -U vllm --pre
--extra-index-url https://wheels.vllm.ai/nightly/cu130
--extra-index-url https://download.pytorch.org/whl/cu130
--index-strategy unsafe-best-match

and launched the model, following the recipe (for B200 cards) on vllm website:

export VLLM_ENGINE_READY_TIMEOUT_S=3600
vllm serve zai-org/GLM-5.3-Flash
--tensor-parallel-size 4
--max-model-len 262144
--kv-cache-dtype fp8
--tool-call-parser glm47
--enable-auto-tool-choice
--reasoning-parser glm45

The resulting versions of vllm and transformers are:
transformers==5.16.1
vllm==0.28.1rc1.dev7+g4a6a3272e

I got this error:
ValueError: There is no module or parameter named 'model.language_model.layers.0.self_attn.k_conv1d' in TransformersMultiModalMoEForCausalLM. The available parameters belonging to model.language_model.layers.0.self_attn (Glm5NextTextLinearAttention) are: {'model.language_model.layers.0.self_attn.v_proj.weight', 'model.language_model.layers.0.self_attn.forget_gate.A_log', 'model.language_model.layers.0.self_attn.forget_gate.f_b_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.o_norm.weight', 'model.language_model.layers.0.self_attn.b_proj.weight', 'model.language_model.layers.0.self_attn.conv1d.weight', 'model.language_model.layers.0.self_attn.v_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.forget_gate.f_b_proj.weight', 'model.language_model.layers.0.self_attn.forget_gate.dt_bias', 'model.language_model.layers.0.self_attn.g_a_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.q_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.o_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.forget_gate.f_a_proj.weight', 'model.language_model.layers.0.self_attn.k_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.k_proj.weight', 'model.language_model.layers.0.self_attn.b_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.g_a_proj.weight', 'model.language_model.layers.0.self_attn.forget_gate.f_a_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.q_proj.weight', 'model.language_model.layers.0.self_attn.g_b_proj.weight', 'model.language_model.layers.0.self_attn.g_b_proj.weight_scale_inv', 'model.language_model.layers.0.self_attn.o_proj.weight'}

It looks vllm still lacks support to the model architecture.

you should use this
docker pull vllm/vllm-openai:glm53-flash

Trying the docker image, I got:
FileNotFoundError: [Errno 2] No such file or directory: 'zai-org/GLM-5.3-Flash/processor_config.json'

this may be you not download the full model,as config is here
https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/processor_config.json

Yep, the file is there; however, vllm fails to resolve the file path. Passing the local path to the downloaded files, instead of the HF model name solved the issue.

Sign up or log in to comment