Q4_K_M fails to load: missing tensor 'blk.40.ssm_conv1d.weight' -- is this a known issue or something on my end?

#1
by G-R-A-V-I-T-Y - opened

Trying to run fable-coder-35B-A3B-Q4_K_M.gguf and it fails to load with:

llama_model_loader: loading model tensors, this can take a while... (mmap = true, direct_io = false)
llama_model_load: error loading model: missing tensor 'blk.40.ssm_conv1d.weight'
llama_model_load_from_file_impl: failed to load model

Setup:

  • llama-server build b9075-4f331667d
  • --n-gpu-layers 999 --n-cpu-moe 32 (same flags I use successfully for other Qwen3.6-35B-A3B GGUFs)
  • Downloaded via hf download, verified the local file's SHA256 matches the repo's LFS hash exactly (f88b25416bff442627fcbcc04d3191b83922ac394575c66f3c0813bc6fe06b92), so it's not a corrupted/incomplete download.

For context: I run the base Qwen3.6-35B-A3B architecture (hybrid MoE + SSM/Gated-DeltaNet layers) all the time on this exact same llama-server build with no issues, so my binary does have working hybrid-SSM support. This looks like layer 40's ssm_conv1d weight may have been dropped somewhere in the LoRA merge / GGUF conversion pipeline for this specific model.

Is this:

  1. Something I'm doing wrong on my end (missing flag, wrong build, etc.)?
  2. A known issue with this repo's conversion that's already being looked at?
  3. A new one for you to dig into?

Happy to share more logs/repro details if useful. Thanks for the model!

Hey my man i appreciate the question

On the missing-tensor theory: blk.40 is the MTP / next-token-prediction block, and it's attention-style by design β€” the stock Qwen3.6-35B-A3B GGUF has no ssm_conv1d there either. I read the tensor index straight out of the uploaded Q4_K_M and diffed it against the base:

base Qwen3.6-35B-A3B this repo's Q4_K_M
total tensors 753 753
blk.40 tensors 20 20, identical names
blk.40.ssm_conv1d absent absent
block_count 41 41
nextn_predict_layers 1 1
SSM blocks 30 (0,1,2,4,5,6,…) 30, same pattern

The hybrid pattern puts full-attention layers at blocks 3, 7, 11 … 39, with block 40 as the MTP block on top of the 40 real layers. So nothing was dropped in the LoRA merge or the GGUF conversion β€” a reasonable hypothesis, just not what happened. Build 9075 appears to type block 40 as a regular hybrid layer and go looking for ssm_conv1d; the newer build handles it.

That also explains why the base works for you: most base GGUF uploads ship with the MTP block stripped (block_count=40, nextn_predict_layers=0), so there's no block 40 to mis-type. This repo keeps it, which needs a newer loader. If you want to confirm, on the base that works for you:

python -c "from gguf import GGUFReader as G; r=G('base.gguf'); \
print({k: v.contents() for k, v in r.fields.items() if 'block_count' in k or 'nextn' in k})"

Sign up or log in to comment