Zamba2-VL-2.7B โ€” GGUF

Our own GGUF conversion of Zyphra's Zamba2-VL-2.7B: the Zamba2 language model with a Qwen2.5-VL vision tower. The converter is ours (llama.cpp #16, #17), including a chat template that places images the way Zyphra's does.

Contents

  • Zamba2-VL-2.7B-Q8_0.gguf: the language model.
  • mmproj-Zamba2-VL-2.7B-F16.gguf: the vision tower.

Validation

  • Vision embeddings against transformers: per-token cosine 0.99989 on the CPU and 0.99944 on Vulkan.
  • On a synthetic image (a red square, a blue circle and "HELLO 42") the model answers "a red square ... a blue circle ... a black text", with the layout right.

Running it

With the 1bit engine:

1bit serve -m Zamba2-VL-2.7B-Q8_0.gguf --mmproj mmproj-Zamba2-VL-2.7B-F16.gguf --device vulkan

With llama.cpp (our fork), pass --jinja so the embedded chat template is used.

Attribution

Downloads last month
60
GGUF
Model size
4B params
Architecture
zamba2
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for 1bit-MONSTER/Zamba2-VL-2.7B-GGUF

Quantized
(1)
this model