Instructions to use 1bit-MONSTER/Zamba2-VL-2.7B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 1bit-MONSTER/Zamba2-VL-2.7B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
Use Docker
docker model run hf.co/1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use 1bit-MONSTER/Zamba2-VL-2.7B-GGUF with Ollama:
ollama run hf.co/1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use 1bit-MONSTER/Zamba2-VL-2.7B-GGUF with Docker Model Runner:
docker model run hf.co/1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
- Lemonade
How to use 1bit-MONSTER/Zamba2-VL-2.7B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 1bit-MONSTER/Zamba2-VL-2.7B-GGUF:Q8_0
Run and chat with the model
lemonade run user.Zamba2-VL-2.7B-GGUF-Q8_0
List all available models
lemonade list
- Atomic Chat
Zamba2-VL-2.7B โ GGUF
Our own GGUF conversion of Zyphra's Zamba2-VL-2.7B: the Zamba2 language model with a Qwen2.5-VL vision tower. The converter is ours (llama.cpp #16, #17), including a chat template that places images the way Zyphra's does.
Contents
Zamba2-VL-2.7B-Q8_0.gguf: the language model.mmproj-Zamba2-VL-2.7B-F16.gguf: the vision tower.
Validation
- Vision embeddings against transformers: per-token cosine 0.99989 on the CPU and 0.99944 on Vulkan.
- On a synthetic image (a red square, a blue circle and "HELLO 42") the model answers "a red square ... a blue circle ... a black text", with the layout right.
Running it
With the 1bit engine:
1bit serve -m Zamba2-VL-2.7B-Q8_0.gguf --mmproj mmproj-Zamba2-VL-2.7B-F16.gguf --device vulkan
With llama.cpp (our fork), pass --jinja so the embedded chat template is used.
Attribution
- Base model: Zyphra/Zamba2-VL-2.7B, Apache 2.0.
- GGUF conversion and validation: the 1bit engine project, with the converter and model code in our llama.cpp fork.
- License: Apache 2.0, inherited from the base model.
- Downloads last month
- 60
Hardware compatibility
Log In to add your hardware
8-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support