Instructions to use orbcom-pedroferreira/AMALIA-VL-DPO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use orbcom-pedroferreira/AMALIA-VL-DPO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
Use Docker
docker model run hf.co/orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use orbcom-pedroferreira/AMALIA-VL-DPO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "orbcom-pedroferreira/AMALIA-VL-DPO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "orbcom-pedroferreira/AMALIA-VL-DPO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
- Ollama
How to use orbcom-pedroferreira/AMALIA-VL-DPO-GGUF with Ollama:
ollama run hf.co/orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use orbcom-pedroferreira/AMALIA-VL-DPO-GGUF with Docker Model Runner:
docker model run hf.co/orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
- Lemonade
How to use orbcom-pedroferreira/AMALIA-VL-DPO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull orbcom-pedroferreira/AMALIA-VL-DPO-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.AMALIA-VL-DPO-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
AMALIA-VL-DPO — GGUF Quantizations
GGUF quantizations of amalia-llm/AMALIA-VL-DPO, the open-source vision-language model targeting European Portuguese (pt-PT), developed by a consortium of Portuguese universities and research centres and funded by the Government of Portugal.
The original model uses a LLaVA-NeXT architecture with a SigLIP vision encoder (1152 hidden dim, 384px image size, patch size 16) and a LLaMA-based language backbone (9.15B parameters, 42 layers, 32768 context). The base model was trained in 3 stages: modality alignment, visual instruction following, and direct preference optimization (DPO).
Note: These are community-converted GGUF files. The original model was released by the AMALIA LLM Project under the Apache 2.0 license. All quantizations were produced from the official BF16 safetensors checkpoint converted to F16 GGUF, then quantized using llama.cpp.
Available Files
| File | Size | Precision | Notes |
|---|---|---|---|
mmproj-model-f16.gguf |
~500 MB | F16 | Required for vision — use alongside any quant below |
AMALIA-VL-DPO-Q8_0.gguf |
~9.5 GB | 8-bit | Near-lossless quality |
AMALIA-VL-DPO-Q6_K.gguf |
~7.5 GB | 6-bit | High quality, recommended if VRAM allows |
AMALIA-VL-DPO-Q5_K_M.gguf |
~6.5 GB | 5-bit | Very good quality |
AMALIA-VL-DPO-Q4_K_M.gguf |
~5.5 GB | 4-bit | Recommended — best quality/size balance |
AMALIA-VL-DPO-Q3_K_M.gguf |
~4.5 GB | 3-bit | Smaller, slight quality drop |
AMALIA-VL-DPO-Q2_K.gguf |
~3.5 GB | 2-bit | Smallest, noticeable quality loss |
The
mmproj-model-f16.gguffile is shared across all quantizations — you only need one copy regardless of which quant you run.
Usage
With llama-mtmd-cli (recommended for local inference)
llama-mtmd-cli \
-m AMALIA-VL-DPO-Q4_K_M.gguf \
--mmproj mmproj-model-f16.gguf \
-ngl 99 \
-fa on \
--image /path/to/image.jpg \
-p "Descreve esta imagem em português"
Interactive chat mode (no image or prompt required at launch):
llama-mtmd-cli \
-m AMALIA-VL-DPO-Q4_K_M.gguf \
--mmproj mmproj-model-f16.gguf \
-ngl 99 \
-fa on
With video:
llama-mtmd-cli \
-m AMALIA-VL-DPO-Q4_K_M.gguf \
--mmproj mmproj-model-f16.gguf \
-ngl 99 \
--video /path/to/video.mp4 \
-p "O que acontece neste vídeo?"
With llama-server (OpenAI-compatible API)
llama-server \
-m AMALIA-VL-DPO-Q4_K_M.gguf \
--mmproj mmproj-model-f16.gguf \
-ngl 99 \
-fa on \
--host 0.0.0.0 \
--port 8080
Recommended sampling parameters
The original model uses ChatML format with the following recommended settings:
--chat-template chatml \
--temp 0.7 \
--top-p 0.9 \
--repeat-penalty 1.1
VRAM Requirements
| Quant | Model VRAM | + mmproj | Total (approx.) |
|---|---|---|---|
| Q8_0 | ~9.5 GB | ~0.5 GB | ~10 GB |
| Q6_K | ~7.5 GB | ~0.5 GB | ~8 GB |
| Q5_K_M | ~6.5 GB | ~0.5 GB | ~7 GB |
| Q4_K_M | ~5.5 GB | ~0.5 GB | ~6 GB |
| Q3_K_M | ~4.5 GB | ~0.5 GB | ~5 GB |
| Q2_K | ~3.5 GB | ~0.5 GB | ~4 GB |
Context length also consumes VRAM. The above estimates are for the model weights + mmproj at short context. KV cache adds approximately 1–5 GB depending on context length used.
Conversion Details
The GGUFs were produced from the official amalia-llm/AMALIA-VL-DPO checkpoint using llama.cpp.
Process:
- The LLM backbone was extracted from the LLaVA-NeXT wrapper using
transformersand converted to F16 GGUF viaconvert_hf_to_gguf.py - The SigLIP vision encoder and MLP projector were extracted using
llava_surgery_v2.pyand converted to an mmproj GGUF viaconvert_image_encoder_to_gguf.py --clip-model-is-siglip - All quantizations were produced from the F16 GGUF master using
llama-quantize
Vision encoder specs:
- Architecture: SigLIP (
siglip_vision_model) - Hidden size: 1152
- Image size: 384×384
- Patch size: 16
- Sequence length: 576 tokens per image
About AMALIA-VL
AMALIA-VL is developed by a consortium of Portuguese universities and research centres including NOVA University Lisbon, Instituto Superior Técnico, the University of Coimbra, the University of Porto, the University of Minho, and the Foundation for Science and Technology (FCT), funded by the Government of Portugal.
Training was carried out on the MareNostrum5 supercomputer and the DEUCALION supercomputer using 64 NVIDIA H100 GPUs across three training stages.
For full details, refer to the technical report and the official model card.
Citation
If you use these GGUF files in your work, please cite the original AMALIA-VL paper:
@article{gloria2026amalia,
title={AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model},
author={Glória-Silva, Diogo and Cardeira, João and da Luz, Manuel Letras and Simplício, Afonso and Vinagre, Gonçalo and Tavares, Diogo and Ferreira, Rafael and Calvo, Inês and Vieira, Inês and Semedo, David and others},
journal={arXiv preprint arXiv:2606.19100},
year={2026}
}
- Downloads last month
- 114
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for orbcom-pedroferreira/AMALIA-VL-DPO-GGUF
Base model
amalia-llm/AMALIA-9B-1225-DPO