Quantized version of: meta-llama/Llama-Guard-4-12B

Quantization Details

  • Language Model Body (Text Body): All Self-Attention (q_proj, k_proj, v_proj, o_proj) and MLP (gate_proj, up_proj, down_proj) layers are successfully compressed to NVFP4.
  • Unquantized (Retained in BF16):
    • lm_head (retained in BF16 to resolve vLLM loading compatibility issues)
    • embed_tokens (token embedding layer)
    • vision_model & multi_modal_projector (vision-related components)
Downloads last month
40
Safetensors
Model size
12B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tengsnake/Llama-Guard-4-12B-NVFP4

Quantized
(2)
this model

Dataset used to train tengsnake/Llama-Guard-4-12B-NVFP4