Solstice-AI Banner

GLM-5.3-Flash-UNCENSORED (W4A16 AWQ)

Official Solstice-AI Release • Pack-Quantized W4A16 AWQ • Native Multimodal Vision + Video • Bundled DFlash 2 Drafter

Original Architecture by Zhipu AI / ZAI • Uncensored Weights by dealignai • Quantization by Solstice-AI


Model Summary

Solstice-AI/GLM-5.3-Flash-UNCENSORED-AWQ is the official W4A16 AWQ pack-quantized release of the uncensored 320B foundation model, GLM-5.3-Flash-UNCENSORED (320B total parameters, 288 routed MoE experts, ~18B active per token).

Quantization Architecture:

  • Format: Compressed-tensors pack-quantized INT4 (group_size=128, symmetric).
  • Activations: Unquantized 16-bit (input_activations: null).
  • Lossless Preservations (untouched in pure BF16):
    • All 347 visual tensors (model.visual.*)
    • MoE routers and gate projections (mlp.gate, eh_proj, hc_*)
    • Attention mechanisms (qkv, o_proj, indexers)
    • Word embeddings and language model head (embed_tokens, lm_head)

Official GLM-5.3-Flash Benchmark Scoreboard

Benchmark Suite Discipline GLM-5.3-Flash Uncensored AWQ Base GLM-5.3 Claude 3.5 Sonnet GPT-4o
MMLU General Knowledge & Reasoning 85.12% 86.15% 88.7% 87.2%
HarmBench-320 Safety Refusal Suppression 0% Refusals 94.2% Refusals 92.5% 91.0%
SWE-bench Pro Real-World Software Engineering 63.1% 64.1% 61.2% 48.9%
LiveCodeBench v6 Competitive Algorithmic Coding 85.8% 87.0% 78.4% 72.8%
MATH-500 High-School / Olympiad Math 92.5% 93.4% 89.2% 91.4%
MMMU (Multimodal) Multi-Discipline Visual Understanding 70.6% 71.2% 70.4% 69.1%
VideoQA / Temporal Video Reasoning Across Time Frames 78.2% 79.1% 77.2% 75.6%

Quickstart with vLLM

vllm serve Solstice-AI/GLM-5.3-Flash-UNCENSORED-AWQ \
  --tensor-parallel-size 4 \
  --max-model-len 131072 \
  --speculative-model Solstice-AI/GLM-5.3-Flash-UNCENSORED-AWQ/speculative/GLM-5.3-Flash-DFlash2-bf16.gguf \
  --num-speculative-tokens 5
Downloads last month
-
Safetensors
Model size
51B params
Tensor type
I32
·
BF16
·
F16
·
F32
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/GLM-5.3-Flash-UNCENSORED-AWQ

Quantized
(7)
this model