Solstice-AI Banner

GLM-5.3-Flash-UNCENSORED (NVFP4 W4A16)

Official Solstice-AI W4A16 NVFP4 Release • Native Multimodal Vision + Video • 1M Context Window (1,048,576 Tokens) • Bundled DFlash 2 Speculative Drafter

Original Architecture by Zhipu AI / ZAI • Uncensored Weights by dealignai • NVFP4 W4A16 Packaging & Curation by Solstice-AI

Solstice-AI License Format Precision Context Hardware


Model Summary

Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4 is the official W4A16 NVFP4 mixed-precision release of the 320B foundation model, GLM-5.3-Flash-UNCENSORED (320B total parameters, 288 routed MoE experts, ~18B active per token).

Key Architectural Highlights:

  • W4A16 Mixed-Precision: Routed MoE experts quantized to NVFP4 (4-bit float, e2m1) while all attention layers, shared experts, and input activations remain strictly in full 16-bit (BF16/FP16). Zero activation clipping degradation!
  • Universal GPU Support: Optimized for NVIDIA Blackwell B200 / GB200, Hopper H100/H200, Ada Lovelace RTX 4090/L40S, and Ampere A100.
  • Native Multimodal Vision + Video: Full 24-layer ViT (glm5_next_vision, hidden size 1024) and 10,240-dim projector preserved byte-for-byte in original precision. Handles high-resolution images and temporal video sequences.
  • Weight-Level Uncensored: Refusal directions completely ablated at the weight level (0% refusals on HarmBench-320, MMLU 85.28% preserved).
  • Native 1M Context Window: 1,048,576 tokens native context.
  • Bundled DFlash 2 Speculative Drafter: Pre-packaged in the speculative/ folder (GLM-5.3-Flash-DFlash2-bf16.gguf & Q8_0.gguf) for 2x–3x generation throughput.

Official GLM-5.3-Flash Benchmark Scoreboard

Benchmark Suite Discipline GLM-5.3-Flash Uncensored NVFP4 Base GLM-5.3 Claude 3.5 Sonnet GPT-4o
MMLU General Knowledge & Reasoning 85.28% 86.15% 88.7% 87.2%
HarmBench-320 Safety Refusal Suppression 0% Refusals 94.2% Refusals 92.5% 91.0%
SWE-bench Pro Real-World Software Engineering 63.4% 64.1% 61.2% 48.9%
LiveCodeBench v6 Competitive Algorithmic Coding 86.1% 87.0% 78.4% 72.8%
MATH-500 High-School / Olympiad Math 92.8% 93.4% 89.2% 91.4%
MMMU (Multimodal) Multi-Discipline Visual Understanding 70.8% 71.2% 70.4% 69.1%
VideoQA / Temporal Video Reasoning Across Time Frames 78.5% 79.1% 77.2% 75.6%

Serving Quickstart

1. High-Throughput Serving with vLLM (W4A16 Mode)

vllm serve Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4   --quantization modelopt   --tensor-parallel-size 2   --trust-remote-code   --max-model-len 131072   --gpu-memory-utilization 0.95

2. Speculative Decoding with SGLang + DFlash 2

python3 -m sglang.launch_server   --model-path Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4   --speculative-algorithm DFLASH   --speculative-draft-model-path incoai/GLM-5.3-Flash-DFlash2   --tp 2   --trust-remote-code

Speculative Drafter Files Included

In the speculative/ directory of this repo:

  • speculative/GLM-5.3-Flash-DFlash2-bf16.gguf (Pure BF16 block-diffusion draft head)
  • speculative/GLM-5.3-Flash-DFlash2-Q8_0.gguf (Q8_0 quantized block-diffusion draft head)

License & Attribution

Downloads last month
281
Safetensors
Model size
165B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4-DSpark

Quantized
(7)
this model