Instructions to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use GCSA-AiLab/GLM-5.3-Flash-Uncensored-MLX with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("GCSA-AiLab/GLM-5.3-Flash-Uncensored-MLX") config = load_config("GCSA-AiLab/GLM-5.3-Flash-Uncensored-MLX") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
GLM-5.3-Flash-Uncensored ยท MLX
The intervention source is GLM-5.3-Flash-Ablitered2 (LoRA v2). The files here are complete MLX model weights, not standalone LoRA adapters; do not apply v2 again.
These releases are intended for controlled safety research and red-teaming. Reducing refusal behavior also weakens a safety boundary; read the limitations and disclaimer before use.
Repository contents
.
โโโ README.md
โโโ 8bit/ # v2-merged MLX affine 8-bit model
โโโ 4bit/ # v2-merged MLX affine 4-bit model
The directories are alternative complete model formats, not parts to load together.
Model summary
| Item | Value |
|---|---|
| Base model | zai-org/GLM-5.3-Flash |
| Intervention | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| Released formats | MLX affine 8-bit and 4-bit |
| Quantization group size | 128 |
| Source adapter rank / alpha | r=1 / lora_alpha=1 |
| Main effective target | Routed-expert down_proj |
MLX conversion
| Directory | Format |
|---|---|
8bit/ |
MLX 8-bit affine weights, group size 128. |
4bit/ |
MLX 4-bit affine weights, group size 128. |
The original block-FP8 tensors were dequantized before MLX quantization. Token embeddings and the language-model head remain BF16; affine scales and biases are FP16. Quantization can change behavior. The consuming MLX runtime must support the GLM5-Next architecture; compatibility should be checked for the exact runtime and version used.
Evaluation
The reference LoRA card evaluated the v2 adapter attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge:
| Reference metric | Prompts | LoRA v2 |
|---|---|---|
| SimpleSafetyTests full refusal | 100 | 5.00% |
| SimpleSafetyTests partial refusal | 100 | 14.00% |
| StrongREJECT rubric mean | 180 | 0.972222 |
| StrongREJECT refusal rate | 180 | 1.67% |
Neither MLX variant was tested in that evaluation. MLX conversion, quantization, and serving changes can alter behavior. A higher StrongREJECT rubric means more specific assistance with harmful requests, not better general quality or safety. Warnings were counted separately from refusals; automated labels may be wrong.
Usage
Download one variant, for example:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-MLX \
--include "8bit/*" --local-dir ./glm53-mlx
Load ./glm53-mlx/8bit with an MLX runtime that supports GLM5-Next. For the 4-bit variant, download 4bit/* and load that directory instead. Do not attach the original LoRA. Verify the chosen runtime's architecture and quantization support before deployment.
Limitations
- Reduced refusal does not guarantee correctness, harmlessness, or improved general ability. Some refusals may remain.
- The reference evaluation did not test either MLX variant. Results from the adapter on NVFP4 must not be presented as MLX measurements.
- MLX architecture support varies by runtime and version; prompts, sampling, and reasoning effort also affect behavior.
Disclaimer
These weights are for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and all relevant model and dependency terms. The base model's MIT license applies; no warranty is provided for outputs or downstream use.
GLM-5.3-Flash-Uncensored ยท MLX๏ผไธญๆ๏ผ
ไปๅบๆไปถ็ปๆ
.
โโโ README.md
โโโ 8bit/ # ๅทฒๅๅนถ v2 ็ MLX ไปฟๅฐ 8-bit ๆจกๅ
โโโ 4bit/ # ๅทฒๅๅนถ v2 ็ MLX ไปฟๅฐ 4-bit ๆจกๅ
ไธคไธช็ฎๅฝๆฏๅฏๅๅซไฝฟ็จ็ๆ ผๅผ๏ผไธ้่ฆๆผๆฅๅ ่ฝฝใ
ๆจกๅๆฆ่ฆ
| ้กน็ฎ | ๅ ๅฎน |
|---|---|
| ๅบ็กๆจกๅ | zai-org/GLM-5.3-Flash |
| ๅนฒ้ขๆฅๆบ | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| ๅๅธๆ ผๅผ | MLX ไปฟๅฐ 8-bitใ4-bit |
| ้ๅ group size | 128 |
| ๆฅๆบ LoRA ็งฉ / alpha | r=1 / lora_alpha=1 |
| ไธป่ฆๆๆ็ฎๆ | ่ทฏ็ฑไธๅฎถ down_proj |
MLX ่ฝฌๆข
| ็ฎๅฝ | ๆ ผๅผ |
|---|---|
8bit/ |
MLX ไปฟๅฐ 8-bit๏ผgroup size 128ใ |
4bit/ |
MLX ไปฟๅฐ 4-bit๏ผgroup size 128ใ |
ๅๅงๅๅ FP8 ๅผ ้ๅ ๅ้ๅ๏ผๅ่ฟ่ก MLX ้ๅใ่ฏๅตๅ ฅๅ LM head ไฟ็ BF16๏ผไปฟๅฐ scale ไธ bias ไฟๅญไธบ FP16ใ้ๅๅฏ่ฝๆนๅ่กไธบใไฝฟ็จๆถ้กปๆ ธๅฏนๅ ทไฝ MLX ่ฟ่กๆถๅ็ๆฌๅฏน GLM5-Next ็ๆฏๆใ
่ฏๆต็ปๆ
LoRA ๅ่ๆจกๅๅก็่ฏๆตๆฏๅจ vLLM ไธญๆๅๅง v2 ้้
ๅจๅ ่ฝฝๅฐ RedHatAI/GLM-5.3-Flash-NVFP4 ๅ่ฟ่ก็๏ผ่ฎพ็ฝฎไฝๆ่ๅผบๅบฆ๏ผๅนถไฝฟ็จ่ชๅจ่ฃๅค deepseek-v4-flashใ
| ๅ่ๆๆ | ๆ ทๆฌๆฐ | LoRA v2 |
|---|---|---|
| SimpleSafetyTests ๅฎๅ จๆ็ป็ | 100 | 5.00% |
| SimpleSafetyTests ้จๅๆ็ป็ | 100 | 14.00% |
| StrongREJECT rubric ๅๅ | 180 | 0.972222 |
| StrongREJECT ๆ็ป็ | 180 | 1.67% |
ๆฌไปๅบไธคไธช MLX ็ๆฌๅๆชๅๅ ไธ่ฟฐ่ฏๆตใ MLX ่ฝฌๆขใ้ๅไธ้จ็ฝฒๆนๅผๅฏ่ฝๆนๅ็ปๆใStrongREJECT rubric ่ถ้ซ๏ผ่กจ็คบๅฏนๆๅฎณ่ฏทๆฑ็ๅธฎๅฉ่ถๅ ทไฝ๏ผไธไปฃ่กจ้็จ่ดจ้ๆๅฎๅ จๆง่ถ้ซ๏ผ่ญฆๅไธๆ็ปๅๅซ็ป่ฎก๏ผ่ชๅจ่ฃๅคไนๅฏ่ฝ่ฏฏๅคใ
ไฝฟ็จๆนๆณ
ๆ้ไธ่ฝฝไธไธชๆ ผๅผ๏ผไพๅฆ๏ผ
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-MLX \
--include "8bit/*" --local-dir ./glm53-mlx
็จๆฏๆ GLM5-Next ็ MLX ่ฟ่กๆถๅ ่ฝฝ ./glm53-mlx/8bitใๅฆ้ 4-bit๏ผๆนไธบไธ่ฝฝ 4bit/* ๅนถๅ ่ฝฝๅฏนๅบ็ฎๅฝใไธ่ฆๅๆฌกๅ ๅ LoRA๏ผ้จ็ฝฒๅ้กปๆ ธๅฏนๆถๆๅ้ๅๅ
ผๅฎนๆงใ
ๅฑ้ๆง
- ้ไฝๆ็ปไธไฟ่ฏๆญฃ็กฎใๆ ๅฎณๆ้็จ่ฝๅๆ้ซ๏ผไปๅฏ่ฝๅ็ๆ็ปใ
- ๆฅๆบ่ฏๆตๆชๆต่ฏไธคไธช MLX ็ๆฌ๏ผไธ่ฝๆ้้ ๅจๅจ NVFP4 ไธ็็ปๆๅฝไฝ MLX ๅฎๆตใ
- MLX ๆถๆๆฏๆๅ ่ฟ่กๆถๅ็ๆฌ่ๅผ๏ผๆ็คบ่ฏใ้ๆ ทๅๆ่ๅผบๅบฆไนไผๅฝฑๅ่กไธบใ
ๅ ่ดฃๅฃฐๆ
ๆฌๆจกๅไป ไพๅๆณ็ ็ฉถใๅฎๅ จ่ฏไผฐใ็บข้ๆต่ฏๅๅ ถไปๅ่ง็จ้ใๅฎไผๆๆๅๅผฑๆ็ป่กไธบ๏ผๅฏ่ฝ็ๆไธๅฎๅ จใ่ฟๆณใๆฌบ้ชใไปๆจ็ญๆๅฎณๅ ๅฎนใ่ฏทๅฟๅจ็ผบๅฐ่ฎฟ้ฎๆงๅถใ็ๆงใๅ ๅฎน่ฟๆปคใ้็้ๅถไธไบบๅทฅ็็ฃๆถๅไธๅฏไฟก็จๆทๅผๆพใไฝฟ็จ่ ้กป้ตๅฎ้็จๆณๅพใๅนณๅฐๆฟ็ญๅ็ธๅ ณๆจกๅๅไพ่ต้กน็ๆกๆฌพใๆฌไปๅบๆฒฟ็จๅบ็กๆจกๅ็ MIT ่ฎธๅฏ่ฏ๏ผๅฏน่พๅบไธไธๆธธไฝฟ็จไธไฝไฟ่ฏใ
Quantized
Model tree for GCSA-AiLab/GLM-5.3-Flash-Uncensored-MLX
Base model
zai-org/GLM-5.3-Flash