Instructions to use defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix") model = AutoModelForCausalLM.from_pretrained("defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix
- SGLang
How to use defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix with Docker Model Runner:
docker model run hf.co/defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix
Access request
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This repository is public metadata with a manual download gate. Submit an accurate intended-use statement. Approval is discretionary and may be revoked for misuse or redistribution without authorization.
Log in or Sign Up to review the conditions and access this model content.
GLM-5.3-DERISKED-Int4-Int8Mix
Client drop · full GLM-5.3 MoE · compressed-tensors pack-quant · Blackfrost DWM a=3
Built by Blackfrost · Las Vegas, Nevada
Public metadata, manual download gate. Client delivery. Request access on the Hub. Do not redistribute the weights without authorization.
What this is
Independent Int4-Int8Mix pack-quantized safetensors of GLM-5.3 (full ~753B MoE, not Flash) after a Blackfrost direction-weight modification (DWM) pass. Source mix weights were not overwritten.
The intended behavior is in the weights. Production DWM details are proprietary and are not disclosed beyond the locked recipe below.
This artifact has not been through a judged refusal suite. Do not copy NVFP4/BF16/Q3 refusal percentages onto this checkpoint.
Specifications
| Architecture | GlmMoeDsaForCausalLM (glm_moe_dsa) |
| Quant | compressed-tensors pack-quantized · int8 group-128 linears · int4 group-128 routed experts · layer 0 BF16 · MTP int8 channel (stock) |
| Shards | 282 safetensors |
| Layers | 78 main + MTP layer 78 |
| Hidden size | 6144 |
| Experts | 256 routed · top-8 / token · 1 shared |
| DWM | alpha 3.0 · skip-early 2 · one pass · frozen scales · no norm restore |
| Targets | 76 o_proj + 1 dense down_proj + 19200 expert down_proj + 75 shared down_proj |
| Runtime | vLLM compressed-tensors / pack-quantized (multi-Spark). Not a GGUF. |
| Languages | English and Chinese |
Sibling client GGUF (separate repo, do not mix formats):
defiobi00/GLM-5.3-DERISKED-UD-Q3_K_XL.
- Downloads last month
- 3
Model tree for defiobi00/GLM-5.3-DERISKED-Int4-Int8Mix
Base model
zai-org/GLM-5.3-BF16