Full-text search
Search in
Scope to owner or repo
3,873 results
NVFP4 / Polaris-7B-Preview-FP4
README.md
model
1 matches
NVFP4 / Qwen3-Coder-480B-A35B-Instruct-FP4
README.md
model
1 matches
NVFP4 / DeepSeek-R1-0528-Qwen3-8B-FP4
README.md
model
1 matches
NVFP4 / Qwen3-32B-FP4
README.md
model
1 matches
NVFP4 / DeepSeek-Prover-V2-7B-FP4
README.md
model
1 matches
NVFP4 / Polaris-4B-Preview-FP4
README.md
model
1 matches
NVFP4 / Qwen3-235B-A22B-Instruct-2507-FP4
README.md
model
1 matches
NVFP4 / Qwen3-235B-A22B-Thinking-2507-FP4
README.md
model
1 matches
NVFP4 / Qwen3-30B-A3B-Thinking-2507-FP4
README.md
model
1 matches
NVFP4 / Qwen3-30B-A3B-Instruct-2507-FP4
README.md
model
1 matches
NVFP4 / Qwen3-Coder-30B-A3B-Instruct-FP4
README.md
model
1 matches
NVFP4 / Qwen3-0.6B-FP4
README.md
model
1 matches
NVFP4 / LTX2.3-NVFP4-Diffusers
model
1 matches
Henley04 / SoulX-Singer-nvfp4
README.md
model
44 matches
tags: nvfp4, SoulX-Singer, svs, text-to-speech, zh, en, arxiv:2602.07803, base_model:Soul-AILab/SoulX-Singer, base_model:finetune:Soul-AILab/SoulX-Singer, license:apache-2.0, region:us
14
# SoulX-Singer NVFP4 (Quantized)
15
16
This is the **NVIDIA FP4 (NVFP4)** weight-only quantized version of the original
⋯
19
It is produced by a one-time `torchao` NVFP4 quantization pass over the official
⋯
29
- **Tooling:** `torchao>=0.17.0` (`NVFP4WeightOnlyConfig`)
⋯
33
## Why NVFP4
34
35
NVFP4 is the **native 4-bit floating-point format** introduced with the
36
Blackwell architecture. Compared with INT4 / W4A16 pseudo-quantization,
37
NVFP4 has dedicated hardware units on sm100+ and runs the **actual 4-bit
⋯
45
quantized to NVFP4. The following tensors stay in their original precision:
sroecker / Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4
README.md
model
6 matches
tags: safetensors, qwen3_5_moe, qwen, nvfp4, vllm, compressed-tensors, base_model:RangerX/Qwen3.6-35B-REAP-Pruned-ratio-0.5, base_model:quantized:RangerX/Qwen3.6-35B-REAP-Pruned-ratio-0.5, 8-bit, region:us
13
# NVFP4 Quantized Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4
14
15
This is an NVFP4 quantized version of
16
`RangerX/Qwen3.6-35B-REAP-Pruned-ratio-0.5` with `llm-compressor`.
17
The model has both weights and activations quantized to NVFP4 format in
⋯
26
Tested on an NVIDIA GeForce RTX 5070 Ti 16 GiB with vLLM's NVFP4 linear path
⋯
164
scheme="NVFP4",
prithivMLmods / Q3.6-27B-GLM-5.1-DA-nvfp4
README.md
model
2 matches
tags: transformers, safetensors, qwen3_5, image-text-to-text, text-generation-inference, vllm, glm-5.1, math, reasoning, pytorch, uncensored, abliterated, unfiltered, unredacted, refusal-ablated, alignment-modified, nvfp4, conversational, en, dataset:prithivMLmods/harm_bench, dataset:Jackrong/GLM-5.1-Reasoning-1M-Cleaned, base_model:prithivMLmods/Q3.6-27B-GLM-5.1-DA, base_model:quantized:prithivMLmods/Q3.6-27B-GLM-5.1-DA, license:apache-2.0, endpoints_compatible, compressed-tensors, region:us
28
# **NVFP4 — Q3.6-27B-GLM-5.1-DA**
29
30
> bf16 : https://huggingface.co/prithivMLmods/Q3.6-27B-GLM-5.1-DA
dryade36513 / 10Eros_v1_experimental_nvfp4
README.md
model
2 matches
ApacheOne / ZImageTurbo-nvfp4_mixed
README.md
model
3 matches
tags: custom, quantization, nvfp4, art, en, base_model:Tongyi-MAI/Z-Image-Turbo, base_model:quantized:Tongyi-MAI/Z-Image-Turbo, license:apache-2.0, region:us
14
## nvfp4 model infomation
15
Note: This model has not been personally tested but should work in ComfyUI. Feedback is appreciated.
⋯
30
Note: Custom FP32 mixed nvfp4 version of base ZImageturbo
ApacheOne / ZImageBase-nvfp4_mixed
README.md
model
3 matches
ItBitter / ZImageTurbo-nvfp4_mixed
README.md
model
3 matches
tags: custom, quantization, nvfp4, art, en, base_model:Tongyi-MAI/Z-Image-Turbo, base_model:quantized:Tongyi-MAI/Z-Image-Turbo, license:apache-2.0, region:us
14
## nvfp4 model infomation
15
Note: This model has not been personally tested but should work in ComfyUI. Feedback is appreciated.
⋯
30
Note: Custom FP32 mixed nvfp4 version of base ZImageturbo