Weight-only INT4 for RTX 3090/4090-class cards, where FP8 and NVFP4 are emulated or unusable. Full recipe on every card.
Alex Αdamopoulos
aleada
·
AI & ML interests
Verified open-weight inference for consumer GPUs. W4A16 packs, reproducible benchmarks and tools that inspect what quantized models actually contain.
Recent Activity
updated a model about 23 hours ago
aleada/Llama-3.2-11B-Vision-Instruct-W4A16 updated a model about 23 hours ago
aleada/Pixtral-12B-W4A16 updated a model about 23 hours ago
aleada/SmolLM3-3B-W4A16Organizations
None yet