ExLlamaV3 Wheels for OrcaSAQ2 & Tesla T4 (sm_75)
Precompiled binary wheels of ExLlamaV3 optimized for NVIDIA Tesla T4 (Compute Capability 7.5) and multi-GPU (2x Tesla T4).
Key Features:
- Turing (sm_75) Native Kernels: Built with PR #325 (paired mma.m16n8k8, synchronous copy fallback, 64 KB shared memory budget).
- Fractional Trellis Support (3.5 bpw): Fully supports 3.5 bpw layers used in orcarouter/OrcaSAQ-2-27B without packed dimension 2 is incorrect size errors.
- Continuum int8_embedding Ready: Compatible with Continuum-AI-Corp's packed int8 embedding table.
Quick Install (Google Colab / Kaggle / Linux x86_64):
ash pip install https://huggingface.co/Sigmo23/ext2l-wheels-orca/resolve/main/<wheel_name>.whl
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support