flash-attn prebuilt wheels (Linux x86_64)
Byte-identical mirror of wheels published in the official Dao-AILab/flash-attention releases. Nothing here is rebuilt or repackaged: each file is the release asset as downloaded, and its SHA-256 is recorded below so any consumer can prove it.
This exists because some container build environments cannot reach
github.com/.../releases/download/... reliably, while huggingface.co is reachable.
Pinning a wheel by URL and checksum keeps the resulting image reproducible
regardless of which host served the bytes.
Contents
| file | version | built against | sha256 | size |
|---|---|---|---|---|
flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl |
2.7.4.post1 | CUDA 12.x, torch 2.6, CPython 3.10, _GLIBCXX_USE_CXX11_ABI=0 |
ffe17686fa1a0f288de9eae7c32af209d32a27b037ef28614f042b377af5b15a |
187815087 bytes |
Upstream source of that file:
https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1%2Bcu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl
Verifying
sha256sum flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl
# ffe17686fa1a0f288de9eae7c32af209d32a27b037ef28614f042b377af5b15a
License
flash-attention is released by its authors under the BSD 3-Clause license; the wheels carry that license unchanged. This repository redistributes the official binaries and claims no additional rights over them.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support