flash-attn prebuilt wheels (Linux x86_64)

Byte-identical mirror of wheels published in the official Dao-AILab/flash-attention releases. Nothing here is rebuilt or repackaged: each file is the release asset as downloaded, and its SHA-256 is recorded below so any consumer can prove it.

This exists because some container build environments cannot reach github.com/.../releases/download/... reliably, while huggingface.co is reachable. Pinning a wheel by URL and checksum keeps the resulting image reproducible regardless of which host served the bytes.

Contents

file version built against sha256 size
flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl 2.7.4.post1 CUDA 12.x, torch 2.6, CPython 3.10, _GLIBCXX_USE_CXX11_ABI=0 ffe17686fa1a0f288de9eae7c32af209d32a27b037ef28614f042b377af5b15a 187815087 bytes

Upstream source of that file:

https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1%2Bcu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl

Verifying

sha256sum flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl
# ffe17686fa1a0f288de9eae7c32af209d32a27b037ef28614f042b377af5b15a

License

flash-attention is released by its authors under the BSD 3-Clause license; the wheels carry that license unchanged. This repository redistributes the official binaries and claims no additional rights over them.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support