YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
llama-cpp-python Precompiled Kernels for AMD gfx1151 (RX 9070 XT / Strix Halo)
Pre-compiled HIP/ROCm kernels for llama-cpp-python on AMD gfx1151 GPUs (RDNA 4).
Contents
comgr/- Precompiled LLVM kernel cache (~94MB)libggml-hip.so- HIP-accelerated GGML library compiled for gfx1151libggml*.so- Supporting GGML librarieslibllama.so- llama.cpp librarylibmtmd.so- Multi-modal library
Installation
- Install llama-cpp-python normally (with HIPBLAS support):
CMAKE_ARGS="-DGGML_HIPBLAS=on" AMDGPU_TARGETS="gfx1151" pip install llama-cpp-python
- Copy the precompiled cache:
# Copy comgr cache (saves ~50 mins of JIT compilation!)
cp -r comgr ~/.cache/
# Optionally replace the compiled libs (must match llama-cpp-python version)
# cp lib*.so ~/.local/lib/python3.x/site-packages/llama_cpp/lib/
Environment Setup
Required environment variables for gfx1151:
export HSA_OVERRIDE_GFX_VERSION=11.5.1
export AMDGPU_TARGETS=gfx1151
Tested Configuration
- GPU: AMD RX 9070 XT (gfx1151, RDNA 4)
- ROCm: 7.x
- llama-cpp-python: 0.3.x
- Python: 3.12
- OS: Linux (Ubuntu 24.04)
Models Tested
- openai/gpt-oss-120b (Q4_K_M, ~60GB VRAM)
- DeepSeek-R1-1.5B
Why This Exists
First-time loading of large GGML models on gfx1151 requires JIT compilation of GPU kernels, which can take 30-60+ minutes. This cache contains pre-compiled kernels that skip that initial compilation step.
Notes
- These kernels are specific to gfx1151 (RDNA 4 architecture)
- They may not work on other GPU architectures (gfx1100, gfx90a, etc.)
- The comgr cache is the most important - it contains the actual compiled GPU code
License
MIT
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support