YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

llama-cpp-python Precompiled Kernels for AMD gfx1151 (RX 9070 XT / Strix Halo)

Pre-compiled HIP/ROCm kernels for llama-cpp-python on AMD gfx1151 GPUs (RDNA 4).

Contents

  • comgr/ - Precompiled LLVM kernel cache (~94MB)
  • libggml-hip.so - HIP-accelerated GGML library compiled for gfx1151
  • libggml*.so - Supporting GGML libraries
  • libllama.so - llama.cpp library
  • libmtmd.so - Multi-modal library

Installation

  1. Install llama-cpp-python normally (with HIPBLAS support):
CMAKE_ARGS="-DGGML_HIPBLAS=on" AMDGPU_TARGETS="gfx1151" pip install llama-cpp-python
  1. Copy the precompiled cache:
# Copy comgr cache (saves ~50 mins of JIT compilation!)
cp -r comgr ~/.cache/

# Optionally replace the compiled libs (must match llama-cpp-python version)
# cp lib*.so ~/.local/lib/python3.x/site-packages/llama_cpp/lib/

Environment Setup

Required environment variables for gfx1151:

export HSA_OVERRIDE_GFX_VERSION=11.5.1
export AMDGPU_TARGETS=gfx1151

Tested Configuration

  • GPU: AMD RX 9070 XT (gfx1151, RDNA 4)
  • ROCm: 7.x
  • llama-cpp-python: 0.3.x
  • Python: 3.12
  • OS: Linux (Ubuntu 24.04)

Models Tested

  • openai/gpt-oss-120b (Q4_K_M, ~60GB VRAM)
  • DeepSeek-R1-1.5B

Why This Exists

First-time loading of large GGML models on gfx1151 requires JIT compilation of GPU kernels, which can take 30-60+ minutes. This cache contains pre-compiled kernels that skip that initial compilation step.

Notes

  • These kernels are specific to gfx1151 (RDNA 4 architecture)
  • They may not work on other GPU architectures (gfx1100, gfx90a, etc.)
  • The comgr cache is the most important - it contains the actual compiled GPU code

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support