llama-cpp-python 0.3.36 (CUDA 13) prebuilt wheel

Prebuilt llama-cpp-python with CUDA, so you can install in seconds instead of compiling for 10-20 minutes. Made to run unsloth/embeddinggemma-2-GGUF (architecture gemma-embedding2), which the PyPI build cannot load.

Install

!pip install https://huggingface.co/vinamrag-code/llama-cpp-python-cuda-wheels/resolve/main/llama_cpp_python-0.3.36%2Bcu130.gemma2-py3-none-linux_x86_64.whl

Note: the + in the file name is written as %2B in the URL. llama_cpp.__version__ prints 0.3.36; pip show llama-cpp-python shows 0.3.36+cu130.gemma2.

Works on

  • Linux x86_64
  • Python 3.13 (tested; the wheel is tagged py3, other 3.x versions are untested)
  • CUDA 13 runtime (built with CUDA 13.0)
  • GPUs with compute capability 7.5 (Tesla T4, as in Google Colab)

Does not work on

  • Windows, macOS, ARM
  • Other GPU generations (built only with -DCMAKE_CUDA_ARCHITECTURES=75)

What is inside

  • llama-cpp-python 0.3.36
  • llama.cpp commit fc9ce6b9d52a8504edcb262abc92737c2289f96c (2026-10-08)
  • Built with CMAKE_ARGS="-DGGML_CUDA=on -DCMAKE_CUDA_ARCHITECTURES=75"
  • Two small fixes to the Python bindings, needed because that llama.cpp commit changed its C structs:
    1. Added the missing moe_cache_size field to llama_context_params (without it, embed() crashes with a NULL pointer error).
    2. embed() now uses llama_model_n_embd_out for the vector size (otherwise embeddinggemma-2 returns 512 numbers instead of 768).

Tested

embeddinggemma-2-BF16.gguf with embedding=True, n_gpu_layers=-1 on a Colab T4: 768-number embeddings, all layers offloaded to the GPU.

Credits

Built from llama-cpp-python and llama.cpp, both MIT licensed.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support