Instructions to use vinamrag-code/llama-cpp-python-cuda-wheels with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use vinamrag-code/llama-cpp-python-cuda-wheels with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="vinamrag-code/llama-cpp-python-cuda-wheels", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
llama-cpp-python 0.3.36 (CUDA 13) prebuilt wheel
Prebuilt llama-cpp-python with CUDA, so you can install in seconds instead of compiling for 10-20 minutes. Made to run unsloth/embeddinggemma-2-GGUF (architecture gemma-embedding2), which the PyPI build cannot load.
Install
Note: the + in the file name is written as %2B in the URL. llama_cpp.__version__ prints 0.3.36; pip show llama-cpp-python shows 0.3.36+cu130.gemma2.
Works on
- Linux x86_64
- Python 3.13 (tested; the wheel is tagged
py3, other 3.x versions are untested) - CUDA 13 runtime (built with CUDA 13.0)
- GPUs with compute capability 7.5 (Tesla T4, as in Google Colab)
Does not work on
- Windows, macOS, ARM
- Other GPU generations (built only with
-DCMAKE_CUDA_ARCHITECTURES=75)
What is inside
- llama-cpp-python 0.3.36
- llama.cpp commit
fc9ce6b9d52a8504edcb262abc92737c2289f96c(2026-10-08) - Built with
CMAKE_ARGS="-DGGML_CUDA=on -DCMAKE_CUDA_ARCHITECTURES=75" - Two small fixes to the Python bindings, needed because that llama.cpp commit changed its C structs:
- Added the missing
moe_cache_sizefield tollama_context_params(without it,embed()crashes with a NULL pointer error). embed()now usesllama_model_n_embd_outfor the vector size (otherwise embeddinggemma-2 returns 512 numbers instead of 768).
- Added the missing
Tested
embeddinggemma-2-BF16.gguf with embedding=True, n_gpu_layers=-1 on a Colab T4: 768-number embeddings, all layers offloaded to the GPU.
Credits
Built from llama-cpp-python and llama.cpp, both MIT licensed.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support