Instructions to use Felipe97/llama-cpp-compiled with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Felipe97/llama-cpp-compiled with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Felipe97/llama-cpp-compiled # Run inference directly in the terminal: llama cli -hf Felipe97/llama-cpp-compiled
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Felipe97/llama-cpp-compiled # Run inference directly in the terminal: llama cli -hf Felipe97/llama-cpp-compiled
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Felipe97/llama-cpp-compiled # Run inference directly in the terminal: ./llama-cli -hf Felipe97/llama-cpp-compiled
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Felipe97/llama-cpp-compiled # Run inference directly in the terminal: ./build/bin/llama-cli -hf Felipe97/llama-cpp-compiled
Use Docker
docker model run hf.co/Felipe97/llama-cpp-compiled
- LM Studio
- Jan
- Ollama
How to use Felipe97/llama-cpp-compiled with Ollama:
ollama run hf.co/Felipe97/llama-cpp-compiled
- Unsloth Desktop
- Docker Model Runner
How to use Felipe97/llama-cpp-compiled with Docker Model Runner:
docker model run hf.co/Felipe97/llama-cpp-compiled
- Lemonade
How to use Felipe97/llama-cpp-compiled with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Felipe97/llama-cpp-compiled
Run and chat with the model
lemonade run user.llama-cpp-compiled-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
Release process
llama.cpp uses semantic versioning (MAJOR.MINOR.PATCH).
Version bump guidelines
| Change type | Version component |
|---|---|
Breaking change to the public C API (include/llama.h) |
MAJOR |
| Backward-compatible features, model support, or API addition | MINOR |
| Bug fix with no API change | PATCH |
The version is set in the three variables at the top of the root CMakeLists.txt:
set(LLAMA_VERSION_MAJOR 0)
set(LLAMA_VERSION_MINOR 1)
set(LLAMA_VERSION_PATCH 0)
A version bump should be included in the PR that introduces the change, or in a dedicated bump commit merged before the release is cut.
TODO: add PR labels (semver: patch, semver: minor, semver: major) to help
identify which PRs require a version bump before cutting a release.
Making a release
Releases are created by running the make-release which is a manual workflow.
The workflow runs against the branch selected in the "Run workflow" dialog
(default master) and takes an optional commit SHA. When a commit is given,
the workflow validates that the commit belongs to the branch and is not older
than 3 days from the branch HEAD, then releases that commit instead of the
branch HEAD.
The workflow creates an annotated git tag (e.g. v0.1.0) and pushes it to the
remote. No GitHub Release object is created, the tag is the release artifact.
Building a release
By default, LLAMA_BUILD_IS_DEV=ON which appends a -dev suffix to LLAMA_VERSION,
marking the build as a nightly/development build. Distributors building from a
release tag must pass -DLLAMA_BUILD_IS_DEV=OFF to produce a clean version string
(e.g. 0.1.0 instead of 0.1.0-dev).
How releases reach users
Currently releases are not published to github releases, only nightly/development builds are available there. The way users can access releases are using the following channels:
- llama-install.sh — downloads pre-built binaries built from the release tag.
- Package managers — consume the git tag directly.
- Build from source — users clone the repo and check out the tag.