Instructions to use Alevetto07/centvrio-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Alevetto07/centvrio-4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Alevetto07/centvrio-4b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Alevetto07/centvrio-4b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Alevetto07/centvrio-4b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Alevetto07/centvrio-4b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Alevetto07/centvrio-4b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Alevetto07/centvrio-4b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Alevetto07/centvrio-4b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Alevetto07/centvrio-4b:Q4_K_M
Use Docker
docker model run hf.co/Alevetto07/centvrio-4b:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Alevetto07/centvrio-4b with Ollama:
ollama run hf.co/Alevetto07/centvrio-4b:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Alevetto07/centvrio-4b with Docker Model Runner:
docker model run hf.co/Alevetto07/centvrio-4b:Q4_K_M
- Lemonade
How to use Alevetto07/centvrio-4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Alevetto07/centvrio-4b:Q4_K_M
Run and chat with the model
lemonade run user.centvrio-4b-Q4_K_M
List all available models
lemonade list
- Atomic Chat
CENTVRIO 4B — Latin → Italian literary translation
A QLoRA fine-tune of Qwen3.5-4B for literary Latin → Italian translation, quantised
to GGUF for local inference. Part of the Centvrio project (see also
Alevetto07/legionarius-08b, the
0.8B sibling that runs in the browser).
Verified from the GGUF metadata itself:
| Architecture | qwen35 |
| Parameters | 4.3 B |
| Native context length | 262,144 |
| Embedding length | 2560 |
| Fine-tuning | QLoRA (Unsloth) |
| Training data | 46,095 Latin→Italian pairs |
⚠️ Read this before quoting any score
Two evaluation sets exist for this model and they are not comparable:
- Teacher-style test (64 pairs) — references were produced by the same 7B teacher the students were distilled from. This measures similarity to that teacher's style and is close to a self-similarity score. It reads 57–60 chrF2.
- Human-held-out test — a published human Italian translation of Pliny, Epistulae I.6 (106 Latin words), never in training. Everything drops to ~35.
The reason is not a defect in the students: the 7B teacher itself scores 34.88 on the human test, i.e. below every 4B quant here. When the reference changes, the teacher's advantage disappears — which is exactly what you would expect if the first table was measuring style matching rather than translation quality.
Do not put the two columns side by side. The 57–60 figures are not human-quality scores.
Quantisations
| File | Size | Teacher-style test (n=64) | Human-held-out (n=1 doc, 3 samples) |
|---|---|---|---|
centvrio-4b-q3_k_m.gguf |
2211 MB | 57.19 chrF2 | 35.26 chrF2 · LaBSE 0.8704 |
centvrio-4b-q4_k_m.gguf |
2655 MB | 58.96 chrF2 | 35.78 chrF2 · LaBSE 0.8975 |
centvrio-4b-q5_k_m.gguf |
3015 MB | 59.99 chrF2 | 35.49 chrF2 · LaBSE 0.8929 |
centvrio-4b-q6_k.gguf |
3398 MB | 59.82 chrF2 | 36.92 chrF2 · LaBSE 0.8938 |
centvrio-4b-q8_0.gguf |
4397 MB | 59.70 chrF2 | 36.89 chrF2 · LaBSE 0.8922 |
(Locally these were built as qwen3.5-4b-la-it-v2-<quant>.gguf; the repo uses the
same centvrio-4b-<quant>.gguf convention as the 0.8B repo. Bytes are identical.)
On human-held-out data the quants are statistically tied: within-model sampling noise across 3 samples is 2.66 chrF2, while the whole q3→q8 spread is 4.25 with overlapping ranges. Q3_K_M is the rational default — same measured quality at 2.2 GB.
Usage
Ollama
ollama pull hf.co/Alevetto07/centvrio-4b:Q3_K_M
ollama run hf.co/Alevetto07/centvrio-4b:Q3_K_M "Gallia est omnis divisa in partes tres."
The model expects this system prompt (the sampler defaults it was trained and
served with are temperature 0.3, top_p 0.9):
Sei un traduttore letterario dal latino all'italiano. Rispondi SOLO con la traduzione italiana, senza spiegazioni ne analisi.
llama.cpp
llama-cli -m qwen3.5-4b-la-it-v2-q3_k_m.gguf \
-p "Sei un traduttore letterario dal latino all'italiano. Rispondi SOLO con la traduzione italiana.\n\nRidebis, et licet rideas..." \
-n 768 --temp 0.3 --top-p 0.9
Known limitations
- Long inputs are not truncated by the model, but the serving default is. The
shipped endpoint uses
num_predict 768; no model here hit that cap on a 106-word passage, but a longer document can. - Real error modes observed on the human test include hallucinated nouns (apros "boars" → "conigli" "rabbits"), a reversed comparative in the closing clause of the Pliny letter, an untranslated Latin word left in the output, and occasional person/agreement slips (sedebam 1sg rendered as 3sg).
- The human-held-out set is currently one document. Treat the ranking as a hypothesis until more human pairs are added.
- Instruction-following is narrow: this model translates. It is not a general assistant.
Provenance — stated precisely
- Base model:
unsloth/Qwen3.5-4B(Apache-2.0). - Training: QLoRA via Unsloth 2026.9.2, Transformers 5.5.0, on 46,095 pairs.
- Adapter rank for the
v2weights could not be recovered from the training artefacts. The directories named*-r32/*-r64contain LoRA adapters whose configs declareQwen/Qwen3.5-0.8Bas their base model — those belong to 0.8B runs that reused the same output path. The only genuine 4B adapter on disk isr=16, alpha=32. Treat the v2 recipe as unconfirmed rather than assumed. Q4_K_Min the sizes column above refers to the file, not the adapter.
Licence
Apache-2.0, inherited from the base model. Training data is a mixture of curated parallel text and machine-generated material; check the project notes before redistributing for commercial use.
- Downloads last month
- 138
3-bit
4-bit
5-bit
6-bit
8-bit