CENTVRIO 4B — Latin → Italian literary translation

A QLoRA fine-tune of Qwen3.5-4B for literary Latin → Italian translation, quantised to GGUF for local inference. Part of the Centvrio project (see also Alevetto07/legionarius-08b, the 0.8B sibling that runs in the browser).

Verified from the GGUF metadata itself:

Architecture qwen35
Parameters 4.3 B
Native context length 262,144
Embedding length 2560
Fine-tuning QLoRA (Unsloth)
Training data 46,095 Latin→Italian pairs

⚠️ Read this before quoting any score

Two evaluation sets exist for this model and they are not comparable:

  1. Teacher-style test (64 pairs) — references were produced by the same 7B teacher the students were distilled from. This measures similarity to that teacher's style and is close to a self-similarity score. It reads 57–60 chrF2.
  2. Human-held-out test — a published human Italian translation of Pliny, Epistulae I.6 (106 Latin words), never in training. Everything drops to ~35.

The reason is not a defect in the students: the 7B teacher itself scores 34.88 on the human test, i.e. below every 4B quant here. When the reference changes, the teacher's advantage disappears — which is exactly what you would expect if the first table was measuring style matching rather than translation quality.

Do not put the two columns side by side. The 57–60 figures are not human-quality scores.

Quantisations

File Size Teacher-style test (n=64) Human-held-out (n=1 doc, 3 samples)
centvrio-4b-q3_k_m.gguf 2211 MB 57.19 chrF2 35.26 chrF2 · LaBSE 0.8704
centvrio-4b-q4_k_m.gguf 2655 MB 58.96 chrF2 35.78 chrF2 · LaBSE 0.8975
centvrio-4b-q5_k_m.gguf 3015 MB 59.99 chrF2 35.49 chrF2 · LaBSE 0.8929
centvrio-4b-q6_k.gguf 3398 MB 59.82 chrF2 36.92 chrF2 · LaBSE 0.8938
centvrio-4b-q8_0.gguf 4397 MB 59.70 chrF2 36.89 chrF2 · LaBSE 0.8922

(Locally these were built as qwen3.5-4b-la-it-v2-<quant>.gguf; the repo uses the same centvrio-4b-<quant>.gguf convention as the 0.8B repo. Bytes are identical.)

On human-held-out data the quants are statistically tied: within-model sampling noise across 3 samples is 2.66 chrF2, while the whole q3→q8 spread is 4.25 with overlapping ranges. Q3_K_M is the rational default — same measured quality at 2.2 GB.

Usage

Ollama

ollama pull hf.co/Alevetto07/centvrio-4b:Q3_K_M
ollama run hf.co/Alevetto07/centvrio-4b:Q3_K_M "Gallia est omnis divisa in partes tres."

The model expects this system prompt (the sampler defaults it was trained and served with are temperature 0.3, top_p 0.9):

Sei un traduttore letterario dal latino all'italiano. Rispondi SOLO con la traduzione italiana, senza spiegazioni ne analisi.

llama.cpp

llama-cli -m qwen3.5-4b-la-it-v2-q3_k_m.gguf \
  -p "Sei un traduttore letterario dal latino all'italiano. Rispondi SOLO con la traduzione italiana.\n\nRidebis, et licet rideas..." \
  -n 768 --temp 0.3 --top-p 0.9

Known limitations

  • Long inputs are not truncated by the model, but the serving default is. The shipped endpoint uses num_predict 768; no model here hit that cap on a 106-word passage, but a longer document can.
  • Real error modes observed on the human test include hallucinated nouns (apros "boars" → "conigli" "rabbits"), a reversed comparative in the closing clause of the Pliny letter, an untranslated Latin word left in the output, and occasional person/agreement slips (sedebam 1sg rendered as 3sg).
  • The human-held-out set is currently one document. Treat the ranking as a hypothesis until more human pairs are added.
  • Instruction-following is narrow: this model translates. It is not a general assistant.

Provenance — stated precisely

  • Base model: unsloth/Qwen3.5-4B (Apache-2.0).
  • Training: QLoRA via Unsloth 2026.9.2, Transformers 5.5.0, on 46,095 pairs.
  • Adapter rank for the v2 weights could not be recovered from the training artefacts. The directories named *-r32 / *-r64 contain LoRA adapters whose configs declare Qwen/Qwen3.5-0.8B as their base model — those belong to 0.8B runs that reused the same output path. The only genuine 4B adapter on disk is r=16, alpha=32. Treat the v2 recipe as unconfirmed rather than assumed.
  • Q4_K_M in the sizes column above refers to the file, not the adapter.

Licence

Apache-2.0, inherited from the base model. Training data is a mixture of curated parallel text and machine-generated material; check the project notes before redistributing for commercial use.

Downloads last month
138
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Alevetto07/centvrio-4b

Finetuned
Qwen/Qwen3.5-4B
Quantized
(25)
this model