YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators).

Founder: Ilia Bolotnikov
Organization: AMAImedia.com
X (Twitter): @AMAImediacom
LinkedIn: Ilia Bolotnikov
Telegram: @djbionicl
NOESIS version: v16.1
Release date: 2026-08

Qwopus3.5-9B-Translate-v3.5-BF16

Role: Dedicated translator (9B). Think-model — for deterministic translation use a closed-think prefill (<think>\n\n</think>\n\n after assistant\n) to suppress reasoning, otherwise it emits a reasoning trace instead of the translation. BF16 = PRIMARY. Sibling: -GGUF-Q4_K_M.gguf (5.24 GB, fits 6GB GPU).

Test results (2026-06-17) — Q4_K_M, FLORES devtest n=20, no-think

Direction chrF++ BLEU
eng→rus 54.9 25.9
eng→cmn 32.8 7.2
AVG 43.8 16.5

Sample (eng→rus): «Теперь у нас есть мыши в возрасте четырёх месяцев, которые ранее страдали диабетом, но сейчас не болеют им», — добавил он.

Comparison (same FLORES n=20)

Model Translate AVG chrF++/BLEU Supervisor-12
This (9B-Translate Q4) 43.8 / 16.5 5/12 (not a supervisor)
NOESIS-4B-LongCtx Q8 41.7 / 16.5 11/12
base 4B Q8 42.4 / 15.4 5/12

Best pure translator (marginal: +2.1 chrF++ over the 4B, BLEU tied) but NOT a supervisor. Use 9B for the final translation pass when VRAM allows; the 4B-LongCtx is the all-rounder.

Speed (RTX 3060 Laptop 6GB, GPU, 33/33 layers offloaded)

  • Q4_K_M: gen 49.1 tok/s, prompt eval 307 tok/s. (vs 4B-LongCtx Q8: 53.5 / 366 — the 4B is ~9% faster despite the 9B being lighter-per-param in Q4.)

⚠️ n=20 quick estimate (partly within noise). Eval on GPU via llama-completion.exe -ngl 99. Written: 2026-06-17

MT benchmark — FLORES-200 devtest (2026-06-17)

Real eval (not smoke): n=100 × 4 directions (eng↔rus, eng↔cmn), GPU via resident llama-server -ngl 99. Primary metric COMET (wmt22-comet-da, neural — how "best translator" is judged), plus chrF++ / BLEU / length-ratio. Each model prompted in its own native format (MT2 = dubbing ChatML "SOURCE (lang): … Только перевод"; 9B = ChatML + no-think). Data + COMET checkpoint: D:/models/by_expert/07_MT_TRANSLATION.

Model Size COMET avg chrF++ BLEU gen tok/s
Qwopus3.5-9B-Translate Q4 5.24 GB 0.8870 50.7 22.5 49
NOESIS-Hy-MT2-7.5B Q5 5.0 GB 0.8709 46.2 21.4 52
NOESIS-Hy-MT2-1.8B Q8 1.78 GB 0.8481 43.9 19.1 121

Per-direction COMET — 9B-Translate wins all 4 (eng-rus .902 / eng-cmn .897 / rus-eng .872 / cmn-eng .877); MT2-7.5B 2nd, MT2-1.8B 3rd.

Notes:

  • MT2 is a dubbing translator (isochrony): its outputs are shorter (len_ratio ~0.87-0.89 vs 9B ~1.0) because it compresses to fit speech slots → lower chrF on literal FLORES news. FLORES does NOT measure MT2's slot-fit strength, so it under-rates MT2 for its actual job.
  • 1.8B→7.5B degradation: COMET +0.023, chrF +2.3, BLEU +2.3 — modest; 1.8B is 2.4× faster and 2.8× smaller (good lightweight tradeoff).
  • BLEU for eng-cmn is low for all (Chinese needs char-tokenization); use chrF++/COMET there.
Downloads last month
108
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including AMAImedia/NOESIS-Qwopus3.5-9B-Translate-v3.5-BF16